EDBT 2026 Demo / reviewers in the wild / expert
Rui Wang 0014
dblp:w/RuiWang14
· DBLP profile ↗
54ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-2741-6033ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorComputer networks · 4Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Parameter Tuning for Multi-FPGA Partitioning: A Preference-Guided ApproachabstractParameter tuning for multi-FPGA partitioning algorithms represents a bottleneck in modern chip emulation and verification workflows. Current multilevel partitioning tools require manual configuration of various parameters, where each evaluation can take tens of seconds to minutes, making exhaustive search impractical and expert-driven tuning both time-consuming and suboptimal. To automate this process, we propose a preference-guided Bayesian optimization framework specifically designed for industrial FPGA partitioning parameter tuning under limited evaluation budgets. Our approach maximizes the minimum timing slack by incorporating domain-specific insights: we exploit the strong correlation between cutsize and timing performance through a priority-based ranking scheme that guides a pairwise Gaussian process to learn configuration preferences. Additionally, we introduce a kernel input transformation that properly handles the mixed discrete-continuous parameter space typical in EDA tools. Our method converges faster with fewer evaluations and achieves the best timing slack in 60–70% of cases on industrial circuit benchmarks compared to existing methods including standard Bayesian optimization, quasi-random sampling, and state-of-the-art preference learning techniques. The proposed framework reduces parameter tuning from days of manual effort to hours of automated optimization, offering practitioners a deployment-ready solution that improves both design quality and engineering productivity. Yutao Dai, Shengbo Tong, Chunyan Pei, Zhuohua Liu, Yi Liu 0013, Rui Wang 0014, Wenjian Yu |
ASP-DAC | 7 |
| 2026 | TC_SpGEMM: High Performance Sparse General Matrix Multiplication with Tensor Core-Accelerated
FuKai Sun, Xing Cong, Chenhao Xie 0001, Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
IPDPS | 4 |
| 2026 | Folivora: Ultralow-Power Microprocessor Design With Nano-Electromechanical Relay and Nanotube MemoryabstractIn post-Moore era, CMOS technology scaling has encountered enormous design and fabrication challenges. “Power Wall” limits the further increase of integration density. Emerging AI computing and data center deployments aggravate the power consumption problem further. In the pursuit of efficient computing paradigm, Nano Electro-mechanical (NEM) relay and Nanotube Random Access Memory (NRAM) technology have attracted enormous attention and have ultra-low power consumption compared to CMOS counterparts. NEM relay is a kind of device based on electronic and mechanical interaction switching, characterized by remarkably low power consumption. This article explores the application of NEM relay and NRAM technology to build a complex RISC processor, aiming to achieve much lower power without degrading performance. The controller and data path can be implemented with primitive logic gates made of NEM relays, and on-chip cache can be implemented with NRAM. Experimental results show that the energy efficiency of the processor design based on NEM relay and NRAM can be improved by 88.2% and 78.9% compared with CMOS technology based in-order and out-of-order microprocessors, respectively. Meanwhile, the performance can be improved by 42.9% and the instruction execution time can be reduced by more than 17.9%, which implies the potentials of NEM relay and NRAM for emerging ultra-low power applications. Yuanqing Cheng, Ying Wang 0001, Rui Wang 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | MultiSky: Dynamic Resource Allocation Framework for High-Throughput CGRA Multitask ExecutionabstractCoarse-grained reconfigurable arrays (CGRAs) offer a promising balance between high performance and flexibility, yet dynamic resource allocation in multi-task scenarios remains challenging due to unpredictable task creation/destruction. Existing static approaches lack flexibility, while dynamic methods suffer from high latency or limited applicability. This paper presents MultiSky, a framework for CGRA multi-task dynamic resource allocation, combining a hardware controller and a software pre-mapper. The hardware controller dynamically allocates resources within hundreds of cycles by calculating tile allocation for each task via weighted averaging, and generating tile shapes using a lightweight heuristic algorithm. The software pre-mapper employs incremental compilation to pre-generate configurations, avoiding online transformation overhead. Evaluations on a real-world multi-task scenario demonstrate that MultiSky achieves 1.72× higher throughput than baselines by maintaining 82.7% average resource utilization. The framework scales efficiently with larger CGRAs and task counts, with hardware overhead decreasing to 1% for 16×16 CGRAs. These results highlight MultiSky’s ability to balance flexibility, efficiency, and practicality in dynamic computing environments. Chenhao Xie 0001, Rui Wang 0014, Liansheng Liu, Xiyuan Peng, Yu Peng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | QDLoRA: Enhanced LoRA Fine-Tuning on Quantized LLMs via Integrated Low-Rank Decomposition
Xingyi Su, Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
APPT | 2 |
| 2025 | MEFusion: Memory-Efficient Data Fusion for Real-Time 3D Reconstruction On Resource-Constrained DevicesabstractOnline semantic 3D modeling from streaming RGB-D data fundamentally requires consistent fusion of 2D segmentation. Popular approaches address segmentation inconsistencies through histogram-based label aggregation, where each 3D element (point/voxel) maintains the frequency of candidate labels, which introduces prohibitive memory and computational overhead for resource-constrained devices. In response to this challenge, we propose MEFusion, a memory-efficient probabilistic fusion framework to avoid element-wise histogram aggregation. Specifically, we propose an element-wise probability update algorithm based on Bayesian Estimation, where each voxel stores only one instance label and updates it based on a posterior probability to maintain segmentation consistency. Following 3D segmentation, we establish a segment-wise voting framework to aggregate the semantic labels from historical data, where co-segment voxels share the semantic voting histogram, for semantic consistency. Our experiments demonstrate that our method achieves a memory reduction of 77%(85%) and a speed improvement of 58%(6.12x) on the desktop (embedded) platform while maintaining comparable reconstruction accuracy to the state-of-the-art point-cloud-based method. Ruizhi Cao, Rui Wang 0014, Yu Wen 0003, Chenhao Xie 0001 |
IROS | 2 |
| 2025 | Towards Efficient LLM Inference via Collective and Adaptive Speculative DecodingabstractLarge language models (LLMs) have gained considerable attention for their remarkable performance across a wide range of tasks. However, efficient LLM inference remains challenging because of the autoregressive decoding process, which generates only one token at a time. Speculative decoding has been introduced to address the limitation by using small speculative models (SSMs) to speed up LLM inference. However, the low acceptance rate of SSMs and the high verification cost of LLM prohibit further performance improvement. In this paper, we present Smurfs, an LLM inference system designed to accelerate LLM inference through collective and adaptive speculative decoding. Smurfs adopts a majority-voted mechanism that harnesses multiple SSMs to collaboratively predict LLM outputs in multi-task scenarios, while avoiding high verification cost. It also decouples SSM speculation from LLM verification and uses a pipelined execution to hide the latency of SSM speculation. Additionally, Smurfs proposes a mechanism to dynamically determine the optimal speculation length of SSM at runtime, balancing the performance impact of accepted tokens and verification cost. The experimental results demonstrate the superiority of Smurfs in terms of inference throughput and latency compared to the state-of-the-art LLM inference systems. Hailong Yang 0002, Tongxuan Liu, Yufan Xu 0001, Xuning Liang, Kejie Ma, Tianyu Feng, Xin You 0001, Ruihao Gong, Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
SC | 12 |
| 2025 | Testing and fault tolerance techniques for carbon nanotube-based FPGAs
Kangwei Xu, Rui Wang 0014, Yuanqing Cheng |
Integr. | 4 |
| 2025 | Deep Learning Operators Performance Tuning for Changeable Sized Input Data on Tensor Accelerate HardwareabstractThe operator library is the fundamental infrastructure of deep learning acceleration hardware. Automatically generating the library and tuning its performance is promising because the manual development by well-trained and skillful programmers is costly in terms of both time and money. Tensor hardware has the best computing efficiency for deep learning applications, but the operator library programs are hard to tune because the tensor hardware primitives have many limitations. Otherwise, the performance is difficult to be fully explored. The recent advancement in LLM exacerbates this problem because the size of input data is not fixed. Therefore, mapping the computing tasks of operators to tensor hardware units is a significant challenge when the shape of the input tensor is unknown before the runtime. We propose DSAT, a deep learning operator performance autotuning technique for changeable-sized input data on tensor hardware. To match the input tensor's undetermined shape, we choose a group of abstract computing units as the basic building blocks of operators for changeable-sized input tensor shapes. We design a group of programming tuning rules to construct a large exploration space of the variant implementation of the operator programs. Based on these rules, we construct an intermediate representation of computing and memory access to describe the computing process and use it to map the abstract computing units to tensor primitives. To speed up the tuning process, we narrow down the optimization space by predicting the actual hardware resource requirement and providing an optimized cost model for performance prediction. DSAT achieves performance comparable to the vendor's manually tuned operator libraries. Compared to state-of-the-art deep learning compilers, it improves the performance of inference by 13% on average and decreases the tuning time by an order of magnitude. Pengyu Mu, Yi Liu 0013, Rui Wang 0014, Hangcheng An, Qianhe Zhao, Hailong Yang 0002, Chenhao Xie 0001, Zhongzhi Luan, Chunye Gong, Depei Qian 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Sifter: An Efficient Operator Auto-Tuner With Speculative Design Space Exploration for Deep Learning CompilerabstractDeep learning compiler can automatically optimize operators. It provides higher flexibility compared to vendor libraries. However, existing DNN operator tuning methods mostly rely on search-based approaches, which still face challenges such as large design spaces and long tuning times. To address these issues, we propose Sifter, an efficient DNN operator auto-tuner with speculative design space exploration. By training and analyzing decision trees, we extract shared characteristics of high-quality schedules and summarize them as pruning rules. Applying these rules during the optimization allows us to speculatively explore the design space, minimize unnecessary hardware measurements, and shorten the optimization time without compromising the optimization result. We conducted experiments on three different platforms with various operators and models. The results demonstrate that Sifter reduces 52% of redundant schedules and shortens the optimization time by 41% while maintaining operator optimization performance at the state-of-the-art level. Qianhe Zhao, Rui Wang 0014, Yi Liu 0013, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
IEEE Trans. Computers | 2 |
| 2024 | ViTa: Optimizing Vision Transformer Operators on Domain-Specific AcceleratorsabstractDomain-specific devices accelerate the inference of deep learning models, whose performance is sensitive to hardware resources. Large artificial intelligence models, especially those involving self-attention, have extensive computational demands that challenge efficient execution on accelerators. Due to restricted search space on resource-constrained hardware, some approaches require significant engineering effort to develop platform-specific optimization code or find suboptimal programs with search-based automatic compilers.This paper proposes ViTa, a framework for deploying vision Transformer models on resource-constrained hardware. First, ViTa adopts an analytical approach to modify the model structure, reducing computational load without sacrificing accuracy. Secondly, we divide computations into tiles and map them to computing units to enhance memory access and computational efficiency. Providing an abstract hardware layer to guide tiling and mapping can improve the model deployment efficiency. Experiments demonstrate that ViTa can reduce memory bandwidth by 43% and FLOPs by 56% while maintaining accuracy. Compared with the Vendor library, the inference time on various accelerators is reduced by an average of 24%. Pengyu Mu, Yi Liu 0013, Rui Wang 0014 |
ISPA | 4 |
| 2023 | HAOTuner: A Hardware Adaptive Operator Auto-Tuner for Dynamic Shape Tensor CompilersabstractDeep learning compilers with auto-tuners have the ability to generate high-performance programs, particularly tensor programs on accelerators. However, the performance of these tensor programs is shape-sensitive and hardware resource-sensitive. When the tensor shape is only known at runtime instead of compile time, auto-tuners must tune the tensor programs for every possible shape, leading to significant time and cost overhead. Additionally, if a tensor program tuned for one device is deployed on a different device, the performance may not be as optimal as before. To address these challenges, we propose HAOTuner, a hardware-adaptive deep learning operator auto-tuner specifically designed for dynamic shape tensors. We leverage the concept of micro-kernels as the unit of task allocation and have observed that the size of the micro-kernel greatly impacts performance. In HAOTuner, we determine the size of micro-kernels based not only on the tensor shapes but also on the available hardware resources. Specifically, we present an algorithm to select hardware-friendly micro-kernels as candidates, reducing the tuning time. We also design a cost model that is sensitive to hardware resources to support various hardware architectures. Furthermore, we provide a model transfer solution to enable fast deployment of the cost model on different hardware platforms. We evaluate HAOTuner on six different types of GPUs. The experiments demonstrate that HAOTuner surpasses the state-of-the-art dynamic shape tensor auto-tuner in terms of running time by an average of 26% and tuning time by 25%. Moreover, HAOTuner outperforms the state-of-the-art compiler with padding in terms of running time by an average of 39% and tuning time by 6×. Pengyu Mu, Yi Liu 0013, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
IEEE Trans. Computers | 3 |
| 2022 | Passive Motion Detection via mmWave Communication SystemabstractIn this paper, an integrated passive sensing and communication system working in 60 GHz band is elaborated, and the sensing performance is investigated in an application of hand gesture recognition. Specifically, in this integrated system, there are two radio frequency (RF) chains at the receiver and one at the transmitter. Each RF chain is connected with one phased array for analog beamforming. To facilitate simultaneous sensing and communication, the transmitter delivers one stream of information-bearing signals via two beam lobes, one is aligned with the main signal propagation path and the other is directed to the sensing target. Signals from the two lobes are received by the two RF chains at the receiver, respectively. By cross ambiguity coherent processing, the time-Doppler spectrograms of hand gestures can be obtained. Relying on the passive sensing system, a dataset of received signals, where three types of hand gestures are sensed, is collected by using Line-of-Sight (LoS) and Non-Lineof-Sight (NLoS) paths as the reference channel respectively. Then a neural network is trained by the dataset for motion detection. It is shown that the classification accuracy rate is high as long as sufficient sensing time is assured. Finally, an empirical model characterizing the relation between the classification accuracy and sensing duration is derived analytically. Chao Yu 0001, Yifei Sun 0003, Rui Wang 0014 |
VTC Spring | 5 |
| 2021 | Mutual calibration training: Training deep neural networks with noisy labels using dual-models
Rui Liu 0020, Yi Liu 0013, Rui Wang 0014, Yucong Zhou |
Comput. Vis. Image Underst. | 3 |
| 2021 | Guardauto: A Decentralized Runtime Protection System for Autonomous DrivingabstractDue to the broad attack surface and the lack of runtime protection, potential safety and security threats hinder the real-life adoption of autonomous vehicles. Although efforts have been made to mitigate some specific attacks, there are few works on the protection of the autonomous driving system, i.e., the control software system performing such as perception, decision making, and motion tracking. This article presents a decentralized self-protection framework called Guardauto to protect the autonomous driving system against runtime threats. First, Guardauto proposes an isolation model to decouple the autonomous driving system and isolate its components with a set of partitions. Second, Guardauto provides self-protection mechanisms for each target component, which combines different methods to monitor the target execution and plan adaption actions accordingly. Third, Guardauto provides cooperation among local self-protection mechanisms to identify the root-cause component in the case of cascading failures affecting multiple components. A prototype has been implemented and evaluated on the open-source autonomous driving system Autoware. Results show that Guardauto could effectively mitigate runtime failures and attacks, and protect the control system with acceptable performance overhead. Yuan Zhou 0005, Bihuan Chen 0001, Rui Wang 0014, Yuebin Bai, Yang Liu 0003 |
IEEE Trans. Computers | 4 |
| 2021 | MIPSGPU: Minimizing Pipeline Stalls for GPUs With Non-Blocking ExecutionabstractImproving the latency hiding ability is important for GPU performance. Although existing works, which mainly target on either improving thread level parallelism or optimizing memory hierarchy, are effective at improving GPUs’ latency hiding ability, warps are still blocked after executing long latency operations, reducing the number of schedulable warps. This article revisits the recently proposed non-blocking execution for GPUs to improve the latency hiding ability of GPUs. With non-blocking execution, instructions from warps blocked by long latency operations can be pre-executed to make full use of GPU resources. However, we find that the state-of-the-art non-blocking GPU architecture gains limited performance improvement. Through in-depth analysis, we observe that the poor performance is largely due to inefficient pre-execution state management, duplicate instruction extraction, frequent early eviction and severe resource congestion. To make non-blocking execution actually useful for GPUs and minimize hardware overheads, we carefully redesign the non-blocking architecture for GPUs based on our analysis and proposeMIPSGPU. Our evaluations show thatMIPSGPU, relative to the state-of-the-art non-blocking GPU architecture, improves performance of memory intensive applications by 19.05 percent, and reduces memory to SM traffics by 14 percent. Chao Yu 0001, Yuebin Bai, Rui Wang 0014 |
IEEE Trans. Computers | 3 |
| 2020 | Extremely Low-bit Convolution Optimization for Quantized Neural Network on Modern Computer ArchitecturesabstractWith the continuous demand for higher accuracy of deep neural networks, the model size has increased significantly. Quantization is one of the most widely used model compression methods, which can effectively reduce the model size without severe accuracy loss. Modern processors such as ARM CPU and NVIDIA GPU have already provided the support of low-bit arithmetic instructions. However, there lack efficient and practical optimizations for convolution computation towards extremely low-bit on ARM CPU (e.g., 2 ∼ 8-bit) and NVIDIA GPU (e.g., 4-bit and 8-bit). This paper explores the performance optimization methods of extremely low-bit convolution on diverse architectures. On ARM CPU, we propose two instruction schemes for 2 ∼ 3-bit and 4 ∼ 8-bit convolution with corresponding register allocation methods. In addition, we re-design the GEMM computation with data padding and packing optimizations. We also implement winograd algorithm for convolution with some specific bit width (e.g., 4 ∼ 6-bit) to achieve higher performance. On NVIDIA GPU, we propose a data partition mechanism and multi-level memory access optimizations, to better adapt the computation to GPU thread and memory hierarchy. We also propose quantization fusion to eliminate unnecessary data access. The experiment results demonstrate our implementations achieve better performance of extremely low-bit convolution compared to the state-of-the-art frameworks and libraries such as ncnn and cuDNN. To the best of our knowledge, this is the first work that provides efficient implementations of extremely low-bit convolutions covering 2 ∼ 8-bit on ARM CPU and 4-bit/8-bit on NVIDIA GPU. Qingchang Han, Yongmin Hu, Fengwei Yu, Hailong Yang 0002, Ruihao Gong, Rui Wang 0014, Zhongzhi Luan, Depei Qian 0001 |
ICPP | 9 |
| 2020 | Temperature-Aware DRAM Cache Management - Relaxing Thermal Constraints in 3-D SystemsabstractHigh bandwidth 3-D-stacked dynamic random access memory (DRAM) has been proposed to address the memory wall in modern systems, especially when it is used as a large last-level cache (LLC). However, stacking DRAM directly on top of the processor significantly impedes the efficiency of cooling, potentially causing thermal issues both in the processor and DRAM. Dynamic thermal management (DTM) based on DRAM temperature can be heavily intrusive because the normal working temperature for DRAM is lower than the processor temperature limit. This paper shows that in many cases it is better to disable hot portions of the cache rather than apply DTM and slow down the processor. Three temperature-aware cache management mechanisms are proposed to decrease the performance impact of DTM on 3-D systems. Our experiments show these techniques can improve the performance of DRAM-targeted DTM by 26.1% on average which make 3-D systems more practical for the future high-performance computing. Minxuan Zhou, Andreas Prodromou, Rui Wang 0014, Hailong Yang 0002, Depei Qian 0001, Dean M. Tullsen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Thread-Level Locking for SIMT ArchitecturesabstractAs more emerging applications are moving to GPUs, thread-level synchronization has become a requirement. However, GPUs only provide warp-level and thread-block-level rather than thread-level synchronization. Moreover, it is highly possible to cause live-locks by using CPU synchronization mechanisms to implement thread-level synchronization for GPUs. In this article, we first propose a software-based thread-level synchronization mechanism called lock stealing for GPUs to avoid live-locks. We then describe how to implement our lock stealing algorithm in mutual exclusive locks and readers-writer locks with high performance. Finally, by putting it all together, we develop a thread-level locking library (TLLL) for commercial GPUs. To evaluate TLLL and show its general applicability, we use it to implement six widely used programs. We compare TLLL against the state-of-the-art ad-hoc GPU synchronization, GPU software transactional memory (STM), and CPU hardware transactional memory (HTM), respectively. The results show that, compared with the ad-hoc GPU synchronization for Delaunay mesh refinement (DMR), TLLL improves the performance by 22 percent on average on a GTX970 GPU, and shows up to 11 percent of performance improvement on a Volta V100 GPU. Moreover, it significantly reduces the required memory size. Such low memory consumption enables DMR to successfully run on the GTX970 GPU with the 10-million mesh size, and the V100 GPU with the 40-million mesh size, with which the ad-hoc synchronization can not run successfully. In addition, TLLL outperforms the GPU STM by 65 percent, and the CPU HTM (running on a Xeon E5-2620 v4 CPU with 16 hardware threads) by 43 percent on average. Lan Gao 0004, Rui Wang 0014, Zhongzhi Luan, Zhibin Yu 0001, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | LADet: A Light-weight and Adaptive Network for Multi-scale Object DetectionabstractScale variation is one of the most significant challenges for object detection task. In comparison with previous one-stage object detectors that simply make feature pyramid network deeper without consideration of speed, we propose a novel one-stage object detector called LADet, which consists of two parts, Adaptive Feature Pyramid Module(AFPM) and Light-weight Classification Function Module(LCFM). Adaptive Feature Pyramid Module generates complementary semantic information for each level feature map by jointly utilizing multi-level feature maps from backbone network, which is different from the top-down manner. Light-weight Classification Function Module is able to exploit more type of anchor boxes without a dramatic increase of parameters because of the utilization of interleaved group convolution. Extensive experiments on PASCAL VOC and MS COCO benchmark demonstrate that our model achieves a better trade-off between accuracy and efficiency over the comparable state-of-the-art detection methods. Yuqiao Tian, Weicheng Li, Rui Wang 0014, Zhongzhi Luan, Depei Qian 0001 |
ACML | 4 |
| 2019 | GraphQ: Scalable PIM-Based Graph ProcessingabstractProcessing-In-Memory (PIM) architectures based on recent technology advances (e.g., Hybrid Memory Cube) demonstrate great potential for graph processing. However, existing solutions did not address the key challenge of graph processing---irregular data movements. Youwei Zhuo, Chao Wang 0051, Rui Wang 0014, Dimin Niu, Yanzhi Wang 0001, Xuehai Qian |
MICRO | 4 |
| 2019 | Multiple Algorithms Against Multiple Hardware Architectures: Data-Driven Exploration on Deep Convolution Neural Network
Chongyang Xu, Zhongzhi Luan, Lan Gao 0004, Rui Wang 0014, Lianyi Zhang, Yi Liu 0013, Depei Qian 0001 |
NPC | 4 |
| 2019 | A novel index system describing program runtime characteristics for workload consolidation
Lin Wang 0112, Depei Qian 0001, Rui Wang 0014, Zhongzhi Luan, Hailong Yang 0002, Huaxiang Zhang 0001 |
Frontiers Comput. Sci. | 3 |
| 2019 | Accelerating in-memory transaction processing using general purpose graphics processing units
Lan Gao 0004, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | Research on Asynchronous Inter-VM Communication Mechanism Based on Embedded HypervisorabstractVirtualization technology, which has achieved great success in server and desktop environments in the last few years, is currently extending itself towards a new territory: embedded system. OKL4 from Open Kernel Labs is a leading virtualization software for embedded systems. Its microkernel approach can improve plain virtualization technologies, but it also brings incomplete IPC (Inter-Process Communication) issues. This paper proposes a novel asynchronous communication mechanism that can generate multiple event channels and efficiently manage concurrent communication requests between virtual machines. This mechanism optimizes the original IPC mechanism, and this paper also proposes a shared-memorybased bulk data transmission mechanism. The final experiments prove its feasibility and demonstrate the specific effect. Rui Wang 0014, Libin Xu, Yuebin Bai, Zhongzhao Wang, Guangqiang Luan, Hailong Yang 0002 |
COMPSAC (1) | 1 |
| 2018 | Network Alarm Flood Pattern Mining Algorithm Based on Multi-dimensional AssociationabstractIn the process of network operation, a large number of alerts are generated every day, which reflect the occurrence of some abnormal conditions. Traditional methods depend too much on the knowledge of equipment manufacturers and industry experts, so we need some novel ways to overcome this problem in the network management. The application of data mining technology to alarm pattern analysis has become the focus of current research. Researchers developed many kinds of algorithms fitting different application characteristics. This paper proposes the concept of association matrix pattern mining, which means that before mining the data, we use the multi-dimensional information of the data to construct the association matrices between the items. And we develop a conditional pattern mining algorithm based on the association matrix which aims to find out less but more meaning results. Our experiments validate that with the multi-dimensional information stored in association matrix, the algorithm performs better than traditional pattern mining methods in finding out the detailed alarm pattern from network alarm flood. Yuebin Bai, Peng Feng 0003, Junfang Zeng, Rui Wang 0014 |
MSWiM | 8 |
| 2018 | Nodes contact probability estimation approach based on Bayesian network for DTNabstractDelay tolerant network (DTN) known as suffering from frequent disruption, high latency and heterogeneous, resulting in low network availability. To improve DTN availability, routing protocols typically need to predict the probability of encountering the nodes. In this paper, we use the Bayesian Network (BN) to construct the knowledge base, which is an unique tool for creating a representation of the dependence relationships among DTN parameters. Then developed a Bayesian network- based approach to estimate the contact probability among nodes of DTN. We conducted an experiment to compare our approach against its counterparts in PROPHET routing protocol and power law distribution-based method. The experiment shows our approach is superior to other methods in both recall ratio and precision in all four datasets, including HAGGLE, NUS, REALITY and SASSY. Yuebin Bai, Xu Shao, Wentao Yang 0001, Peng Feng 0003, Rui Wang 0014 |
NOMS | 8 |
| 2018 | A network traffic flow prediction with deep learning approach for large-scale metropolitan area networkabstractAccurate and timely internet traffic information is important for many applications, such as bandwidth allocation, anomaly detection, congestion control and admission control. Over the last few years, internet flow data have been exploding, and we have truly entered the era of big data. Existing traffic flow prediction methods mainly use simple traffic prediction models and are still unsatisfying for many real-world applications. This situation inspires us to rethink the internet traffic flow prediction problem based on deep architecture models with big traffic data. In this paper, we propose a novel deep-learning-based internet traffic flow prediction method, which is called SDAPM. It consider the spatial and temporal correlations inherently and internet flow data character. A stacked denoising autoencoder prediction model (SDA) is used to learn generic internet traffic flow features, and it is trained in a greedy layer-wise fashion. Moreover, experiments demonstrate that the SDAPM for traffic flow prediction has effective performance. Our prediction model is in production as part of the traffic scheduling system at China Unicom, one of the largest Internet companies in China, helping improving the network bandwidth utilization. Yuebin Bai, Chao Yu 0001, Yuhao Gu, Peng Feng 0003, Rui Wang 0014 |
NOMS | 7 |
| 2018 | T1000: Mitigating the memory footprint of convolution neural networks with decomposition and re-fusion
Changxi Liu, Hailong Yang 0002, Rui Wang 0014, Zhongzhi Luan, Depei Qian 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | SRAM- and STT-RAM-based hybrid, shared last-level cache for on-chip CPU-GPU heterogeneous architectures
Lan Gao 0004, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001, Jihong Cai |
J. Supercomput. | 2 |
| 2017 | Achieving Versatile and Simultaneous Cache Optimizations With Nonvolatile SRAMabstractThe efficiency of caches plays a vital role in microprocessors. In this paper, we introduce a novel and flexible cache substrate, which integrates nonvolatile memory devices into the standard SRAM cells. By allowing this nonvolatile SRAM (NV-SRAM) cell to store inconsistent data between SRAM portion and NV portion, we show that the proposed NV2-SRAM cache not only provides enriched functionalities, but also allows simultaneous multiple optimizations. For example, the NV2-SRAM cache can reduce cache misses caused by context-switching and improve the performance by 15%. It can also save up to 67% energy over the SRAM-based cache, outperforming the drowsy cache in terms of both power efficiency and reliability. Moreover, the proposed cache architecture can be used to improve the performance of prefetching by 10%. Comparing with a conventional cache (equipped with a victim buffer) that occupies the same die area, the NV2-SRAM cache gains an 11% performance benefit. To achieve simultaneous optimizations, we propose architecture and OS support to optimize the cache power, performance and reliability concurrently on multicore-based systems. Rui Wang 0014, Dan Jia, Tao Li 0006, Depei Qian 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Scheduling Tasks with Mixed Timing Constraints in GPU-Powered Real-Time SystemsabstractDue to the cost-effective, massive computational power of graphics processing units (GPUs), there is a growing interest of utilizing GPUs in real-time systems. For example GPUs have been applied to automotive systems to enable new advanced and intelligent driver assistance technologies, accelerating the path to self-driving cars. In such systems, GPUs are shared among tasks with mixed timing constraints: real-time (RT) tasks that have to be accomplished before specified deadlines, and non-real-time, best-effort (BE) tasks. In this paper, (1) we propose resource-aware non-uniform slack distribution to enhance the schedulability of RT tasks (the total amount of work of RT tasks whose deadlines can be satisfied on a given amount of resources) in GPU-enabled systems; (2) we propose deadline-aware dynamic GPU partitioning to allow RT and BE tasks to run on a GPU simultaneously, such that BE tasks are not blocked for a long time. Rui Wang 0014, Tao Li 0006, Mingcong Song, Lan Gao 0004, Zhongzhi Luan, Depei Qian 0001 |
ICS | 2 |
| 2016 | QIM: Quantifying Hyperparameter Importance for Deep Learning
Dan Jia, Rui Wang 0014, Cheng-Zhong Xu 0001, Zhibin Yu 0001 |
NPC | 2 |
| 2016 | Coordinating workload balancing and power switching in renewable energy powered data center
Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
Frontiers Comput. Sci. | 2 |
| 2016 | Managing Server Clusters on Renewable Energy MixabstractAs climate change has become a global concern and server energy demand continues to soar, many IT companies have started to explore server clusters running on various renewable energy sources. Existing green data center designs often yield suboptimal performance as they only look at a certain specific type of energy source. This article explores data centers powered by hybrid renewable energy systems. We propose GreenWorks, a framework for HPC data centers running on a renewable energy mix. Specifically, GreenWorks features a cross-layer power management scheme tailored to the timing behaviors and capacity constraints of different energy sources. Using realistic workload traces and renewable energy data, we show that GreenWorks could provide a near-optimal workload performance (within 3% difference) on average. It can also reduce the worst-case performance degradation by 43% compared to the state-of-the-art design. Moreover, the performance improvements are based on carbon-neutral operations and are not at the cost of significant efficiency degradation and reduced battery lifecycle. Our technique becomes more efficient when servers become more energy proportional and can effectively handle the ever-increasing depth of renewable power penetration in green data centers. Chao Li 0009, Rui Wang 0014, Depei Qian 0001, Tao Li 0006 |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2015 | Optimizing Soft Real-Time Scheduling Performance for Virtual Machines with SRT-XenabstractMultimedia applications are an important part of today's Internet. However, currently most virtualization solutions, including Xen, lack adequate support for soft real-time tasks. Soft real-time applications, e.g. media workloads, are impeded by components of virtualization, such as the increase of scheduling latency. This paper focuses on improving scheduling scheme to support soft real-time workloads in virtualization systems. In this paper, we present an enhanced scheduler SRT-Xen. SRT-Xen can promote the soft real-time domain's performance compared with Xen's existing scheduling. It focuses on not only bringing a new realtime-friendly scheduling framework with corresponding strategies but also improving the management of the virtual CPUs' queuing in order to implement a fair scheduling mechanism for both real-time and non-real-time tasks. Finally, we use PESQ (Perceptual Evaluation of Speech Quality) and other benchmarks to evaluate and compare SRT-Xen with some other works. The results show that SRT-Xen supports soft real-time domains well without penalizing non-real-time ones. Yuebin Bai, Rui Wang 0014 |
CCGRID | 3 |
| 2015 | An Efficient Transmission Method for Bulk Data Based on Network Coding in Delay Tolerant NetworkabstractWith nodes in Delay Tolerant Network(DTN) distributing sparsely and moving rapidly, they usually suffer from intermittent connections and communications, thus bringing about limited message forwarding opportunities. All these could lead to inefficient forwarding, low delivery, long latency and limited transmission capacity in performance. In this paper, efficient encoding and decision methods are presented and integrated into the DTN routing strategy. The custody-encoding-forwarding mode is designed by merging the random linear network coding into the DTN routing, together with the replica re-allocation and memory management, and built on that encoding scheme, the intra/inter flow adaptive collaborative network coding is elaborated to implement the fresh custody-decision-encoding-forwarding mode. A decision-making strategy based on Bayesian Network(BN) measures the ``degree'' in a specific generation in networks by taking comprehensive consideration of current and historical network conditions to enhance the network robustness and self-adaptivity. By evaluating the delivery, delay and overhead performance on ONE and MATLAB platforms, the effectiveness of proposed strategies is validated in the end. Wancheng Chen, Yuebin Bai, Jiaojiao Liang, Wenjia Liu, Rui Wang 0014, Xiaoyun Mo, Ziming Luo |
MSWiM | 5 |
| 2015 | Improving multiprocessor performance with fine-grain coherence bypass
Rui Wang 0014, Zhongzhi Luan, Xuehai Qian, Depei Qian 0001 |
Sci. China Inf. Sci. | 2 |
| 2015 | Reducing DRAM refreshing in an error correction manner
Danfeng Zhu, Rui Wang 0014, Yanjiang Wei, Depei Qian 0001 |
Sci. China Inf. Sci. | 2 |
| 2014 | Software Transactional Memory for GPU Architectures
Rui Wang 0014, Nilanjan Goswami, Tao Li 0006, Lan Gao 0004, Depei Qian 0001 |
CGO | 2 |
| 2014 | Speedup Critical Stage of Machine Learning with Batch Scheduling in GPU
Rui Wang 0014, Yanjiang Wei, Depei Qian 0001 |
NPC | 2 |
| 2014 | Towards Automated Provisioning and Emergency Handling in Renewable Energy Powered Datacenters
Chao Li 0009, Rui Wang 0014, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006, Depei Qian 0001 |
J. Comput. Sci. Technol. | 2 |
| 2014 | Lightweight dynamic partitioning for last-level cache of multicore processor on real system
Ludan Zhang, Yi Liu 0013, Rui Wang 0014, Depei Qian 0001 |
J. Supercomput. | 3 |
| 2013 | Interference-Aware Program Scheduling for Multicore Processors
Lin Wang 0112, Rui Wang 0014, Cuijiao Fu, Zhongzhi Luan, Depei Qian 0001 |
ICA3PP (1) | 2 |
| 2013 | M&C: A Software Solution to Reduce Errors Caused by Incoherent Caches on GPUs in Unstructured Graphic Algorithm
Rui Wang 0014, Zhongzhi Luan, Depei Qian 0001 |
ICA3PP (1) | 2 |
| 2013 | Chameleon: Adapting throughput server to time-varying green power budget using online learningabstractEco-friendly energy sources (i.e. green power) attract great attention as lowering computer carbon footprint has become a necessity. Existing proposals on managing green energy powered systems show sub-optimal results since they either use rigid load power capping or heavily rely on backup power. We propose Chameleon, a novel adaptive green throughput server. Chameleon comprises of multiple flexible power management policies and leverages learning algorithm to select the optimal operating mode during runtime. The proposed design outperforms the state-of-the-art approach by 13% on performance, improves system MTBF by 42%, and still maintains up to 95% green energy utilization. Chao Li 0009, Rui Wang 0014, Tao Li 0006, Nilanjan Goswami, Depei Qian 0001 |
ISLPED | 3 |
| 2012 | Lightweight Dynamic Partitioning for Last Level Cache of Multicore Processor on Real SystemabstractAs multi-core/many-core becomes the trend of processor architecture, conflict in shared cache has become more and more serious that restricts performance improvement of parallel program. Recent research has employed page coloring mechanism to realizing cache partitioning on real system for the purpose of decline shared cache conflict. However, page coloring-based cache partitioning has some side-effects, one is page coloring restricts memory space an application can allocate from which may lead to memory pressure, another is changing cache partition dynamically need massive page copying which will incur large overhead and may go against with application's performance. To make page coloring based cache partition more practical, we proposed a malloc allocator based dynamic cache partitioning mechanism with page coloring. Memory allocated by our malloc allocator can be partitioned among different applications according to the cache partitioning policy. Our partition policy is based on a type recognition approach. Cache partition can be adjusted at run-time by changing the color of the pages allocated by the malloc allocator. Only coloring the dynamic allocated pages can remission memory pressure and reduce page copying overhead lead by re-coloring compared to all-page coloring. To further alleviate the overhead, we introduced minimum distance page copying strategy and lazy flush strategy. These policies yield performance improvements for co-running applications as high as 14.28% through cache partitioning and reduce the overhead of re-coloring by 55% on average when partitioning frequency is high. Our results demonstrate that only partitioning the dynamically allocated memory can reach the purpose of reducing cache conflict miss and the minimum distance page copying strategy is more beneficial to application with larger data-set and shorter data reuse distance. Ludan Zhang, Yi Liu 0013, Rui Wang 0014, Depei Qian 0001 |
PDCAT | 3 |
| 2012 | MANET adaptive structured P2P overlay
Nadir Shah, Depei Qian 0001, Rui Wang 0014 |
Peer-to-Peer Netw. Appl. | 3 |
| 2011 | Enhancing cooperation with multiple stage auctions in opportunistic routing for wireless mesh networksabstractOpportunistic routing significantly increases throughput in wireless mesh networks (WMNs) by utilizing the wireless broadcast medium. Most opportunistic routing protocols assume all nodes are cooperative. But in WMNs, one realistic problem is that nodes perform in their own interests and refuse to forward packets for other nodes. Game theory approach is an efficient way used in routing protocol to incentive nodes to forward other's packets. In this paper, we develop an auction incentive mechanism (AIM) for WMNs with opportunistic routing to encourage cooperation and balance energy consumption between nodes. In AIM, a fair pricing mechanism is used to incentive nodes and the pricing process is modeled as auction game reached Bayesian Nash equilibrium which maximizes the profit of each relay node. The energy status and throughput are considered in the bidding process; this not only ensures the high throughput but also balances the energy cost to reduce invalid nodes. Simulations are presented to complement our theoretical and evaluation results demonstrate its high performance in terms of stability, throughput and fairness. Rui Wang 0014, Depei Qian 0001, Zhongzhi Luan |
Integrated Network Management | 2 |
| 2010 | A Fair Thread-Aware Memory Scheduling Algorithm for Chip Multiprocessor
Danfeng Zhu, Rui Wang 0014, Depei Qian 0001, Zhongzhi Luan |
ICA3PP (1) | 2 |
| 2009 | Optimizing Transmission in Multi-Flow Streaming Overlay NetworksabstractAiming at improving the performance of relay transmission in multi-flow streaming overlay networks, a performance evaluation model to optimize the global weighted average latency was proposed. This model supports multiple senders and multiple receivers. A distributed heuristic allocation algorithm was proposed to optimize the transmission performance. In this algorithm, the bottleneck bandwidth shared by multiple flows is allocated based on the path weighted average latency which is computed between any pair of nodes in the network. Source node uses a two stages feedback pattern to allocate the outgoing flow to multiple available paths. Simulation results show that there is linear dependence between the weighted average latency and the data slip hit ratio which is a metric of streaming network performance, the heuristic allocation algorithm can effectively lower the network overall transmission latency and can effectively adjust the traffic allocation. It gets 4%-17% lower in the network overall weighted average latency comparing with that of the pattern of average allocation shared bottle bandwidth. Rui Wang 0014, Depei Qian 0001, Danfeng Zhu, Qinglin Zhu, Zhongzhi Luan |
NPC | 1 |
| 2009 | Re-exploring the Potential of Using Tree Structure in P2P Live Streaming NetworksabstractThe current peer-to-peer (P2P) live streaming networks can be generally classified into two categories: tree-based and data-driven. The tree-based approach suffers from three limitations: interruptive delivery due to failures of high level nodes, unfair uploading (out-going) bandwidth utilization in leaf nodes and bandwidth bottleneck in nodes near the root. The data driven approach has been widely studied recently to tackle the defects of the tree-based approach mentioned above. However the tree-based approach still has its advantages: deterministic delivery path length and predictable delay, and natural support to PUSH mode content delivery. Because of these advantages the tree-based approach will not be simply replaced by the data-driven approach. Based on this consideration, we propose a cluster-based approach to remedy the disadvantages of the normal tree-based approach and meanwhile retain its advantages as much as possible. By grouping peers into clusters, the content delivery tree constructed by clusters can maintain a stable overlay structure and transmission direction in a dynamic network environment. Simulation results show that our approach can effectively overcome the shortages of the single tree-based approach and outperform the data-driven approach in terms of deterministic content delivery path and predictable path length. Qinglin Zhu, Rui Wang 0014, Depei Qian 0001 |
NPC | 2 |
| 2008 | An evolutionary node architecture and performance optimizationabstractThe boom of Internet applications has resulted in ever increasing demands for new network services support. Different applications require different QoS, protocols and security mechanisms to be deployed in the network. These demand the network to evolve to keep pace with the change of application requirements. As the current Internet’s functions are very hard to expand and cope with the changing environment. A new evolutionary node architecture is proposed in this paper. New network services, protocols, control and management functions can be dynamically deployed on the ENN node so that the network can exhibit a flexible and application sensitive behavior. Major technical issues related to ENN are discussed.A prototyping evolutionary network called FAN is designed and implemented. The performance of the mobile agent execution environment was analyzed, evaluated and optimized. Experiments on FAN show that the ENN architecture is effective in promoting the evolution of the network. Tao Liu 0033, Depei Qian 0001, Yongxiang Huang, Ying He 0002, Rui Wang 0014 |
AICCSA | 5 |
| 2008 | An Architecture for Distributed Controllable Networks and Manageable Node Based on Network Processor
Tao Liu 0033, Depei Qian 0001, Yongxiang Huang, Rui Wang 0014 |
APWeb | 4 |