VLDB 2026 Research / reviewers in the wild / expert
Guibin Wang
dblp:59/6714
· DBLP profile ↗
31ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 9 since 2021Systems, architecture and hardware · 12 · 5 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic disentanglement: A contrastive causal framework for deepfake detection
Guibin Wang |
Comput. Vis. Image Underst. | 2 |
| 2026 | Real-Time Resilient Power System Operation With Defender-Attacker Soft Actor-Critic Reinforcement LearningabstractThreatened by weather disasters and operational uncertainties, power systems require resilient and cost-effective decision making to ensure security. This article proposes a novel deep reinforcement learning algorithm, namely defender–attacker soft actor-critic (DA-SAC), designed for contingency-constrained optimal power flow underN–ksecurity criteria. A two-agent Markov decision process is formulated, where the defender learns robust control actions and the attacker identifies worst-case contingencies. The core soft actor-critic algorithm is enhanced by integrating constraint violation levels into the reward function and employing a two-timescale learning scheme to improve feasibility and stability. The proposed method is validated on the IEEE 30-bus and 118-bus systems. Simulation results show that DA-SAC significantly reduces unserved energy, load shedding, and constraint violations, outperforming conventional and deep-reinforcement-learning-based benchmarks underN–1,N–2, andN–3scenarios. These results demonstrate that DA-SAC offers a fast, resilient, and practical solution for real-time power system operation under severe contingencies. Ka Wing Chan, Khaled Al Jaafari, Xian Zhang 0003, Guibin Wang, Ahmed Rabee Sayed |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Full-defense framework: multi-level deepfake detection and source tracingabstractDeepfake poses significant threats to various fields, including politics, journalism, and entertainment. Although many defense methods against deepfake have been proposed based on either passive detection or proactive defense, few have achieved both passive detection and proactive defense. To address this issue, we propose a full-defense framework (FDF) based on cross-domain feature fusion and separable watermarks (SepMark) to achieve copyright protection and deepfake detection, combining the ideas of passive detection and proactive defense. The proactive defense module consists of one encoder and two separable decoders, where the encoder embeds one watermark into the protected face, and two decoders separately extract two watermarks with different robustness. The robust watermark can reliably trace the trusted marked face while the semi-robust watermark is sensitive to malicious distortions that make the watermark disappear after deepfake or watermark removal attack. The passive detection module fuses spatial- and frequency-domain features to further differentiate between deepfake content and watermark removal attacks in the absence of watermarks. The proposed cross-domain feature fusion involves substituting the “secondary” channels of spatial-domain features with the “primary” channels of frequency-domain features. Subsequently, the “primary” channels of spatial-domain features are used to replace the “secondary” channels of frequency-domain features. Extensive experiments demonstrate that our approach not only offers proactive defense mechanisms by using extracted watermarks, i.e., source tracing and copyright protection, but also achieves passive detection when there are no watermarks, to further differentiate between deepfake content and watermark removal attacks, thereby offering a full-defense approach. Guibin Wang, Yanni Li, Rujia Qi |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | Competitive Pricing Strategy for the Wireless Charging Lane Operator Considering Range Anxiety of Electric Vehicle UsersabstractOn-road wireless charging is an emerging charging method, in addition to plug-in charging, that is a promising application in the future smart grid. Hence, in this paper, a competitive pricing strategy for the wireless charging lane (WCL) is proposed to maximize the economic benefits of the WCL operator. First, the competitive pricing strategy is formulated based on a non-cooperative game between the WCL operator and the charging station (CS). An iterative optimal pricing searching algorithm is developed to find the Nash equilibrium of the game. Second, a tri-level framework is established to derive the optimal competitive price considering the interaction among the WCL operator, the power distribution network (PDN) operator, and EV users. Third, the range anxiety of EV users is mathematically modeled based on Prospect theory. Numerical results indicate that the pricing strategy is effective in enhancing the attractiveness and profitability of the WCL operator. In addition, the utility of EV users is increased as well. Moreover, the PDN loss cost can be reduced, and downward voltage violation can be avoided. Shuying Lai, Zhao Yang Dong, Jing Qiu 0001, Yuechuan Tao, Junhua Zhao 0001, Guibin Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | An Efficient Gated-Attention Spatiotemporal Convolutional Network for Economical Operation of Electric Vehicle Charging StationsabstractThe rapid development of electric vehicles raises a higher requirement for charging station operation and management. Therefore, this work proposes a data-driven method aimed at enhancing the economical operation of charging stations. Considering the privacy of charging data and the influence of traffic flow, this data-driven method simulates the spatiotemporal charging demand based on the predicted traffic flow. To obtain precise prediction, this work develops an efficient gated-attention spatiotemporal convolutional network (GSTCN) to explore the long-term spatial and temporal dependence of traffic flow. GSTCN is constructed by two main components: a spatial gated attention (SGA) unit and a temporal gated (T-Gated) attention layer. The spatial pattern of traffic flow is unearthed by the well-designed SGA unit, while the temporal correlation between different time steps is captured through the proposed T-Gated attention layer. Then, an energy storage system (ESS) is employed to improve the effective management of charging stations. Numerical results demonstrate the efficiency of GSTCN in charging station management. GSTCN can produce more accurate prediction data, which leads to a more economical operation of the energy storage system compared with the benchmarks. Xian Zhang 0003, Guibin Wang, Fushuan Wen, Ziyuan Pu, Edward Chung 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | STPNet: Quantifying the Uncertainty of Electric Vehicle Charging Demand via Long-Term Spatiotemporal Traffic Flow Prediction IntervalsabstractConstructing the prediction intervals (PIs) for electric vehicle (EV) charging demand based on traffic flow information is crucial for the efficient operation of EV charging stations. However, due to the volatile nature of traffic flow, obtaining long-term traffic flow information (e.g., one week in advance) is challenging, particularly for multiple neighbored sites. To address this issue, this work establishes a deep learning prediction framework called Spatiotemporal Periodic Network (STPNet), which utilizes an encoder-decoder architecture. The STPNet incorporates a spatiotemporal and periodic pattern learning technique, and leverages the advantages of convolutional long short-term memory units (ConvLSTM) to quantify the uncertainty of traffic flow. Furthermore, to improve the performance of traffic flow prediction, a spatiotemporal series decomposition strategy based on Seasonal and Trend decomposition using Loess (STL) is employed, and a spatiotemporal PI performance-based loss function is creatively developed in this work. Then, the PIs of the EV charging demand are obtained based on the predicted traffic flow information and an M/M/C/K queuing model. Validated using a real-world dataset, the proposed model has been demonstrated to exhibit effectiveness in generating high-quality EV charging demand PIs for multiple locations. Songjian Chai, Xian Zhang 0003, Guibin Wang, Rongwu Zhu, Edward Chung 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Low-Carbon Charging Facilities Planning for Electric Vehicles Based on a Novel Travel Route Choice ModelabstractPromotion of charging facilities (CFs) can ameliorate range anxiety and facilitate long-distance travel by electric vehicles (EVs). In this context, a multistage low-carbon EV CFs planning model is proposed in this paper for the coupled transportation and power systems. This model not only takes into account the construction costs of newly-built CFs and their adverse impacts on the distribution system, but also considers the carbon emissions of CFs and the penalty due to the inconvenience for EVs to be recharged. In addition, a rational travel route choice model is significantly important to accurately evaluate the service capability of CFs to be constructed and thereby to obtain an optimal CF planning result. Therefore, a novel travel route choice model developed in this work allows EV drivers to take detours according to the CF locations to charge their EVs multiple times and then finish their trips. Carbon emission flow (CEF) model is innovatively employed to precisely calculate CFs’ carbon emission amount from the perspective of consumption side. Subsequently, uncertainties involved in CF planning, i.e., distribution and growth rate of traffic flow and electric load, popularity of different types of EVs, as well as location of the connected node of clean electricity, are fully considered to obtain a robust CF planning scheme that can achieve a good performance in the current stage and exhibit robustness for the uncertainties in the future stage. Finally, numerical experiments are conducted to verify the effectiveness of the proposed model. The impacts of future carbon price and EV cruising range on the planning results are also comprehensively evaluated. Ting Wu 0007, Guibin Wang, Xian Zhang 0003, Jing Qiu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Human-Machine Reinforcement Learning Method for Cooperative Energy ManagementabstractThe increasing penetration of distributed energy resources and a large volume of unprecedented data from smart metering infrastructure can help consumers transit to an active role in the smart grid. In this article, we propose a human-machine reinforcement learning (RL) framework in the smart grid context to formulate an energy management strategy for electric vehicles and thermostatically controlled loads aggregators. The proposed model-free method accelerates the decision-making speed by substituting the conventional optimization process, and it is more capable of coping with the diverse system environment via online learning. The human intervention is coordinated with machine learning to: 1) prevent the huge loss during the learning process; 2) realize emergency control; and 3) find preferable control policy. The performance of the proposed human-machine RL framework is verified in case studies. It can be concluded that our proposed method performs better than the conventional deep Q-learning and deep deterministic policy gradient in terms of convergence capability and preferable result exploration. Besides, the proposed method can better deal with emergent events, such as a sudden drop of photovoltaic (PV) output. Compared with the conventional model-based method, there are slight deviations between our method and the optimal solution, but the decision-making time is significantly reduced. Yuechuan Tao, Jing Qiu 0001, Shuying Lai, Xian Zhang 0003, Guibin Wang |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | Quantifying the Uncertainty in Long-Term Traffic Prediction Based on PI-ConvLSTM NetworkabstractThis work proposes a novel uncertainty quantification framework for long-term traffic flow prediction (TFP) based on a sequential deep learning model. Quantifying the uncertainty of TFP is crucial for intelligent transportation system (ITS) to make robust traffic congestion analysis and efficient traffic management due to the inherent uncertain and fluctuating nature of traffic flow. However, the performance (e.g., reliability and sharpness) of uncertainty quantification is hard to guarantee, particularly for long-term traffic flow (e.g., one week or two weeks in advance). To this end, this work develops a nonparametric performance-oriented prediction interval (PI) construction approach based on an enhanced sequential convolutional long short-term memory units (ConvLSTM) model, which is named as PI-ConvLSTM. This model can well learn the temporal correlations involved in the multivariate explanatory samples. Specifically, a periodic pattern learning strategy and a performance-oriented loss function are developed to ensure the quality of the derived PIs. Through validating on the real-life England freeway traffic flow dataset, the proposed PI-ConvLSTM proves to be capable of producing the skillful PIs for long-term TFP. For instance, the performance of derived PIs for two-week ahead is 0.175%, 0.198 and 1957.127 in average in terms of reliability, average width and sharpness, respectively. As compared to the benchmark models the proposed model shows at least 68.1% improvement on reliability, 3.4% on average width and 1.7% on sharpness. Songjian Chai, Guibin Wang, Xian Zhang 0003, Jing Qiu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Deep-Learning-Based Probabilistic Forecasting of Electric Vehicle Charging Load With a Novel Queuing ModelabstractWith the emerging electric vehicle (EV) and fast charging technologies, EV load forecasting has become a concern for planners and operators of EV charging stations (CSs). Due to the nonstationary feature of the traffic flow (TF) and the erratic nature of the charging procedures, EV charging load is difficult to accurately forecast. In this article, TF is first predicted using a deep-learning-based convolutional neural network (CNN), and different forecast uncertainties are evaluated to formulate the TF prediction intervals (PIs). Then, the EV arrival rates are calculated according to the historical data and the proposed mixture model. Based on TF forecasting and arrival rate results, the EV charging process is studied to convert the TF to the charging load using a novel probabilistic queuing model that takes into consideration charging service limitations and driver behaviors. The proposed models are assessed using the actual TF data, and the results show that the uncertainties of the EV charging load can be learned comprehensively, indicating significant potential for practical applications. Xian Zhang 0003, Ka Wing Chan, Hairong Li, Huaizhi Wang, Jing Qiu 0001, Guibin Wang |
IEEE Trans. Cybern. | 6 |
| 2021 | Extreme Learning Machine-Based State Reconstruction for Automatic Attack Filtering in Cyber Physical Power SystemabstractSuccessful detection of false data injection attacks (FDIAs) and removal of state bias due to FDIAs are essential for ensuring secure power grids operation and control. This article first extends the approximate dc model of FDIA to a more general ac model that can handle both traditional and synchronized measurements. To automatically filter out the established FDIAs, we propose a state reconstruction scheme consisting of a contaminated state separation method, an enhanced bad data identification approach and a state recovery algorithm. In this scheme, a classifier is developed by aggregating a series of extreme learning machines (ELMs) to detect anomaly states caused by FDIAs. Gaussian random distribution and Latin hypercube sampling are adopted to initialize the input weights of base ELMs, which can provide more diversities to enhance the ensemble performance. Then, to identify the exact locations of the compromised measurements, a state forecasting-based bad data identification approach is proposed by exploiting the consistency between the forecasted and the received measurements. Finally, an effective state recovery algorithm applies quasi-Newton method and Armijo line search to address the possible system unobservable problem due to the removal of attacked measurements. Numerical tests on serval IEEE standard test systems verify the efficiency of the proposed FDIA model and state reconstruction scheme. Ting Wu 0007, Wenli Xue, Huaizhi Wang, C. Y. Chung 0001, Guibin Wang, Jian-Chun Peng, Qiang Yang 0004 |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Deep Learning-Based Interval State Estimation of AC Smart Grids Against Sparse Cyber AttacksabstractDue to the aging of electric infrastructures, conventional power grid is being modernized toward smart grid that enables two-way communications between consumer and utility, and thus more vulnerable to cyber-attacks. However, due to the attacking cost, the attack strategy may vary a lot from one operation scenario to another from the perspective of adversary, which is not considered in previous studies. Therefore, in this paper, scenario-based two-stage sparse cyber-attack models for smart grid with complete and incomplete network information are proposed. Then, in order to effectively detect the established cyber-attacks, an interval state estimation-based defense mechanism is developed innovatively. In this mechanism, the lower and upper bounds of each state variable are modeled as a dual optimization problem that aims to maximize the variation intervals of the system variable. At last, a typical deep learning, i.e., stacked auto-encoder, is designed to properly extract the nonlinear and nonstationary features in electric load data. These features are then applied to improve the accuracy for electric load forecasting, resulting in a more narrow width of state variables. The uncertainty with respect to forecasting errors is modeled as a parametric Gaussian distribution. The validation of the proposed cyber-attack models and defense mechanism have been demonstrated via comprehensive tests on various IEEE benchmarks. Huaizhi Wang, Jiaqi Ruan, Guibin Wang, Bin Zhou 0005, Yitao Liu, Xueqian Fu, Jian-Chun Peng |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Robust Planning of Electric Vehicle Charging Facilities With an Advanced Evaluation MethodabstractThe planning of charging facilities (CFs) for electric vehicles (EVs) plays an important role for the extensive applications of EVs. Uncertainties existing in the development of future EV technology should be properly modeled to ensure the robustness of the planning scheme. The uncertainties concerned include EV development types, growth rate of load and traffic flow, and distributions of load and traffic flow in the smart grid. The existing single-stage planning model cannot fully evaluate the risks brought by all kinds of uncertainties. Given this background, the multistage CF planning problem considering uncertainties is studied in this work. First, several typical uncertainties in future smart grid with a high penetration of EVs are considered to generate development scenarios for multistage planning. Then, the well-established data envelopment analysis is utilized to evaluate the planning schemes while the novel EV expected energy not supplied cost is defined to measure the service ability of CFs. The final planning result obtained by the proposed framework will not only have good performance in the current stage but also exhibit robustness for all the considered scenarios in the future stage with respect to uncertainties. The application potential of the designed multistage planning framework is proved by an example with both the distribution network and traffic network included. Guibin Wang, Xian Zhang 0003, Huaizhi Wang, Jian-Chun Peng, Hui Jiang 0006, Yitao Liu, Zhao Xu 0002, Wenxin Liu 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Double closed-loop control strategy of LCL three-phase grid-connected inverterabstractGrid-connected inverter is an important part of the grid-connected system. Compared with the traditional L or LC filter, LCL filter has a better high-frequency harmonic attenuation performance. However, LCL filter has resonant peak, which has a great influence on the stability of the system. This paper first analyzes the effect of passive damping method on the resonance peak; then a double closed-loop control strategy with the inner loop of capacitor current and the outer loop of grid current is introduced; finally, the LCL three-phase grid-connected inverter with double-loop current control significantly reduces the Total Harmonic Distortion (THD), which verifies the feasibility of parameters design scheme of control strategy. Yitao Liu, Dianheng Jin, Huaizhi Wang, Guibin Wang, Jian-Chun Peng |
IECON | 4 |
| 2017 | Comparative evaluation of WBG and Si power devices for the flyback converterabstractPower switch devices which base on wide band-gap (WBG) semiconductor, such as silicon carbide metal-oxide-semiconductor-field-effect-transistor (SiC MOSFET) and gallium nitride high-electron-mobility transistor (GaN-HEMT) perform superior performance as compared with silicon (Si) MOSFET in high switching frequency, high blocking voltage, and high temperature operation. In this paper, a series of characteristic test and comparative evaluation of power loss and efficiency which based on Si MOSFET (IPW65R080CFDA), SiC MOSFET (C2M0080120D), and GaN HEMT (TPH3212PS) with similar current rating are conducted in the single-ended flyback converter. It shows that the WBG switch devices can reduce more power loss, improve efficiency, and enhance more power density instead of Si MOSFET. Yitao Liu, Zhendong Song, Huaizhi Wang, Guibin Wang, Jian-Chun Peng |
IECON | 4 |
| 2014 | The TH Express high performance interconnect networks
Zhengbin Pang, Guibin Wang, Dezun Dong, Guang Suo |
Frontiers Comput. Sci. | 5 |
| 2012 | MPtostream: an OpenMP compiler for CPU-GPU heterogeneous parallel systems
Xuejun Yang, Tao Tang 0001, Guibin Wang, Jia Jia 0004, Xinhai Xu |
Sci. China Inf. Sci. | 3 |
| 2011 | Coordinate strip-mining and kernel fusion to lower power consumption on GPUabstractAlthough general purpose GPUs have relatively high computing capacity, they also introduce high power consumption compared with general purpose CPUs. Therefore low-power techniques targeted for GPUs will be one of the most hot topics in the future. On the other hand, in several application domains, users are unwilling to sacrifice performance to save power. In this paper, we propose an effective kernel fusion method to reduce the power consumption for GPUs without performance loss. Different from executing multiple kernels serially, the proposed method fuses several kernels into one larger kernel. Owing to the fact that most consecutive kernels in an application have data dependency and could not be fused directly, we split large kernel into multiple slices with strip-mining method, then fuse independent sliced kernels into one kernel. Based on the CUDA programming model, we propose three different kernel fusion implementations, with each one targeting for a special case. Based on the different strip-ming methods, we also propose two fusion mechanisms, which are called invariant-slice fusion and variant-slice fusion. The latter one could be better adapted to the requirements of the kernels to be fused. The experimental results validate that the proposed kernel fusion method could effectively reduce the power consumption for GPU. Guibin Wang |
DATE | 1 |
| 2011 | Heterogeneity-Aware Peak Power Management for Accelerator-Based SystemsabstractPower management has become one of the first-order considerations in high performance computing field. Many recent studies focus on optimizing the performance of a computer system within a given power budget. However, most existing solutions adopt fixed period control mechanism and are transparent to the running applications. Although the application-transparent control mechanism has relatively good portability, it exhibits low efficiency in accelerator-based heterogeneous parallel systems. In typical accelerator-based parallel systems, different processing units have largely different processing speeds and power consumption. Under a given power constraint, how to choose the processor to be slowed down and how to schedule a parallel task onto different processors for the maximum performance are different from those in homogeneous systems and have not been well studied. From the motivating example in this paper, we could find that in order to efficiently harness the heterogeneous parallel processing, one should not only perform dynamic voltage/frequency scaling (DVFS) to meet the power budget, but also tune the parallel task scheduling to adapt to the changes. In this paper, we propose a heterogeneity-aware peak power management, which extends existing application-transparent power controller with an application-aware power controller. Firstly, we theoretically analyze the conditions for the maximum performance given a power budget for heterogeneous systems. Based on this result, we provide a power-constrained parallel task partition algorithm, which coordinates parallel task partition and voltage scaling for heterogeneous processing units to achieve the optimal performance given a system power budget. Finally, we evaluate the proposed method on a typical CPU-GPU heterogeneous system, and validate the superiority of application-aware power controller over the existing method. Guibin Wang, Yisong Lin |
ICPADS | 1 |
| 2011 | Communication-Aware Task Partition and Voltage Scaling for Energy Minimization on Heterogeneous Parallel SystemsabstractHeterogeneous parallel systems have become popular in general purpose computing and even high performance computing fields. There are many studies focused on harnessing heterogeneous parallel processing for better performance. However the energy optimization for heterogeneous system has not been well studied. Owing to the differences in performance and energy consumption, the energy optimization technique for heterogeneous system is different from the existing methods designed for homogeneous system. Besides typical voltage scaling method, reasonable task partitioning is also an essential method for optimizing energy consumption on heterogeneous systems. Through partitioning a data parallel task and mapping sub-tasks onto several processors, one could achieve better performance and reduced energy consumption. As the computation cost reduces with specific accelerators, the communication overhead becomes more prominent. Therefore, the task partition optimization should holistically consider the computation improvement and communication overhead to achieve higher energy efficiency. Typically, task partition and voltage scaling are not orthogonal and influence the effect of each other in the energy optimization problem. In order to harness both two knobs efficiently, this paper proposes an integer linear programming (ILP) based energy-optimal solution designed for heterogeneous system. We present a case study of optimizing MGRID benchmark on a typical CPU-GPU heterogeneous system. The experimental results demonstrate that the proposed method could exploit the heterogeneity in different processors and achieve improved energy efficiency. Guibin Wang |
PDCAT | 1 |
| 2011 | Power Optimization for GPU Programs Based on Software PrefetchingabstractGPUs render higher computing unit density than contemporary CPUs and thus exhibit much higher power consumption despite its higher power efficiency. The power consumption has become an important issue that impacts CPU's applications, thereby necessitating the low power optimization technology for GPUs. Software prefetching is an efficient way to alleviate the memory wall problem which overlaps the computing and memory access latencies. However, software prefetching will cause some power overhead because it increases the number and density of the instructions. Thus, we should consider the balance between the performance income and the power overhead when applying the optimization. To address this problem, in this paper we first analyze the multi-thread execution model of GPU and validate the potential space of software prefetching optimization. Then we give the software prefetching method for GPU programs to improve the performance. Aiming at two different objects: energy optimization under performance constraint and performance optimization under power constraint, we discuss the optimization methods based on software prefetching and dynamic voltage scaling technologies. The experimental results show that our method can efficiently optimize the energy consumption (performance) under the performance (power) constraint. Yisong Lin, Tao Tang 0001, Guibin Wang |
TrustCom | 3 |
| 2010 | Power-Efficient Work Distribution Method for CPU-GPU Heterogeneous SystemabstractAs the system scales up continuously, the problem of power consumption for high performance computing (HPC) system becomes more severe. Heterogeneous system integrating two or more kinds of processors, could be better adapted to heterogeneity in applications and provide much higher energy efficiency in theory. Many studies have shown heterogeneous system is preferable on energy consumption to homogeneous system in a multi-programmed computing environment. However, how to exploit energy efficiency (Flops/Watt) of heterogeneous system for a single application or even for a single phase in an application has not been well studied. This paper proposes a power-efficient work distribution method for single application on a CPU-GPU heterogeneous system. The proposed method could coordinate inter-processor work distribution and per-processor's frequency scaling to minimize energy consumption under a given scheduling length constraint. We conduct our experiment on a real system, which equips with a multi-core CPU and a multi-threaded GPU. Experimental results show that, with reasonably distributing work over CPU and GPU, the method achieves 14% reduction in energy consumption than static mappings for several typical benchmarks. We also demonstrate that our method could adapt to changes in scheduling length constraint and hardware configurations. Guibin Wang, Xiaoguang Ren |
ISPA | 1 |
| 2010 | Exploiting the reuse supplied by loop-dependent stream references for stream processorsabstractMemory accesses limit the performance of stream processors. By exploiting the reuse of data held in the Stream Register File (SRF), an on-chip, software controlled storage, the number of memory accesses can be reduced. In current stream compilers, reuse exploitation is only attempted for simple stream references, those whose start and end are known. Compiler analysis, from outside of stream processors, does not directly enable the consideration of other more complex stream references. In this article, we propose a transformation to automatically optimize stream programs to exploit the reuse supplied by loop-dependent stream references. The transformation is based on three results: lemmas identifying the reuse supplied by stream references, a new abstract representation called the Stream Reuse Graph (SRG) depicting the identified reuse, and the optimization of the SRG for our transformation. Both the reuse between the whole sequences accessed by stream references and between partial sequences is exploited in the article. In particular, partial reuse and its treatment are quite new and have never, to the best of our knowledge, appeared in scalar and vector processing. At the same time, reusing streams increases the pressure on the SRF, and this presents a problem of which reuse should be exploited within limited SRF capacity. We extend our analysis to achieve this objective. Finally, we implement our techniques based on the StreamC/KernelC compiler that has been optimized with the best existing compilation techniques for stream processors. Experimental results show a resultant speed-up of 1.14 to 2.54 times using a range of benchmarks. Xuejun Yang, Ying Zhang 0032, Xicheng Lu, Jingling Xue, Ian Rogers, Gen Li 0002, Guibin Wang, Xudong Fang |
ACM Trans. Archit. Code Optim. | 7 |
| 2009 | Program Optimization of Array-Intensive SPEC2k Benchmarks on Multithreaded GPU Using CUDA and Brook+abstractGraphic Processing Unit (GPU), with many light-weight data-parallel cores, can provide substantial parallel computing power to accelerate several general purpose applications. Both the AMD and NVIDIA corps provide their specific high performance GPUs and software platforms. As the floating-point computing capacity increases continually, the problem of ``memory-wall'' becomes more serious, especially for array-intensive applications. In this paper, we optimize and implement two SPEC2k benchmarks mgrid and swim on multithreaded GPU using CUDA and Brook+. In order to reduce the pressure on off-chip memory, we make use of data locality in multi-level memory hierarchies and hide long memory access latency via double-buffers. To balance inter-thread parallelism and intra-thread locality, we further tune thread granularity for each kernel and empirically study the best equilibrium point for this problem. Flow control instruction can significantly impact the effective instruction throughput. Oriented to this problem, we introduce a diverge elimination technology to convert condition expression into computing operation. Through all the optimizations, we gain the speedup of 10×-34× to the CPU implementation on the GPUs of AMD and NVIDIA respectively. Finally, we summarize and compares the GPUs from AMD and NVIDIA in hardware and software. Guibin Wang, Tao Tang 0001, Xudong Fang, Xiaoguang Ren |
ICPADS | 1 |
| 2009 | Program Optimization of Stencil Based Application on the GPU-Accelerated SystemabstractGraphic Processing Unit (GPU), with many light-weight data-parallel cores, can provide substantial parallel computational power to accelerate general purpose applications. But the powerful computing capacity could not be fully utilized for memory-intensive applications, which are limited by off-chip memory bandwidth and latency. Stencil computation has abundant parallelism and low computational intensity which make it a useful architectural evaluation benchmark. In this paper, we propose some memory optimizations for a stencil based application mgrid from SPEC 2K benchmarks. Through exploiting data locality in 3-level memory hierarchies and tuning the thread granularity, we reduce the pressure on the off-chip memory bandwidth. To hide the long off-chip memory access latency, we further prefetch data during computation through double-buffer. In order to fully exploit the CPU-GPU heterogeneous system, we redistribute the computation between these two computing resource. Through all these optimizations, we gain 24.2x speedup compared to the simple mapping version, and get as high as 34.3x speedup when compared with a CPU implementation. Guibin Wang, Xuejun Yang, Ying Zhang 0032, Tao Tang 0001, Xudong Fang |
ISPA | 1 |
| 2009 | SRF Coloring: Stream Register File Allocation via Graph Coloring
Xuejun Yang, Yu Deng 0001, Li Wang 0027, Xiaobo Yan, Jing Du 0002, Ying Zhang 0032, Guibin Wang, Tao Tang 0001 |
J. Comput. Sci. Technol. | 7 |
| 2008 | Exploiting loop-dependent stream reuse for stream processorsabstractThe memory access limits the performance of stream processors. By exploiting the reuse of data held in the Stream Register File (SRF), an on-chip storage, the number of memory accesses can be reduced. In current stream compilers reuse is only attempted for simple stream references, those whose start and end are known. Compiler analysis from outside of stream processors does not directly enable the consideration of other complex stream references. In this paper we propose a transformation to automatically optimize stream programs to exploit the reuse supplied by loop-dependent stream references. The transformation is based on three results: algorithms to recognize the reuse supplied by stream references, a new abstract expression called the Stream Reuse Graph (SRG) to depict the reuse and the optimization of the SRG for the transformation. Both the reuse between whole sequences accessed by stream references and that between partial sequences are exploited in the paper. In particular, the problem of exploiting partial stream reuse does not have its parallel in the traditional data reuse exploitation setting (for scalars and arrays). Finally, we have implemented our techniques using the StreamC/KernelC compiler for Imagine. Experimental results show a resultant speedup of 1.14 to 2.54 times using a range of typical stream processing application kernels. Xuejun Yang, Ying Zhang 0032, Jingling Xue, Ian Rogers, Gen Li 0002, Guibin Wang |
PACT | 6 |
| 2008 | Scientific Computing Applications on a Stream ProcessorabstractStream processors, developed for the stream programming model, perform well on media applications. In this paper we examine the applicability of a stream processor to scientific computing applications. Eight scientific applications, each having different performance characteristics, are mapped to a stream processor. Due to the novelty of the stream programming model, we show how to map programs in a traditional language, such as FORTRAN. In a stream processor system, the management of system resources is the programmers' responsibility. We present several optimizations, which enable mapped programs to exploit various aspects of the stream processor architecture. Finally, we analyze the performance of the stream processor and the presented optimizations on a set of scientific computing applications. The stream programs are from 1.67 to 32.5 times faster than the corresponding FORTRAN programs on an Itanium 2 processor, with the optimizations playing an important role in realizing the performance improvement. Ying Zhang 0032, Xuejun Yang, Guibin Wang, Ian Rogers, Gen Li 0002, Yu Deng 0001, Xiaobo Yan |
ISPASS | 3 |
| 2007 | Implementation and Evaluation of Jacobi Iteration on the Imagine Stream Processor
Jing Du 0002, Xuejun Yang, Tao Tang 0001, Guibin Wang |
HiPC | 5 |
| 2007 | Architecture-Based Optimization for Mapping Scientific Applications to Imagine
Jing Du 0002, Xuejun Yang, Guibin Wang, Tao Tang 0001 |
ISPA | 3 |
| 2007 | Implementation and Optimization of Sparse Matrix-Vector Multiplication on Imagine Stream Processor
Li Wang 0027, Xuejun Yang, Guibin Wang, Xiaobo Yan, Yu Deng 0001, Jing Du 0002, Ying Zhang 0032, Tao Tang 0001 |
ISPA | 3 |