EDBT 2026 Demo / reviewers in the wild / expert
Qingbo Wu 0003
dblp:07/8007-3
· DBLP profile ↗
49ranked-venue papers
2as first author
19since 2021 · last 2026
0009-0001-9102-9890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Operator Fusion for LLM Inference on the Tensix Architecture
Qingbo Wu 0003, Ke Li 0026, Wenzhu Wang, Jie Yu 0008, Ruian Zhang |
ICIC (23) | 1 |
| 2026 | Emp: enhance memory in data pruning
Jinying Xiao, Ping Li 0034, Jie Nie, Bin Ji 0002, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Qingbo Wu 0003, Jie Yu 0008 |
Data Min. Knowl. Discov. | 8 |
| 2026 | Adaptive CPU sharing for co-located latency-critical JVM applications and batch jobs under dynamic workloads
Dishi Xu, Fagui Liu, Bin Wang 0048, Xuhao Tang 0001, Qingbo Wu 0003 |
Future Gener. Comput. Syst. | 5 |
| 2026 | CoreScaler: A Resource-Efficient Hybrid Scaling Framework for Dynamic Workloads in Cloud
Dinghao Zeng, Fagui Liu, Runbin Chen, Jingwei Tan, Dishi Xu, Qingbo Wu 0003, C. L. Philip Chen |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2025 | EC5: Edge-cloud collaborative computing framework with compressive communication
Jingwei Tan, Fagui Liu, Bin Wang 0048, Qingbo Wu 0003, C. L. Philip Chen |
Future Gener. Comput. Syst. | 4 |
| 2025 | GenesisRM: A state-driven approach to resource management for distributed JVM web applications
Dishi Xu, Fagui Liu, Bin Wang 0048, Xuhao Tang 0001, Dinghao Zeng, Huaiji Gao, Runbin Chen, Qingbo Wu 0003 |
Future Gener. Comput. Syst. | 8 |
| 2025 | A Survey of AI Inference Technologies for On-Device SystemsabstractIn recent years, artificial intelligence(AI) technologies represented by foundation models have experienced rapid development. Concurrently, On-device AI inference has become the primary approach for intelligent technology applications, offering advantages such as low latency, high security, and personalization. However, due to the limited resources of on-device systems, on-device AI inference faces new challenges, including improving computational efficiency, optimizing task parallelism, and model optimization. This survey addresses these challenges from a software and algorithmic perspective, focusing on three key areas: Operator Computation: Explores methods to accelerate matrix multiplication and convolution, as well as techniques like operator fusion and vectorized computation. Task Inference: Analyzes heterogeneous and distributed computing, memory allocation, and energy-efficient tuning to improve the parallel execution and energy efficiency of inference tasks. AI Models: Covers model compression, lookup table quantization, and model architecture design to reduce computational complexity and storage requirements. By analyzing these areas, the survey aims to improve inference speed, reduce resource dependency, and provide insights into the future trends of on-device AI technology. Wenzhu Wang, Ke Li 0026, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Qingbo Wu 0003 |
IEEE Internet Things J. | 6 |
| 2025 | Cost and Makespan-Aware Task Scheduling With Deep Reinforcement Learning in Multicloud EnvironmentsabstractThe multicloud environments (MCE) represent a novel paradigm encompassing multiple infrastructure as a service (IaaS) providers, enabling users to tailor and optimize cloud services according to their specific requirements. This approach effectively addresses the limitations of a single cloud environment (SCE) regarding technical constraints, geographical coverage deficiencies, and cost-effectiveness concerns while catering to the increasingly diverse and expanding user demands. In MCE, users must employ appropriate strategies to efficiently allocate diverse tasks across multiple cloud service providers (CSPs) by leveraging the best available resources. Traditional scheduling algorithms are inadequate for addressing the complexities of such MCE. This study introduces a framework for the task scheduling procedure in MCE, treating independent task scheduling as a Markov decision process (MDP). We propose a novel agent environment framework that is designed based on the distinctive characteristics of MCE and enables independent task scheduling. Furthermore, we propose a task scheduling algorithm for MCE based on deep reinforcement learning (DRL) to optimize cost and makespan according to diverse user requirements. The simulation experiments are conducted using both simulated datasets and real-world datasets, demonstrating that our proposed algorithm surpasses the other five algorithms in terms of cost minimization and makespan optimization. Xuhao Tang 0001, Fagui Liu, Bin Wang 0048, Jun Jiang 0003, Quan Tang 0001, Qingbo Wu 0003, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2024 | DGLP: Incorporating Orientation Information for Enhanced Link Prediction in Directed GraphsabstractLink prediction in directed graphs offers a solution for uncovering detailed and accurate relationships among distinct entities. Unlike conventional link prediction in undirected graphs, the task becomes more intricate in directed graphs as it involves predicting both associations and orientations. Existing methods simply apply classic graph embedding techniques to learn node representations, followed by mapping representations of corresponding node pairs into probabilities indicating potential links. However, the inadequate capture of orientation information and sole reliance on node representations for prediction hinder the effective differentiation of orientation, thereby impeding the link prediction accuracy. In response, we introduce DGLP, an orientation-aware link prediction method tailored for directed graphs. DGLP utilizes the incidence matrix to learn both node and edge representations, effectively capturing structural and orientation information. By leveraging edge representations, DGLP achieves accurate link prediction in directed graphs without relying solely on implicit node representations. Experiments across six datasets demonstrate the effectiveness of DGLP, achieving a 1.2x improvement in prediction results. Yusen Zhang 0007, Yusong Tan, Songlei Jian, Qingbo Wu 0003, Kenli Li 0001 |
ICASSP | 4 |
| 2024 | Workflow scheduling based on asynchronous advantage actor-critic algorithm in multi-cloud environmentabstractRecently, the multi-cloud environment (MCE) has increasingly become the preferred choice of users. As with the cloud environment, efficient workflow scheduling in a MCE remains crucial for identifying the cost efficiency and overall performance of the MCE. In MCE, the resources exhibit heterogeneity, complexity, and dynamism. Simultaneously, the intricate inter-task dependencies among workflow tasks, diverse Quality of Service (QoS) metrics for users, and multiple cloud service providers’ (CSPs) billing mechanisms significantly amplify the workflow scheduling challenge. Motivated by the application of reinforcement learning (RL) in workflow scheduling in a cloud environment, this paper proposes a scheduling algorithm that takes advantage of the asynchronous advantage actor–critic algorithm (A3C) to balance cost, makespan and resource utilization in workflow scheduling in a MCE. By analyzing the elements in the MCE, we design and define multiple agents in the MCE, and each cloud service provider will have an agent to record the state and update the local parameters. For the workflow task submitted by the user, the action is selected according to the initialization policy and submitted to the scheduling action to allocate the task to a designated virtual machine in the MCE so that each agent can more clearly perceive the environment change and adapt to the MCE. In contrast to the traditional A3C algorithm, we design a new critic network according to the data characteristics of real-world scientific workflows so that each agent is more suitable for real-world scientific workflow data. Through multiple sets of simulation experiments, the workflow scheduling algorithm based on the A3C algorithm in the MCE (MCWS-A3C) was compared with three benchmark methods. The experimental results show that the proposed method has better advantages than other methods in terms of cost, makespan, and resource utilization . Specifically, on the Montage_100 dataset, the average cost was reduced by 55.12% compared to other methods. The pioneering introduction of the A3C algorithm that adapts to the dynamic environment into the MCE brings more possibilities to address the issue of workflow scheduling in the MCE. Xuhao Tang 0001, Fagui Liu, Bin Wang 0048, Dishi Xu, Jun Jiang 0003, Qingbo Wu 0003, C. L. Philip Chen |
Expert Syst. Appl. | 6 |
| 2024 | GPU and VPU Enabled Virtual Mobile Infrastructure for 3-D Image Rendering and its Application in TelemedicineabstractTelemedicine for 3D images on mobile devices presents promising development opportunities. Being constrained by computing power and storage capacity on mobile devices, the processing performance of 3D medical images is insufficient for more demanding tasks. Using virtual mobile infrastructure technology to utilize cloud resources is a common solution. But it encounters the challenge of poor performance in data transmission, image rendering and image coding. This paper presents a GPU and VPU enabled Open Virtual Mobile Infrastructure (OpenVMI) for 3D image rendering to solve the challenge. It makes two improvements. First, a bespoke GPU driver is developed in the Android Docker, optimizing the transmission workflow for data transmission and image rendering. Second, a Video Process Unit (VPU) is added to the hardware layer to code rendered results in H.264 format, replacing CPU coding which consumes a large amount of CPU resources. By adopting the OpenVMI, the telemedicine training system proposed in this paper presents an easy-to-set up, cheap and low latency solution that is particularly helpful for telemedicine training in remote and underdeveloped areas. Performance experiments suggest that the OpenVMI delivers better performance than existing state-of-the-art systems, even in mobile devices with weaker hardware capabilities. Concurrency experiment suggests that a single host server can support up to 24 concurrent training sessions, which makes the OpenVMI very helpful for telemedicine training that demands high concurrency. The OpenVMI-based solution proposed in this paper is not restricted to the use of telemedicine training, but also suitable for other application areas such as Virtual Reality and Augmented Reality in mobile environments. Zhipeng Fu, Wanpeng Xu, Changguo Guo, Qingbo Wu 0003 |
IEEE Internet Things J. | 5 |
| 2023 | Robust unsupervised network intrusion detection with self-supervised masked context reconstruction
Wei Wang 0130, Songlei Jian, Yusong Tan, Qingbo Wu 0003, Chenlin Huang |
Comput. Secur. | 4 |
| 2022 | Fine-tuning more stable neural text classifiers for defending word level adversarial attacks
Zibo Yi, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003 |
Appl. Intell. | 4 |
| 2022 | Representation learning-based network intrusion detection system by capturing explicit and implicit feature interactions
Wei Wang 0130, Songlei Jian, Yusong Tan, Qingbo Wu 0003, Chenlin Huang |
Comput. Secur. | 4 |
| 2022 | Towards an Efficient and Robust Adversarial Attack Against Neural Text ClassifierabstractAdversarial attack is a serious threat to neural network-based natural language processing applications. Adversarial attack uses tiny well-crafted perturbations to mislead neural networks. While existing adversarial text attacks can achieve good attack effects, they still do not guarantee efficiency and robustness. The adversarial text attacks are more efficient if they use less perturbation to achieve a higher attack success rate. The attacks are more robust if they can achieve a higher success rate when defense strategies are applied. To improve the efficiency and robustness of the adversarial attack, we propose SMAL: Saliency Map Attack with Levenshtein-similarity. The proposed attack consists of two parts: (1) The saliency map measures the perturbation priority of each word. It considers not only the influence of each word on the classification result but also how to maintain the misled classification result to improve the robustness of the attack. (2) Levenshtein-similarity network embeds words into edit distance space. When perturbing sentences, some words are replaced by substitutions with less edit distance. This can reduce the amount of modification, which improves the efficiency of the attack. Since the words are embedded in edit distance space rather than semantic space, the semantic-based defense is not effective for this attack, which improves the robustness. The experiments show that SMAL achieves a higher attack success rate with fewer perturbations. Also, the proposed attack is better when attacking a classifier defended by adversarial training. Zibo Yi, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003 |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2022 | Adaptive Processor Frequency Adjustment for Mobile-Edge Computing With Intermittent Energy SupplyabstractWith astonishing speed, bandwidth, and scale, mobile-edge computing (MEC) has played an increasingly important role in the next generation of connectivity and service delivery. Yet, along with the massive deployment of MEC servers, the ensuing energy issue is now on an increasingly urgent agenda. In the current context, the large-scale deployment of renewable-energy-supplied MEC servers is perhaps the most promising solution for the incoming energy issue. Nonetheless, as a result of the intermittent nature of their power sources, these special design MEC servers must be more cautious about their energy usage, in a bid to maintain their service sustainability as well as service standard. Targeting optimization on a single-server MEC scenario, we, in this article, propose neural network-based adaptive frequency adjustment (NAFA), an adaptive processor frequency adjustment solution, to enable an effective plan of the server’s energy usage. By learning from the historical data revealing request arrival and energy harvest pattern, the deep reinforcement learning-based solution is capable of making intelligent schedules on the server’s processor frequency, so as to strike a good balance between service sustainability and service quality. The superior performance of NAFA is substantiated by real-data-based experiments, wherein NAFA demonstrates up to 20% increase in the average request acceptance ratio and up to 50% reduction in average request processing time. Tiansheng Huang, Weiwei Lin 0001, Xiumin Wang 0005, Qingbo Wu 0003, Rui Li 0047, Ching-Hsien Hsu, Albert Y. Zomaya |
IEEE Internet Things J. | 5 |
| 2021 | Many-To-Many Chinese ICD-9 Terminology Standardization Based on Neural Networks
Shasha Li 0001, Jie Yu 0008, Yusong Tan, Jun Ma 0015, Qingbo Wu 0003 |
ICIC (2) | 6 |
| 2021 | Span Representation Generation Method in Entity-Relation Joint Extraction
Yongtao Tang, Jie Yu 0008, Shasha Li 0001, Bin Ji 0002, Yusong Tan, Qingbo Wu 0003 |
ICIC (2) | 6 |
| 2021 | On-demand cut off the covert channel to mitigate meltdown
Yusong Tan, Baozi Chen, Liehuang Zhu, Qingbo Wu 0003, Yuanzhang Li 0001 |
Sci. China Inf. Sci. | 4 |
| 2020 | Span-based Joint Entity and Relation Extraction with Attention-based Span-specific and Contextual Semantic RepresentationsabstractSpan-based joint extraction models have shown their efficiency on entity recognition and relation extraction.These models regard text spans as candidate entities and span tuples as candidate relation tuples.Span semantic representations are shared in both entity recognition and relation extraction, while existing models cannot well capture semantics of these candidate entities and relations.To address these problems, we introduce a span-based joint extraction framework with attention-based semantic representations.Specially, attentions are utilized to calculate semantic representations, including span-specific and contextual ones.We further investigate effects of four attention variants in generating contextual semantic representations.Experiments show that our model outperforms previous systems and achieves state-of-the-art results on ACE2005, CoNLL2004 and ADE. Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003 |
COLING | 5 |
| 2020 | Research on Chinese medical named entity recognition based on collaborative cooperation of multiple neural network models
Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Jintao Tang, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003, Yun Ji |
J. Biomed. Informatics | 6 |
| 2019 | CLASC: A Changelog Based Automatic Code Source Classification Method for Operating System PackagesabstractOpen source represents an important way in which today's software is developed. The adoption of open source software continues to accelerate because of the great potential it offers, such as productivity improvement, cost savings and quicker innovation. While the complexity and the size of software composition grow, it becomes difficult to effectively scan and track the code source, especially for software with tremendous scale of code, such as operating systems. So far, existing work on open source components mainly focus on how to mitigate potential license incompliance, to reduce potential security risks introduced by open source vulnerabilities, and to detect and match open source components in the code. To ensure code traceability and manageability for large scale mixed-source operating system, we believe it is beneficial to automatically distinguish sources of the system code in the granularity of software packages and manage them separately. However, according to the literature, there is a lack of relevant work in this area. In this paper, we first classify the packages into three categories in terms of code source from the perspective of OS developers and maintainers. Then we propose CLASC, an efficient code source classification algorithm. With the capability of package info extraction and analysis, CLASC can classify software packages into the defined categories according to their changelog info. And we design and implement KyAnalyzer, a Web-based package management and code source analysis platform. It provides automatic code source analyzing services and is capable of managing OS packages differentially according to their different categories of code source with CLASC incorporated as a component of it. Experimental results show the correctness and efficiency of the Web-enabled package source classifier. Yi Ren 0008, Jianbo Guan, Jun Ma 0015, Yusong Tan, Qingbo Wu 0003 |
APSEC | 5 |
| 2019 | Incremental Learning of GAN for Detecting Multiple Adversarial Attacks
Zibo Yi, Jie Yu 0008, Shasha Li 0001, Yusong Tan, Qingbo Wu 0003 |
ICANN (3) | 5 |
| 2019 | An Efficient and Transparent Approach for Adaptive Intra-and Inter-Node Virtual Machine Communication in Virtualized CloudsabstractNetwork I/O workloads are dominating as one of the leading costs for most of the virtualized clouds. One way to improve the inter virtual machine (VM) inefficiency is to build shared memory channels between VMs co-located on the same physical node to by-pass traditional TCP/IP network stack, so that the overhead is reduced by shorter communication path and fewer kernel interactions. However, it is a key challenge for existing work to achieve high performance inter-VM communication while keeping the capability of VM live migration, and most of existing work are neither seamlessly agile in the presence of VM live migration nor transparent to upper users as well as to operating system kernels, which limits the application of current co-location aware shared-memory based approaches. In this paper, we present the design and implementation of XenVMC, an adaptive and transparent inter-VM communication system for high performance network I/O in virtualized clouds. With proposed dynamic co-located VM membership update mechanism, XenVMC is applicable not only to intra-node VM communication, but also to cross-node communication. It also supports adaptive switching between shared-memory based channel and traditional network-based channel in case of VM live migration, with the aid of proposed VM migration perception and handling mechanisms. XenVMC enables efficient data transmission for both TCP and UDP workloads, with multilevel transparency guaranteed. Extensive experiments show that XenVMC achieves better performance for both TCP and UDP workloads with high transparency, compared with both native virtualized environment and representative existing work. Experimental results also show that it is capable of automatically handling VM migration correctly with acceptable latency. Yi Ren 0008, Renshi Liu, Qi Zhang 0009, Jianbo Guan, Ziqi You, Yusong Tan, Qingbo Wu 0003 |
ICPADS | 7 |
| 2019 | Resource stealing: a resource multiplexing method for mix workloads in cloud system
Yusong Tan, Fuhui Wu, Qingbo Wu 0003, Xiangke Liao |
J. Supercomput. | 3 |
| 2018 | A virtual cluster embedding approach by coordinating virtual network and software-defined network
Yusong Tan, Rongzhen Li, Qingbo Wu 0003 |
Soft Comput. | 3 |
| 2017 | Drug-Drug Interaction Extraction via Recurrent Neural Network with Multiple Attention Layers
Zibo Yi, Shasha Li 0001, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003, Ting Wang 0009 |
ADMA | 5 |
| 2017 | MicRun: A framework for scale-free graph algorithms on SIMD architecture of the Xeon PhiabstractGraph algorithms currently play increasingly important roles, especially in social networks and language modeling scenarios. Recently, accelerating graph algorithms by heterogeneous high performance computers with the integrated cores and expanded SIMD lanes has been becoming the mainstream. However, the existing methods, restricted by the low-efficiency grouping strategy and the non-optimized selection mechanism of tile size of a graph, are far below our expectations in many ways. Moreover, there are few convenient integrated tools provided for deploying the graph algorithms on MIC architecture. In this paper, we propose a high-efficiency framework MicRun, which is flexible to be used for graph algorithms on SIMD architecture of the Xeon Phi. There are two key components in MicRun, the Bucket Grouping module and Auto-tuning module. In the Grouping module, an optimization algorithm is designed for splitting graph tiles into conflict-free groups, which can be directly processed on SIMD parallelism. In the Auto-tuning module, a novel strategy is proposed for optimizing the tile size to boost execution efficiency of the graph computation. MicRun currently supports Bellman-Ford and PageRank algorithms, we also conduct extensive validation experiments on MicRun. Experimental results show that MicRun outperforms existing mechanisms in terms of storage and time overhead. As a consequence, both graph algorithms achieve an average speedup of 1.1× by MicRun, compared with the state-of-the-art. Qingbo Wu 0003, Yusong Tan, Jie Yu 0008, Qi Zhang 0028, Xiaoling Li 0002, Lei Luo 0002 |
ASAP | 2 |
| 2016 | A novel optimization scheme for caching in locality-aware P2P networksabstractDeploying cache has been generally adopted by Internet service providers (ISPs) to mitigate P2P traffic in recent years. Most traditional caching algorithms are designed for locality-unaware P2P networks, which mainly consider the requested frequency of contents as the principle of caching policies. However, in more prevalent locality-aware conditions with biased neighbor-selection policies, the existing caching schemes can hardly optimize the situation. In this paper we show that, what need to be cached in locality-aware conditions are the contents that can not be well provided by local neighbors, rather than the contents which are requested most frequently. Therefore, states of local neighbors should be taken into consideration in caching policies. We first present a new model in which P2P cache and locality-aware neighbor selection work together. We focus on inter-ISP traffic and available bandwidth of users in order to benefit both ISPs and users. Based on the mathematical model, a novel caching algorithm is proposed which considers replacement and allocation policies together. According to trace-driven simulations, the proposed algorithm outperforms other two representative caching algorithms in various scenarios. Shaoduo Gan, Jiexin Zhang 0001, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003 |
ISCC | 7 |
| 2016 | An Optimized DHT for Linux Package DistributionabstractThe rapid rising of Linux users requires P2P, an efficient content transport method, to distribute packages. Different from traditional streaming P2P systems, a P2P package distribution system is hazarded by the special characteristics of small package size and hot packages. Due to the small package size, DHT search performance, which is rarely considered in traditional P2P system, comes to be an important factor in package distribution process. To improve the DHT search performance in this circumstance, we propose four kinds of optimizations as follow. Firstly, Fast-Response is proposed to eliminate useless searches after having found the target pair. Secondly, LRU Cache is proposed to reduce search hops on the same package. Thirdly, Leap Cache is proposed to reduce the cache redundancy. Finally, Probability Cache, gathering those optimizing above and considering hot packages in addition, is proposed to get increase of cache hit rate and overall efficiency improvement. We simulate our optimizations using PeerSim platform. The results show that all the four optimizations get considerable improvement in performance. With the best situation of Probability Cache, 86.22% delay time of original Kademlia is saved. Qi Zhang 0028, Jie Yu 0008, Lei Luo 0002, Jun Ma 0015, Qingbo Wu 0003, Shasha Li 0001 |
ISPDC | 5 |
| 2016 | ERPC: An Edge-Resources Based Framework to Reduce Bandwidth Cost in the Personal Cloud
Shaoduo Gan, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003, Shasha Li 0001 |
WAIM (2) | 6 |
| 2016 | PCP-B2: Partial critical path budget balanced scheduling algorithms for scientific workflow applications
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Rongzhen Li, Wei Wang 0130 |
Future Gener. Comput. Syst. | 2 |
| 2016 | micMR: An efficient MapReduce framework for CPU-MIC heterogeneous architecture
Wenzhu Wang, Yusong Tan, Qingbo Wu 0003, Yaoxue Zhang |
J. Parallel Distributed Comput. | 3 |
| 2015 | Optimizing the MapReduce Framework for CPU-MIC Heterogeneous Cluster
Wenzhu Wang, Qingbo Wu 0003, Yusong Tan, Yaoxue Zhang |
APPT | 2 |
| 2015 | Maximize Throughput Scheduling and Cost-Fairness Optimization for Multiple DAGs with Deadline Constraint
Wei Wang 0130, Qingbo Wu 0003, Yusong Tan, Fuhui Wu |
ICA3PP (2) | 2 |
| 2015 | Unified Multi-constraint and Multi-objective Workflow Scheduling for Cloud System
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Wei Wang 0130 |
ICA3PP (2) | 2 |
| 2015 | Workflow scheduling in cloud: a survey
Fuhui Wu, Qingbo Wu 0003, Yusong Tan |
J. Supercomput. | 2 |
| 2015 | Complementary Synthesis for Encoder with Flow Control MechanismabstractComplementary synthesis automatically generates an encoder's decoder with the assumption that the encoder's all input variables can always be uniquely determined by its output symbol sequence. However, to prevent the faster encoder from overwhelming the slower decoder, many encoders employ flow control mechanism that fails this assumption. Such encoders, when their output symbol sequences are too fast to be processed by the decoders, will stop transmitting data symbols, but instead transmitting idle symbols that can only uniquely determine a subset of the encoder's input variables. And the decoder should recognize and discard these idle symbols. This mechanism fails the assumption of all complementary synthesis algorithms, because some input variables can't be uniquely determined by the idle symbol. A novel algorithm is proposed to handle such encoders. First, it identifies all input variables that can be uniquely determined, and takes them as flow control variables. Second, it infers a predicate over these flow control variables that enables all other input variables to be uniquely determined. Third, it characterizes the decoder's Boolean function with Craig interpolant. Experimental results on several complex encoders indicate that this algorithm can always correctly identify the flow control variables, infer the predicates and generate the decoder's Boolean functions. ShengYu Shen, Qingbo Wu 0003, Huadong Dai, Yan Jia 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2013 | Residency-Aware Virtual Machine Communication Optimization: Design Choices and TechniquesabstractNetwork I/O workloads are dominating in many data centers and cloud computing environments today. One way to improve inter Virtual Machine (VM) communication efficiency is to support co-resident VM communication by using shared memory based approaches and to resort to the traditional TCP/IP for inter-VM communications between VMs that are located on different physical hosts. Although a number of independent efforts are dedicated to improving communication efficiency between co-resident VMs, they differ from one another in terms of how the inter-VM communication optimization is carried out and where in the software stack the shared memory channel is established. In this paper, we provide an in-depth overview of the design choices and techniques for optimizing the performance of the co-resident inter-VM communication, with dual objectives. First, we describe the core design guidelines and key issues for optimizing inter-VM communication by using shared memory based mechanisms. Typical issues include choices of implementation layer in the software stack, seamless agility for VM live migration and VM dynamic deployment support, multilevel transparency. Second, we conduct a comprehensive analysis of representative state-of-the-art research efforts and implementation techniques based on the core design guidelines. We also give an analysis of future requirements in advanced features such as reliability, security and stability. The research reported in this paper not only provides the reference for developing the next generation of inter-VM communication optimization mechanisms, but also offers opportunities for both cloud infrastructure providers and cloud service consumers to improve inter-VM communication efficiency in virtualized platforms. Yi Ren 0008, Ling Liu 0001, Qi Zhang 0009, Qingbo Wu 0003, Jinzhu Kong, Jianbo Guan, Huadong Dai |
IEEE CLOUD | 4 |
| 2013 | A Vectorized K-Means Algorithm for Intel Many Integrated Core Architecture
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Lifeng Wei, Lisong Shao |
APPT | 2 |
| 2013 | Speeding Up SIFT Algorithm by Multi-core Processor Supporting SIMD Instruction SetsabstractScale Invariant Feature Transform (SIFT) method plays a critical role in a wide variety of vision applications. But it is now facing the real-time computational challenge. Parallel computing is one of the most promising solutions to overcome the computational challenge. In this paper, we target at parallelizing SIFT by multi-core architecture with per-core SIMD support. We focus on the SIMDization of data parallel parts of SIFT to fully utilize per-core computing power. At Orientation Assignment and Key point Descriptor stages, we observe that load balance is an important factor. We also implement the optimized algorithm on multi-core system with SIMD support from Tianhe-2 Supercomputer and make comparison with the State-of-the-Art parallel SIFT algorithms. Fuhui Wu, Qingbo Wu 0003, Yusong Tan |
CAD/Graphics | 2 |
| 2013 | Multi-resource Aware Congestion Control in Data CentersabstractNetwork has been widely reported as a bottleneck of data center applications. However, current researches of congestion control are unaware of multiple resources consuming and decrease all flows when congestion, ignoring some involved flows may not be the faults. In this paper, we propose a novel multi-resources aware congestion control framework MRTCP to provide a fine-grain control on flows when congestions appear. MRTCP exploits a multi-tuple vector model to measure multi-resources provision and consumption, and develops a novel metric RB (Resource Balance) to denote the heterogeneous amounts of resources employed by each flow. It analyzes which resources are being the bottlenecks that lead to congestions, calculates the responsibility of each flow to this congestion, and then adjusts their sending rates respectively. Our experiment results demonstrate that MRTCP is able to optimize network multi-resources utilization and improve network throughput without adding obvious packets delays. Deke Guo, Qingbo Wu 0003, Shanshan Li 0001, Yusong Tan, Quanyuan Wu |
ICPADS | 2 |
| 2012 | A fast and transparent communication protocol for co-resident virtual machinesabstractNetwork I/O workloads are dominating in most of the Cloud datacenters today. One way to improve inter-VM communication efficiency is to support co-resident VM communication using a faster communication protocol than the traditional TCP/IP commonly used regardless whether VMs are located on the same Yi Ren 0008, Ling Liu 0001, Xiaojian Liu 0005, Jinzhu Kong, Huadong Dai, Qingbo Wu 0003, Yuan Li 0011 |
CollaborateCom | 6 |
| 2012 | Comparability Graph Coloring for Optimizing Utilization of Software-Managed Stream Register Files for Stream ProcessorsabstractThe stream processors represent a promising alternative to traditional cache-based general-purpose processors in achieving high performance in stream applications (media and some scientific applications). In a stream programming model for stream processors, an application is decomposed into a sequence of kernels operating on streams of data. During the execution of a kernel on a stream processor, all streams accessed must be communicated through a nonbypassing software-managed on-chip memory, the SRF (Stream Register File). Optimizing utilization of the scarce on-chip memory is crucial for good performance. The key insight is that the interference graphs (IGs) formed by the streams in stream applications tend to be comparability graphs or decomposable into a set of comparability graphs. We present a compiler algorithm for finding optimal or near-optimal colorings, that is, SRF allocations in stream IGs, by computing a maximum spanning forest of the sub-IG formed by long live ranges, if necessary. Our experimental results validate the optimality and near-optimality of our algorithm by comparing it with an ILP solver, and show that our algorithm yields improved SRF utilization over the First-Fit bin-packing algorithm, the best in the literature. Xuejun Yang, Li Wang 0027, Jingling Xue, Qingbo Wu 0003 |
ACM Trans. Archit. Code Optim. | 4 |
| 2010 | Vapor: Virtual Machine Based Parallel Program Profiling FrameworkabstractIt is hard to execute parallel program efficiently on man-core platform because we could not divide program into appropriate granularity executed simultaneously. Based on virtual machine and binary translation technologies the article proposes the vapor profiling framework that uses SBIRP instruction in-place replacement method to collect program's run-time control flow and data flow information precisely. Moreover, it explains how to create control flow and data flow dependency graphs. Experiment results prove that vapor has better performance than traditional methods. Yusong Tan, Wei Chen 0009, Qingbo Wu 0003 |
ICPADS | 3 |
| 2010 | Scalability comparison of commodity operating systems on multi-coresabstractIn this paper, we evaluate and compare the parallel scalability of three commodity operating systems (Linux, Solaris and FreeBSD) on an AMD 32-core platform. Measurements of microbenchmarks and a real-life application reveal that no operating system scales totally better than another for microbenchmarks; for the real-life application, Linux and Solaris are competitive in scalability and perform better than FreeBSD. Related kernel source analysis and performance data suggest that synchronization primitives protecting the shared data structure in kernels are the root cause of the poor scalability on multi-cores. Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003 |
ISPASS | 4 |
| 2009 | CFS Optimizations to KVM Threads on Multi-Core EnvironmentabstractMulti-core architecture provides more on-chip parallelism and powerful computational capability. It helps virtualization achieve scalable performance. KVM (kernel based virtual machine) is different from other virtualization solutions which can make use of the Linux kernel components such as completely fair scheduler (CFS). However, CFS treats the KVM threads as normal tasks without considering about their unique features such as thread allocation mechanism and lock inside guest virtual machine, which may harm the KVM virtualization performance. In this paper, we analyze a phenomenon that some guest multi-threaded applications have very low performance when scheduled by CFS. As a solution to this problem, we introduce two kinds of optimizations in CFS: (1) configuration optimizations (2) lock optimizations. Our contributions are: (1) implement 5 original and 2 newest proposed optimizations in the newest Linux kernel. (2) Classify and compare them, a brief analysis is also given. They are all very simple and general to other virtual machine monitors such as Xen and schedulers as O(1). The performance of our CFS optimizations to KVM threads is measured by running some well-known benchmarks in two guest virtual machines on an 8-core server which models the real world applications. The results indicate our scheduling optimizations can improve the overall system performance. This paper can provide useful advices to KVM developers and virtualization data center administrators. Yisu Zhou, Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003 |
ICPADS | 7 |
| 2009 | Block-Based In-Place Replacement Strategy for x86 Sensitive Instructions in Virtual MachineabstractIt is trendy that virtualization technology is adopted by server and desktop computers recently. Binary translation is an important method to implement full virtualization supporting any guest operating system without modification. Traditional methods use trap or interrupt to catch sensitive instruction's execution. Its performance is influenced by trap's context switch overhead. This article proposes a novel code scanning and replacing strategy, named as Block-based In-Place Replacement. BIPR tries to find a code block whose length is longer than 5 bytes and replaces the block with 5-bytes JMP instruction. The translated code block has same run-time mode as original code. As a result, BIPR's cost is lower than traditional trap methods. Moreover, it gives an optimize strategy, i.e. Super Block-based In-Place Replacement, to reduce unnecessary translation overhead of BIPR and get better performances. Experiment results prove that SBIPR performs pretty. Yusong Tan, Qingbo Wu 0003 |
ISPA | 3 |
| 2009 | System Monitoring and Controlling Mechanism Based on HypervisorabstractCurrent commodity operating systems allow a privileged user to run some programs in kernel mode by installing a kernel module or a device driver, but there isnpsilat an available method to verify the reliability of these programs. As a result, malware leverages this way to corrupt system services, defeat anti-malware and even get control of the whole system. It makes operating-system-based security tools undergo all kinds of hardships in face of various attacks. Virtualization technology brings a new chance to protect operating systems out-of-box in hypervisors. This article proposes a system monitoring and controlling framework, named as BMCS, which monitors low-level states and hardware accessing events of operating system in a hypervisor. BMCS abstracts low-level states and events and profiles OS-level semantics with the help of a hosted daemon. Moreover, it can control the execution of operating system software through the hypervisor. Our out-of box method can collect real run-time operating system profile and control the execution of the OS according some defined rules. As a result, it makes the operating system stronger. We implement BMCS based on KVM and experiment results prove that BMCS performs well. Qingbo Wu 0003, Yusong Tan |
ISPA | 1 |