Wenhong Tian

dblp:70/916 · DBLP profile ↗
← Back
63ranked-venue papers
12as first author
41since 2021 · last 2026
0000-0002-5551-9796ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 1 first-author · 22 since 2021Systems, architecture and hardware · 16 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Security and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs
abstract
Large language models (LLMs) have advanced to encompass extensive knowledge across diverse domains. Yet controlling what a LLMs should not know is important for ensuring alignment and thus safe use. However, effective unlearning in LLMs is difficult due to the fuzzy boundary between knowledge retention and forgetting. This challenge is exacerbated by entangled parameter spaces from continuous multi-domain training, often resulting in collateral damage, especially under aggressive unlearning strategies. Furthermore, the computational overhead required to optimize State-of-the-Art (SOTA) models with billions of parameters poses an additional barrier. In this work, we present ALTER, a lightweight unlearning framework for LLMs to address both the challenges of knowledge entanglement and unlearning efficiency. ALTER operates through two phases: (I) high entropy tokens are captured and learned via the shared A matrix in LoRA, followed by (II) an asymmetric LoRA architecture that achieves a specified forgetting objective by parameter isolation and unlearning tokens within the target subdomains. Serving as a new research direction for achieving unlearning via token-level isolation in the asymmetric framework. ALTER achieves SOTA performance on TOFU, WMDP, and MUSE benchmarks with over 95% forget quality and shows minimal side effects through preserving foundational tokens. By decoupling unlearning from LLMs' billion-scale parameters, this framework delivers excellent efficiency while preserving over 90% of model utility, exceeding baseline preservation rates of 47.8-83.6%.
Xunlei Chen, Jinyu Guo, Yuang Li, Zhaokun Wang, Jie Zou 0001, Jiwei Wei, Wenhong Tian
AAAI8
2026 AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache Reuse
abstract
Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Ruiqi Wu, Zhaokun Wang, Wenyi Li, Wenhong Tian. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Zhaokun Wang, Wenhong Tian
ACL (1)8
2026 CAP: Controllable Alignment Prompting for Unlearning in LLMs
abstract
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Xunlei Chen, Jie Ou, Guangchun Luo, Wenhong Tian
ACL (1)10
2026 Adaptive probabilistic transformer for medical image segmentation
Tahir Kamal, Wenhong Tian, Jinyu Guo, Muhammad Shafiq 0006, Shuaihong Jiang, Ruini Xue
Expert Syst. Appl.2
2026 Seeking commonality while preserving diversity: A differential dual-path MoE solution for Multilingual Neural Machine Translation
Jinyu Guo, Yuang Li, Wenxian Liu, Zhaokun Wang, Wenhong Tian
Knowl. Based Syst.7
2026 Accelerating long-context inference of large language models via dynamic attention load balancing
Jie Ou, Jinyu Guo, Shuaihong Jiang, Ruini Xue, Wenhong Tian, Rajkumar Buyya
Knowl. Based Syst.6
2025 Efficient and Automatic 3D Parallelism Strategies Search via Contrastive Reinforcement Learning Pretrained Neural Networks
Jie Ou, Xiaowang Li, Yueming Chen, Jiahong Qian, Teng Su, Wenhong Tian
DSAA6
2025 Enhancing Large Language Model Inference Efficiency via Lookahead Cache Filtering
abstract
The large Key-Value (KV) cache is a significant challenge in deploying Large Language Models (LLMs). Current research addressing these issues employs cache compression techniques, which we find suffer from information loss and the "lost-in-the-middle" problem. We propose the Lookahead Cache Filtering (LCF), which retains the full cache in host memory to keep complete details, using the Important Key-Value Lookahead Prediction by Approximate Sorting and the Gather-based Matrix Multiplication on CPU, to filter important KV cache to reduce the overhead of loading cache to GPU and improve the throughput of inference. Finally, we use the Multi-Scale Pyramid Information Fusion mechanism to enhance information fusion to further improve the effectiveness of inference results. Experiments demonstrate that LCF effectively maintains accuracy without introducing additional latency while reducing memory requirements. Our code will be available on GitHub1.
Jie Ou, Yueming Chen, Shuaihong Jiang, Wenhong Tian
ICASSP4
2025 A Neural Network-Based Pipeline Parallel Strategy Solver for Heterogeneous Environments
abstract
The widespread application of large language models(LLMs) has made distributed training increasingly important, especially pipeline parallelism, which is a fundamental technique for ultra-large-scale LLMs. Current research in this field mainly employs combinatorial optimization algorithms such as dynamic programming. However, as the problem size increases, these methods become difficult to solve quickly in large-scale scenarios due to their high search time. Online optimization algorithms that combine neural networks with reinforcement learning require real-time interaction with the cluster environment to obtain feedback, resulting in high resource overhead and low search efficiency. Moreover, current research lacks studies on heterogeneous computing environments, which are frequently used by small research teams. To address these issues, we designed a novel Neural Network-based Pipeline Parallel strategy solver (NN-Piper) for heterogeneous environments. NN-Piper can perceive computational and communication costs, the number of stages to be divided, and the number of micro-batches. In addition, it can directly provide the strategy for allocating specific devices to each pipeline stage. To avoid an online training process that requires interaction with the cluster environment, we propose the Virtual Contrastive Training Algorithm (VCTA) to enable efficient training of NN-Piper without collecting large amounts of real data. After training, NN-Piper can be transferred to many different scenarios without further training or fine-tuning, and it can search for strategies within a few dozen milliseconds. Compared with the state-of-the-art method, NN-Piper can improve the training speed on average by 16-25% in different environments for the transformer-based models.
Jie Ou, Jinyu Guo, Yueming Chen, Shuaihong Jiang, Ruini Xue, Wenhong Tian
IJCNN6
2025 Low-Rank Decomposition Assisted Quantization and Inference Compensation for Quality Large Language Model Inference
abstract
Large Language Models (LLMs) have demonstrated exceptional performance on natural language processing tasks. However, these models are computationally intensive and require substantial hardware resources for deployment. Quantization has emerged as a popular technique for LLM deployment, reducing memory requirements, but it results in accuracy degradation, particularly when using low-bit quantization. To mitigate this accuracy loss, we introduce Low-Rank Compensation (LoRC), a novel compensation mechanism that aims to recover the performance drop caused by quantization. Additionally, we propose Low-Rank Quantization (LoRQ), which further reduces the quantization-induced loss by adaptively adjusting weights at the element-wise level to help LLMs accommodate quantized computations. LoRC focuses on compensating for accuracy loss during inference, LoRQ integrates low-rank compensation directly into the quantization process, and they do not need end-to-end fine-tuning with LLM. Furthermore, we propose the Rank-α Addition Strategy (RαAS) to combine LoRC into the inference framework, which improves inference accuracy without increasing inference latency. Experimental results show that our method outperforms the state-of-the-art OmniQuant by 1.89% on several common zero-shot datasets under the W4A4 setting of the widely-used LLaMA. Through the joint design of algorithms and systems, our techniques can be easily integrated into the FlexGen inference framework without introducing additional inference latency, thereby maintaining high throughput while improving accuracy.
Jie Ou, Jinyu Guo, Shuaihong Jiang, Zhaokun Wang, Yueming Chen, Ruini Xue, Wenhong Tian
IJCNN7
2025 Noise-Robustness Through Noise: A Framework combining Asymmetric LoRA with Poisoning MoE
abstract
Current parameter-efficient fine-tuning methods for adapting pre-trained language models to downstream tasks are susceptible to interference from noisy data. Conventional noise-handling approaches either rely on laborious data pre-processing or employ model architecture modifications prone to error accumulation. In contrast to existing noise-process paradigms, we propose a noise-robust adaptation method via asymmetric LoRA poisoning experts (LoPE), a novel framework that enhances model robustness to noise only with generated noisy data. Drawing inspiration from the mixture-of-experts architecture, LoPE strategically integrates a dedicated poisoning expert in an asymmetric LoRA configuration. Through a two-stage paradigm, LoPE performs noise injection on the poisoning expert during fine-tuning to enhance its noise discrimination and processing ability. During inference, we selectively mask the dedicated poisoning expert to leverage purified knowledge acquired by normal experts for noise-robust output. Extensive experiments demonstrate that LoPE achieves strong performance and robustness purely through the low-cost noise injection, which completely eliminates the requirement of data cleaning.
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Lingfeng Chen, Hongli Pu, Jie Ou, Libo Qin 0001, Wenhong Tian
NeurIPS8
2025 Multi-Perspective Dialogue Non-Quota Selection with loss monitoring for dialogue state tracking
Jinyu Guo, Zhaokun Wang, Jingwen Pu, Wenhong Tian, Guiduo Duan, Guangchun Luo
Expert Syst. Appl.4
2025 Cross-Search With Improved Multi-Dimensional Dichotomy-Based Joint Optimization for Distributed Parallel Training of DNN
abstract
istributedistributedD parallel training of large-scale deep neural networks (DNN) has attracted the attentions of both artificial intelligence and high-performance distributed computing. One of efficient approaches is the micro-batch-based pipeline parallelism (MBPP), e.g., GPipe and Terapipe. Based on the MBPP, we establish a time-cost model with the basic time function of layers, which considers computing time and communication time simultaneously as well as considers they are nonlinear with the amount of input data. Focusing on the jointly optimal solutions of network division and data partition, we propose a Cross-Search algorithm with Improved Multi-dimensional Dichotomy (CSIMD). Through theoretical derivation, we prove improved multi-dimensional dichotomy (IMD) has appreciable theoretical optimality and linear computational complexity significantly faster than the state-of-the-art methods including dynamic programming and recursive algorithm. Extensive experiments on both CNN-based and transformer-based neural networks demonstrate our proposed CSIMD can obtain optimal network division and data partition schemes under MBPP. On average, the training speeds of CSIMD in CNN- and transformer-based DNNs are respectively (2.0, 2.5)× and (2.66, 5.48)× of (MBPP-R, MBPP-E).
Yiqin Fu, Haocheng Lan, Yuanlun Xie, Wenhong Tian, Rajkumar Buyya, Jianhong Qian, Teng Su
IEEE Trans. Parallel Distributed Syst.5
2025 UMPIPE: Unequal Microbatches-Based Pipeline Parallelism for Deep Neural Network Training
abstract
The increasing need for large-scale deep neural networks (DNN) has made parallel training an area of intensive focus. One effective method, microbatch-based pipeline parallelism (notably GPipe), accelerates parallel training in various architectures. However, existing parallel training architectures normally use equal data partitioning (EDP), where each layer's process maintains identical microbatch-sizes. EDP may hinder training speed because different processes often require varying optimal microbatch-sizes. To address this, we introduce UMPIPE, a novel framework for unequal microbatches-based pipeline parallelism. UMPIPE enables unequal data partitions (UEDP) across processes to optimize resource utilization. We develop a recurrence formula to calculate the time cost in UMPIPE by considering both computation and communication processes. To further enhance UMPIPE's efficiency, we propose the Dual-Chromosome Genetic Algorithm for UMPIPE (DGAP) that accounts for the independent time costs of forward and backward propagation. Furthermore, we present TiDGAP, a two-level improvement on DGAP. TiDGAP accelerates the process by simultaneously calculating the end time for multiple individuals and microbatches using matrix operations. Our extensive experiments validate the dual-chromosome strategy's optimization benefits and TiDGAP's acceleration capabilities. TiDGAP can achieve better training schemes than baselines, such as the local greedy algorithm and the global greedy-based dynamic programming. Compared to (GPipe, PipeDream), UMPIPE achieves increases in training speed:$(13.89,11.09)\%$for GPT1-14,$(17.11, 7.96)\%$for VGG16 and$\geq (170,100)\%$for simulation networks.
Wenhong Tian, Rajkumar Buyya, Kui Wu 0001
IEEE Trans. Parallel Distributed Syst.2
2024 CSIMD: Cross-Search Algorithm with Improved Multi-dimensional Dichotomy for Micro-Batch-Based Pipeline Parallel Training in DNN
Haocheng Lan, Yuanlun Xie, Wenhong Tian, Jiahong Qian, Teng Su
Euro-Par (2)4
2024 Extracting Spatio-Temporal Coupling Feature of Patches for Long-Term Multivariate Time Series Forecasting
Weigang Huo, Yilang Deng, Yuanlun Xie, Zhaokun Wang, Wenhong Tian
ICIC (4)6
2024 Improving Chinese Emotion Classification Based on Bilingual Feature Fusion
Haocheng Lan, Jie Ou, Zhaokun Wang, Wenhong Tian
ICPR (31)4
2024 A Supervised Domain Adaptation Method with Alignment Regularization for Low-Light Facial Expression Recognition
Zhaokun Wang, Yuanlun Xie, Jie Ou, Jiahui Zhong, Wenhong Tian
PRCV (3)5
2024 A joint learning method with consistency-aware for low-resolution facial expression recognition
Yuanlun Xie, Wenhong Tian, Ruini Xue, Zhiyuan Zha, Bihan Wen
Expert Syst. Appl.2
2024 Compound facial expressions recognition approach using DCGAN and CNN
Jie Ou, Yuanlun Xie, Wenhong Tian
Multim. Tools Appl.4
2023 Robust facial expression recognition with Transformer Block Enhancement Module
Yuanlun Xie, Wenhong Tian, Zitong Yu
Eng. Appl. Artif. Intell.2
2023 Multi-search-routes-based methods for minimizing makespan of homogeneous and heterogeneous resources in Cloud computing
Wenhong Tian, Rajkumar Buyya
Future Gener. Comput. Syst.2
2023 A CNN Compression Method via Dynamic Channel Ranking Strategy
abstract
In recent years, the rapid development of mobile devices and embedded system raises a demand for intelligent models to address increasingly complicated problems. However, the complexity of the structure and extensive parameters press significantly on efficiency, storage space, and energy consumption. Additionally, the explosive growth of tasks with enormous model structures and parameters makes it impossible to compress models manually. Thus, a standardized and effective model compression solution achieving lightweight neural networks is established as an urgent demand by the industry. Accordingly, Dynamic Channel Ranking Strategy (DCRS) method is proposed to compress deep convolutional neural networks. DCRS selects channels with high contribution of each prunable layer according to compression ratio searched by reinforcement learning agent. Compared with current model compression methods, DCRS efficaciously applies various channel ranking strategies on prunable layers. Experiments indicate with a 50% compression ratio, compressed MobileNet achieved 70.62% top1 and 88.2% top5 accuracy on ImageNet, and compressed ResNet achieved 92.03% accuracy on CIFAR-10. DCRS reduces more FLOPS in these neural networks. The compressed model achieves the best Top-1 and Top-5 accuracy on ResNet50, the best Top-1 accuracy on MobilNetV1.
Ruiming Wen, Yuanlun Xie, Wenhong Tian
Int. J. Comput. Intell. Appl.4
2023 Prepartition: Load Balancing Approach for Virtual Machine Reservations in a Cloud Data Center
Wenhong Tian, Minxian Xu, Kui Wu 0001, Cheng-Zhong Xu 0001, Rajkumar Buyya
J. Comput. Sci. Technol.1
2023 Facial expression recognition through multi-level features extraction and fusion
Yuanlun Xie, Wenhong Tian, Hengxin Zhang, Tingsong Ma
Soft Comput.2
2022 SPAC: Scalable Pattern Approximate Counting in Graph Mining
Ruini Xue, Shengbo Liu, Wenhong Tian
ICA3PP5
2022 Optimal distributed parallel algorithms for deep learning framework Tensorflow
Yuanlun Xie, Majun He, Tingsong Ma, Wenhong Tian
Appl. Intell.4
2022 Workload forecasting and energy state estimation in cloud data centres: ML-centric approach
Tahseen Khan, Wenhong Tian, Shashikant Ilager, Rajkumar Buyya
Future Gener. Comput. Syst.2
2022 5G-Enabled Medical Data Transmission in Mobile Hospital Systems
abstract
As one of the healthcare development trends, mobile hospital systems nowadays need more advanced means of treatment and diagnosis and require higher data transmission speed and capacity with lower latency. With its ability to simultaneously support a variety of services in a wide range of application scenarios using network slicing technique, the emerging fifth generation communication system (5G) is expected to offer significant communication performances, including higher throughput, higher mobility, higher connection density, higher transmission reliability with lower latency than the current fourth generation (4G). Hence, 5G can serve as an effective transmission means to meet the modern mobile hospital systems requirements. This article proposes a 5G network slicing-based mobile hospital system where two types of slices, namely, enhanced mobile broadband (eMBB) slice and ultrareliable and low-latency communications (uRLLC) slice, are dedicated to the medical data. We propose an optimization method to maximize the throughput of the medical data assigned to the eMBB slice. We also propose a resource allocation algorithm for very high transmission reliability with very low latency of the medical data assigned to the uRLLC slice. Simulation results indicate that our proposed approach can effectively meet mobile hospital systems requirements for data transmission throughput, reliability, and latency from remote sites to hospital centers.
Parfait I. Tebe, Guangjun Wen, Jian Li 0060, Yongjun Yang, Wenhong Tian, Jing Chong
IEEE Internet Things J.5
2022 Machine learning (ML)-centric resource management in cloud computing: A review and future directions
Tahseen Khan, Wenhong Tian, Shashikant Ilager, Mingming Gong, Rajkumar Buyya
J. Netw. Comput. Appl.2
2022 Deep reinforcement learning-based algorithms selectors for the resource scheduling in hierarchical Cloud computing
Ruiming Wen, Wenhong Tian, Rajkumar Buyya
J. Netw. Comput. Appl.3
2022 Multi-level knowledge distillation for low-resolution object detection and facial expression recognition
Tingsong Ma, Wenhong Tian, Yuanlun Xie
Knowl. Based Syst.2
2021 A Novel Cluster Ensemble based on a Single Clustering Algorithm
abstract
In recent years, several cluster ensemble methods have been developed, but they still have some limitations.They commonly use different clustering algorithms in both stages of the clustering ensemble method, such as the ensemble generation step and the consensus function, resulting in a compatibility issue in terms of working functionality between different clustering algorithms.In addition, in a clustering ensemble method, the accuracy of the final results is a major concern.To deal with it, we propose a novel cluster ensemble method based on a single clustering algorithm (CES).In this method, we iterate a clustering algorithm affinity propagation (AP) ten times in the ensemble generation step to obtain multiple base partitions with a high level of diversity in each iteration due to its nature of producing a random number of clusters.Furthermore, with a few modifications, the same algorithm AP is used to propose a novel consensus function for combining these base partitions into a single partition.The proposed consensus function takes advantage of little side-information in the form of partial labels by using pairwise constraints with AP and number of clusters in a dataset.By employing this information, AP is limited to produce an actual number of cluster centres in a dataset rather than a random number of clusters, which considerably enhanced the accuracy of final outcomes.As a result, CES uses the same clustering functionality in both stages of proposed cluster ensemble method and produces the desired number of clusters in the final partition of a dataset which is significantly improving accuracy when compared to state-of-the-art cluster ensemble methods.Furthermore, as a result of these modifications, the CES outperforms AP in terms of accuracy and execution time.Experiments on real-world datasets from various sources show that CES improves accuracy by 5% on average compared to state-of-the-art cluster ensemble methods and by 55.54% compared to AP while consuming 44.60% less execution time.
Tahseen Khan, Wenhong Tian, Kadhim Mustafa Raad Kadhim, Rajkumar Buyya
FedCSIS2
2021 Detecting Anomaly Features in Facial Expression Recognition
Tingsong Ma, Yuanlun Xie, Hengxin Zhang, Wenhong Tian
ICONIP (5)6
2021 Toward perfect neural cascading architecture for grammatical error correction
Kingsley Nketia Acheampong, Wenhong Tian
Appl. Intell.2
2021 Cloud Resource Scheduling With Deep Reinforcement Learning and Imitation Learning
abstract
The cloud resource management belongs to the category of combinatorial optimization problems, most of which have been proven to be NP-hard. In recent years, reinforcement learning (RL), as a special paradigm of machine learning, has been used to tackle these NP-hard problems. In this article, we present a deep RL-based solution, called DeepRM_Plus, to efficiently solve different cloud resource management problems. We use a convolutional neural network to capture the resource management model and utilize imitation learning in the reinforcement process to reduce the training time of the optimal policy. Compared with the state-of-the-art algorithm DeepRM, DeepRM_Plus is 37.5% faster in terms of the convergence rate. Moreover, DeepRM_Plus reduces the average weighted turnaround time and the average cycling time by 51.85% and 11.51%, respectively.
Wenxia Guo, Wenhong Tian, Lingxiao Xu, Kui Wu 0001
IEEE Internet Things J.2
2021 An anchor-free object detector with novel corner matching method
Tingsong Ma, Wenhong Tian, Ping Kuang, Yuanlun Xie
Knowl. Based Syst.2
2021 Classical and modern face recognition approaches: a complete review
Waqar Ali 0001, Wenhong Tian, Desire Iradukunda, Abdullah Aman Khan
Multim. Tools Appl.2
2021 A frequency-aware and energy-saving strategy based on DVFS for Spark
Yaojun Wei, Enjie Ma, Wenhong Tian
J. Supercomput.5
2021 Nonfragile Sampled-Data Filtering of Uncertain Fuzzy Systems With Time-Varying Delays
abstract
This article studies the nonfragile sampled-data filtering problems for Takagi-Sugeno fuzzy systems with uncertainties and delays. By introducing two delay-product-type terms and other augmented terms, a novel Lyapunov-Krasovskii functional containing more detailed information of time delays is constructed. With the help of some new bounding inequalities, improved stability results for the addressed systems are established. Besides, a resilient sampled-data controller is devised with less cost. Furthermore, the efficiency of the obtained criteria is illustrated by a numerical example.
Jinnan Luo, Xinzhi Liu, Wenhong Tian, Shouming Zhong, Kaibo Shi
IEEE Trans. Syst. Man Cybern. Syst.3
2021 Back-projection-based progressive growing generative adversarial network for single image super-resolution
Tingsong Ma, Wenhong Tian
Vis. Comput.2
2020 Energy Efficient Algorithms based on VM Consolidation for Cloud Computing: Comparisons and Evaluations
abstract
Cloud Computing paradigm has revolutionized IT industry and be able to offer computing as the fifth utility. With the pay-as-you-go model, cloud computing enables to offer the resources dynamically for customers anytime. Drawing the attention from both academia and industry, cloud computing is viewed as one of the backbones of the modern economy. However, the high energy consumption of cloud data centers contributes to high operational costs and carbon emission to the environment. Therefore, Green cloud computing is required to ensure energy efficiency and sustainability, which can be achieved via energy efficient techniques. One of the dominant approaches is to apply energy efficient algorithms to optimize resource usage and energy consumption. Currently, various virtual machine consolidation-based energy efficient algorithms have been proposed to reduce the energy of cloud computing environment. However, most of them are not compared comprehensively under the same scenario, and their performance is not evaluated with the same experimental settings. This makes users hard to select the appropriate algorithm for their objectives. To provide insights for existing energy efficient algorithms and help researchers to choose the most suitable algorithm, in this paper, we compare several state-of-the-art energy efficient algorithms in depth from multiple perspectives, including architecture, modelling and metrics. In addition, we also implement and evaluate these algorithms with the same experimental settings in CloudSim toolkit. The experimental results show the performance comparison of these algorithms with comprehensive results. Finally, detailed discussions of these algorithms are provided.
Qiheng Zhou, Minxian Xu, Sukhpal Singh, Chengxi Gao, Wenhong Tian, Cheng-Zhong Xu 0001, Rajkumar Buyya
CCGRID5
2020 A Reinforcement Learning Based Approach to Identify Resource Bottlenecks for Multiple Services Interactions in Cloud Computing Environments
Lingxiao Xu, Minxian Xu, Richard Semmes, Hong Mu, Shuangquan Gui, Wenhong Tian, Kui Wu 0001, Rajkumar Buyya
CollaborateCom (2)7
2020 An improved recurrent neural networks for 3d object reconstruction
Tingsong Ma, Ping Kuang, Wenhong Tian
Appl. Intell.3
2020 Handling data skew at reduce stage in Spark by ReducePartition
abstract
Summary As a typical representative of distributed computing framework, Spark has been continuously developed and popularized. It reduces the data transmission time through efficient memory‐based operations and solves the shortcomings of the traditional MapReduce computation model in iterative computation. In Spark, data skew is very prominent due to the uneven distribution of input data and the unbalanced allocation of default partitioning algorithm. When data skew occurs, the execution efficiency of the program will be reduced, especially in the reduce stage of Spark. Therefore, this paper proposes ReducePartition to solve data skew problem at reduce stage of Spark platform. First, the compute node samples the local data according to the sampling algorithm to predict the overall characteristics of data distribution. Then, to take full use of cluster resources, ReducePartition divides data into multiple partitions evenly. Next, taking into account the differences in computational capabilities among Executors, each task is assigned to Executor with the highest performance factor according to the greedy strategy. Finally, the results of the related algorithms and ReducePartition are compared by using WordCount benchmark and Sort benchmarks on heterogeneous Spark standalone cluster. The performance of the ReducePartition under different degree of data skew and different data size is analyzed. Experimental results show that the proposed algorithm can effectively reduce the impact of data skew on the total makespan of Spark big data applications, and the average total makespan is reduced by 30% to 50% while resource utilization is increased by 20%‐30% on average.
Wenxia Guo, Chaojie Huang, Wenhong Tian
Concurr. Comput. Pract. Exp.3
2020 Arbitrary Back-Projection Networks for Image Super-Resolution
abstract
Recently, a method called Meta-SR has solved the problem of super-resolution of arbitrary scale factor with only one single model. However, it has a limited reconstruction accuracy compared with RDN[Formula: see text] and EDSR[Formula: see text]. Inspired by Meta-SR, we noticed that by combining the core idea of Meta-SR and D-DBPN, we might construct a network that has as good image reconstruction accuracy as D-DBPN’s, at the same time, keeps arbitrary scaling function. According to Meta-SR’s Meta-Upscale Module, we designed a different structure called Meta-Downscale Module. By using these two different modules and back-projection structure, we construct an arbitrary back-projection network, which has the ability to enlarge images with arbitrary scale factor by using only one single model, meanwhile, obtains state-of-the-art reconstruction results. Through extensive experiments, our proposed method performs better reconstruction effect than Meta-SR and more efficient than D-DBPN. Besides that, we also evaluated the proposed method on widely used benchmark dataset on single image super-resolution. The experimental results show the superiority of our model compared to RDN+ and EDSR+.
Tingsong Ma, Wenhong Tian
Int. J. Comput. Intell. Appl.2
2019 Autonomic decentralized elasticity based on a reinforcement learning controller for cloud applications
Seyed Mohammad Reza Nouri, Han Li 0002, Srikumar Venugopal, Wenxia Guo, MingYun He, Wenhong Tian
Future Gener. Comput. Syst.6
2019 SAVE: self-adaptive consolidation of virtual machines for energy efficiency of CPU-intensive applications in the cloud
Wenxia Guo, Ping Kuang, Yaqiu Jiang, Wenhong Tian
J. Supercomput.5
2018 On Minimizing the Makespan of a Set of Offline MapReduce Jobs
abstract
There are quite a few algorithms on minimizing the makespan of a set of offline MapReduce jobs. However, these algorithms are heuristic or suboptimal. The best-known algorithm for minimizing the makespan is 3-approximation by applying Johnson model. In this paper, we evaluate three algorithms in terms of approximation, computational complexity, stability and optimal configuration of the underline cluster for the first time.
Majun He, Wenxia Guo, Houwen Huang, Wenhong Tian
SERVICES6
2018 On minimizing total energy consumption in the scheduling of virtual machine reservations
Wenhong Tian, Majun He, Wenxia Guo, Wenqiang Huang, Xiaoyu Shi 0001, Mingsheng Shang 0001, Adel Nadjaran Toosi, Rajkumar Buyya
J. Netw. Comput. Appl.1
2017 A Near Optimal Approach for Symmetric Traveling Salesman Problem in Euclidean Space
Wenhong Tian, Chaojie Huang
ICORES1
2017 A survey on load balancing algorithms for virtual machines placement in cloud computing
abstract
Summary The emergence of cloud computing based on virtualization technologies brings huge opportunities to host virtual resource at low cost without the need of owning any infrastructure. Virtualization technologies enable users to acquire, configure, and be charged on pay‐per‐use basis. However, cloud data centers mostly comprise heterogeneous commodity servers hosting multiple virtual machines (VMs) with potential various specifications and fluctuating resource usages, which may cause imbalanced resource utilization within servers that may lead to performance degradation and service level agreements violations. So as to achieve efficient scheduling, these challenges should be addressed and solved by using load balancing strategies, which have been proved to be nondeterministic polynomial time (NP)‐hard problem. From multiple perspectives, this work identifies the challenges and analyzes existing algorithms for allocating VMs to hosts in infrastructure clouds, especially focuses on load balancing. A detailed classification targeting load balancing algorithms for VM placement in cloud data centers is investigated, and the surveyed algorithms are classified according to the classification. The goal of this paper is to provide a comprehensive and comparative understanding of existing literature and aid researchers by providing an insight for potential future enhancements.
Minxian Xu, Wenhong Tian, Rajkumar Buyya
Concurr. Comput. Pract. Exp.2
2016 HScheduler: an optimal approach to minimize the makespan of multiple MapReduce jobs
Wenhong Tian, Guozhong Li 0001, Wutong Yang, Rajkumar Buyya
J. Supercomput.1
2015 Minimizing total busy time in offline parallel scheduling with application to energy efficiency in cloud computing
abstract
Summary Our paper considers the following fundamental scheduling problem. There are n deterministic jobs to be scheduled offline on multiple identical machines, which have bounded capacities. Each job is associated with a start‐time, an end‐time, a process time, and demand for machine capacity. The goal is to schedule all of the jobs non‐preemptively in their start‐time‐end‐time windows, subject to machine capacity constraints such that the total busy time of the machines is minimized. We refer to this problem as minimizing the total busy time for the scheduling of multiple identical machines (MinTBT). This problem has important applications in power‐aware scheduling for Cloud computing, optical network design, customer service systems, and other related areas. Scheduling to minimize busy times is already NP‐hard in the special case where all jobs have the same process time and can be scheduled in a fixed time interval. One best‐known result for this problem is a 5‐approximation algorithm for special instances using first‐fit‐decreasing algorithm. In this paper, we propose and prove a 3‐approximation algorithm, modified first‐fit‐decreasing‐earliest for the general case and obtain more results for special cases. We then show how our results are applied in cloud computing to improve the energy efficiency. Copyright © 2013 John Wiley & Sons, Ltd.
Wenhong Tian, Chee Shin Yeo
Concurr. Comput. Pract. Exp.1
2015 Enabling scalable scientific workflow management in the Cloud
Yong Zhao 0009, Youfu Li 0002, Ioan Raicu, Shiyong Lu, Wenhong Tian, Heng Liu 0004
Future Gener. Comput. Syst.5
2015 A Toolkit for Modeling and Simulation of Real-Time Virtual Machine Allocation in a Cloud Data Center
abstract
Resource scheduling in infrastructure as a service (IaaS) is one of the keys for large-scale Cloud applications. Extensive research on all issues in real environment is extremely difficult because it requires developers to consider network infrastructure and the environment, which may be beyond the control. In addition, the network conditions cannot be predicted or controlled. Therefore, performance evaluation of workload models and Cloud provisioning algorithms in a repeatable manner under different configurations and requirements is difficult. There is still lack of tools that enable developers to compare different resource scheduling algorithms in IaaS regarding both computing servers and user workloads. To fill this gap in tools for evaluation and modeling of Cloud environments and applications, we propose CloudSched. CloudSched can help developers identify and explore appropriate solutions considering different resource scheduling algorithms. Unlike traditional scheduling algorithms considering only one factor such as CPU, which can cause hotspots or bottlenecks in many cases, CloudSched treats multidimensional resource such as CPU, memory and network bandwidth integrated for both physical machines and virtual machines (VMs) for different scheduling objectives (algorithms). In this paper, two existing simulation systems at application level for Cloud computing are studied, a novel lightweight simulation system is proposed for real-time VM scheduling in Cloud data centers, and results by applying the proposed simulation system are analyzed and discussed.
Wenhong Tian, Yong Zhao 0009, Minxian Xu, Yuanliang Zhong, Xiashuang Sun
IEEE Trans Autom. Sci. Eng.1
2015 A Service Framework for Scientific Workflow Management in the Cloud
abstract
Cloud computing is an emerging computing paradigm that can offer unprecedented scalability and resources on demand, and is getting more and more adoption in the science community, while scientific workflow management systems provide essential support such as management of data and task dependencies, job scheduling and execution, provenance tracking, etc., to scientific computing. As we are entering into a “big data” era, it is imperative to migrate scientific workflow management systems into the cloud to manage the ever increasing data scale and analysis complexity. We propose a reference service framework for integrating scientific workflow management systems into various cloud platforms, which consists of eight major components, including Cloud Workflow Management Service, Cloud Resource Manager, etc., and six interfaces between them. We also present a reference framework for the implementation of Cloud Resource Manager, which is responsible for the provisioning and management of virtual resources in the cloud. We discuss our implementation of the framework by integrating the Swift scientific workflow management system with the OpenNebula and Eucalyptus cloud platforms, and demonstrate the capability of the solution using a NASA MODIS image processing workflow and a production deployment on the Science@Guoshi network with support for the Montage image mosaic workflow.
Yong Zhao 0009, Youfu Li 0002, Ioan Raicu, Shiyong Lu, Cui Lin, Wenhong Tian, Ruini Xue
IEEE Trans. Serv. Comput.7
2014 Prepartition: A new paradigm for the load balance of virtual machine reservations in data centers
abstract
It is significant to apply load-balancing strategy to improve the performance and reliability of resource in data centers. One of the challenging scheduling problems in Cloud data centers is to take the allocation and migration of reconfigurable virtual machines (VMs) as well as the integrated features of hosting physical machines (PMs) into consideration. In the reservation model, workload of data centers has fixed process interval characteristics. In general, load-balance scheduling is NP-hard problem as proved in many open literatures. Traditionally, for offline load balance without migration, one of the best approaches is LPT (Longest Process Time first), which is well known to have approximation ratio 4/3. With virtualization, reactive (post) migration of VMs after allocation is one popular way for load balance and traffic consolidation. However, reactive migration has difficulty to reach predefined load balance objectives, and may cause interruption and instability of service and other associated costs. In view of this, we propose a new paradigm-Prepartition: it proactively sets process-time bound for each request on each PM and prepares in advance to migrate VMs to achieve the predefined balance goal. Prepartition can reduce process time by preparing VM migration in advance and therefore reduce instability and achieve better load balance as desired. Trace-driven and synthetic simulation results show that Prepartition has 10%-20% better performance than the well known load balancing algorithms with regard to average CPU utilization, makespan as well as capacity makespan.
Wenhong Tian, Minxian Xu, Yong Zhao 0009
ICC1
2014 A Note on "Orchestrating an Ensemble of MapReduce Jobs for Minimizing Their Makespan"
abstract
In paper [1], a scheduling model is considered for multiple MapReduce jobs. The goal in [1] is to design an automatic job scheduler that minimizes the makespan of such a set of MapReduce jobs. In this work, we find that there is a key assumption in [1] which leads to the violation of the conditions for classical Johnson's algorithm and a suboptimal job scheduling for minimizing total makespan. By considering a better strategy and implementation, we can still meet the conditions of classical Johnson's algorithm. Then we can still use Johnson's algorithm for an optimal solution. As for BalancedPools algorithm proposed in paper [1], under our proposed new strategy, it is possible to solve it exactly in linear time, but not NP-hard as suggested in [1], the proof is provided. With the new strategy, results obtained in [1] need reevaluating.
Wenhong Tian
IEEE Trans. Dependable Secur. Comput.1
2013 An Energy-Efficient Online Parallel Scheduling Algorithm for Cloud Data Centers
abstract
This paper considers online energy-efficient scheduling of real-time virtual machines (VMs) for Cloud data centers. Each request is associated with a starttime, a end-time, a processing time and demand for a Physical Machine (PM) capacity. The goal is to schedule all of the requests non-preemptively in their start-timeend- time windows, subjecting to PM capacity constraints, such that total busy time of all used PMs is minimized (called MinTBT-ON for abbreviation). This problem is a fundamental scheduling problem for parallel jobs allocation on mutliple machines, it has important applications in power-aware scheduling in cloud computing, optical network design and customer service systems and other related areas. Offline scheduling to minimize busy time is NP-hard already in the special case where all jobs have the same processing time and can be scheduled in a fixed time interval. One best-known result for MinTBT-ON problem is a g-competitive algorithm for general instances using First-Fit algorithm for unit-size jobs, where g is the total capacity of a PM. In this paper, a B-competitive algorithm, GRID is proposed and proved for general case, where B is a natural number and 1 < B < g. More results are obtained and applied to Cloud computing to improve energy-efficiency.
Wenhong Tian, Ruini Xue, Qin Xiong, Yunjun Hu
SERVICES1
2013 An online parallel scheduling method with application to energy-efficiency in cloud computing
Wenhong Tian, Qin Xiong
J. Supercomput.1
2009 Adaptive Dimensioning of Cloud Data Centers
abstract
Cloud data centers (CDCs) provide key infrastructure for cloud computing. Allocating computing resources is very important for the CDCs to function efficiently. Current allocation of resources in CDCs is mostly dedicated and static. However, workloads for cloud applications are highly variable which cause poor application performance, poor resource utilization or both. In this paper, adaptive dimensioning methods for CDCs are developed so that right amount of computing resources are allocated for variable workloads to meet quality of service requirements.
Wenhong Tian
DASC1
2006 A Dynamic Modeling And Dimensioning Approach For All-Optical Networks
abstract
In this paper we introduce a new dimensioning model for the optical network by the time-dependent blocking probability. This approach is motivated by [6] where the absorption probability based on transient analysis of continuous time Markov chain (CTMC) is used. Unlike the absorption probability model, the blocking probability has been used to dimensioning network for long time and has lower computation complexity. However, using traditional steady-state probability will cause overprovisioning in most cases. So we introduce a time-dependent blocking probability model which is based on the transient analysis of CTMC. Then we propose a novel closed-form transient analysis methods for a single link and and an arbitrary topology network. Applying this approach, the time-dependent provisioning can satisfy dynamically changing traffic demands and avoid overprovisioning problem in optical networks.
Wenhong Tian
BROADNETS1