Song Fu

dblp:57/2335 · DBLP profile ↗
← Back
102ranked-venue papers
20as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 16 since 2021Computer networks · 18 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 16 · 1 first-author · 8 since 2021Security and privacy · 12 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8Software engineering, systems software and programming languages · 3
YearPublicationVenuePosition
2026 Performance evaluation method for modules based on organic fusion of data-driven methods and mechanistic knowledge
Wenhui He, Lin Lin 0014, Song Fu
Adv. Eng. Informatics3
2026 Gene recombination-guided convolution neural network for early fault diagnosis of aero-engines
Jinlei Wu, Lin Lin 0014, Song Fu, Lingyu Yue, Sihao Zhang
Adv. Eng. Informatics4
2026 Engine-specific degradation prediction of aviation engines via transferable snippet augmentation: A trend-grouped fine-tuning perspective
Minghang Zhao, Song Fu
Adv. Eng. Informatics6
2026 MTGFormer: A novel multi-task gated transformer with multi-head selective fusion attention for aeroengine gas-path parameter deviations parallel prediction
Song Fu, Fazhan Han, Lin Lin 0014, Yue Wang 0087, Minghang Zhao
Expert Syst. Appl.1
2026 GIRMSF-global information reconstruction and multi-scale feature sharpening framework for knowledge graph embedding
Lin Lin 0014, Shiwei Suo, Song Fu, Lizheng Zu, Sihao Zhang
Expert Syst. Appl.3
2026 Privacy-Preserving Driver Monitoring on the Edges: Transformer-Based Processing of Secret Shares from Video Streams
abstract
Modern vehicles increasingly rely on advanced driver monitoring systems (DMS) to ensure safety and enhance the driving experience. These systems assess driver status to prevent accidents caused by fatigue, inattentiveness, or intoxication. While some DMS applications process video data on vehicle, many rely on edge or cloud-based solutions, raising significant privacy concerns due to the storage of sensor data from vehicles. Existing approaches, such as de-identification and homomorphic encryption, either impose heavy computational overhead on vehicles or insufficiently address privacy. To overcome these limitations, we present the Privacy-preserving Driver Monitoring System (PDMS), a novel framework based on the additive secret sharing theory and privacy-preserving Transformer-based deep learning models. PDMS creates randomized secret shares from driver’s facial video data on vehicle, processes them independently through privacy-preserving Transformer models on edges, and securely aggregates partial results on vehicle, ensuring vehicles’ sensor data and final results remain protected. This approach reduces the computational load on the vehicle, enabling cost-effective and scalable DMS solutions that protect the privacy of the driver both in transit and in processing. Our contributions include the design and optimization of the PDMS system, incorporating privacy-preserving DNN layers that are capable of processing randomized secret shares. Furthermore, we present a practical system that utilizes a vision transformer (ViT)-based gaze estimation model, demonstrating the effectiveness of PDMS through comprehensive experiments.
Tianyu Bai, Danyang Shao, Qing Yang 0003, Yunhe Feng, Song Fu
ACM Trans. Internet Things6
2025 Collaborative Tree Search for Enhancing Embodied Multi-Agent Collaboration
abstract
Embodied agents based on large language models (LLMs) face significant challenges in collaborative tasks, requiring effective communication and reasonable division of labor to ensure efficient and correct task completion. Previous approaches with simple communication patterns carry erroneous or incoherent agent actions, which can lead to additional risks. To address these problems, we propose Cooperative Tree Search (CoTS), a framework designed to significantly improve collaborative planning and task execution efficiency among embodied agents. CoTS guides multi-agents to discuss long-term strategic plans within a modified Monte Carlo tree, searching along LLM-driven reward functions to provide a more thoughtful and promising approach to cooperation. Another key feature of our method is the introduction of a plan evaluation module, which not only prevents agent action confusion caused by frequent plan updates but also ensures plan updates when the current plan becomes unsuitable. Experimental results show that the proposed method performs excellently in planning, communication, and collaboration on embodied environments (CWAH and TDW-MAT), efficiently completing long-term, complex tasks and significantly outperforming existing methods.
Lizheng Zu, Lin Lin 0014, Song Fu, Na Zhao 0004, Pan Zhou 0002
CVPR3
2025 GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
abstract
In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple modalities, including the point cloud (PC), RGB image, and depth. This allows GSOT3D to support various 3D tracking tasks, such as single-modal 3D SOT on PC and multi-modal 3D SOT on RGB-PC or RGB-D, and thus greatly broadens research directions for 3D object tracking. To provide highquality per-frame 3D annotations, all sequences are labeled manually with multiple rounds of meticulous inspection and refinement. To our best knowledge, GSOT3D is the largest benchmark dedicated to various generic 3D object tracking tasks. To understand how existing 3D trackers perform and to provide comparisons for future research on GSOT3D, we assess eight representative point cloud-based tracking models. Our evaluation results exhibit that these models heavily degrade on GSOT3D, and more efforts are required for robust and generic 3D object tracking. Besides, to encourage future research, we present a simple yet effective generic 3D tracker, named PROT3D, that localizes the target object via a progressive spatial-temporal network and outperforms all current solutions by a large margin. By releasing GSOT3D, we expect to advance further 3D tracking in future research and applications. Our benchmark and model as well as the evaluation results will be publicly released at our webpage https://github.com/ailovejinx/GSOT3D.
Yifan Jiao, Junhua Ding 0001, Qing Yang 0003, Song Fu, Heng Fan 0001, Libo Zhang 0001
ICCV5
2025 FD-LLM: Large language model for fault diagnosis of complex equipment
Sihao Zhang, Song Fu
Adv. Eng. Informatics3
2025 Continual contrastive reinforcement learning: Towards stronger agent for environment-aware fault diagnosis of aero-engines through long-term optimization under highly imbalance scenarios
Minghang Zhao, Song Fu
Adv. Eng. Informatics6
2025 An adaptive integrated learning-based virtual sensing framework for temperature prediction in aircraft brake monitoring engineering
Song Fu, Jinlei Wu
Eng. Appl. Artif. Intell.3
2025 PSTFormer: A novel parallel spatial-temporal transformer for remaining useful life prediction of aeroengine
Song Fu, Yiming Jia, Lin Lin 0014, Shiwei Suo, Sihao Zhang
Expert Syst. Appl.1
2025 Prototype matching-based meta-learning model for few-shot fault diagnosis of mechanical system
Lin Lin 0014, Sihao Zhang, Song Fu, Shiwei Suo, Guolei Hu
Neurocomputing3
2025 CUR-Estimator: Towards reliable missing data imputation for aero-engine degradation process
Minghang Zhao, Song Fu
Neurocomputing6
2025 Safeguarding user data privacy in online Large Language Model services
Tianyu Bai, Yunhe Feng, Song Fu
J. Syst. Archit.3
2025 Dinic's assembly parts matching optimization based on CACGAN gyroscope performance prediction
Wuyang Fan, Song Fu
J. Supercomput.2
2024 SiCP: Simultaneous Individual and Cooperative Perception for 3D Object Detection in Connected and Automated Vehicles
abstract
Cooperative perception for connected and automated vehicles is traditionally achieved through the fusion of feature maps from two or more vehicles. However, the absence of feature maps shared from other vehicles can lead to a significant decline in 3D object detection performance for cooperative perception models compared to standalone 3D detection models. This drawback impedes the adoption of cooperative perception as vehicle resources are often insufficient to concurrently employ two perception models. To tackle this issue, we present Simultaneous Individual and Cooperative Perception (SiCP), a generic framework that supports a wide range of the state-of-the-art standalone perception backbones and enhances them with a novel Dual-Perception Network (DP-Net) designed to facilitate both individual and cooperative perception. In addition to its lightweight nature with only 0.13M parameters, DP-Net is robust and retains crucial gradient information during feature map fusion. As demonstrated in a comprehensive evaluation on the V2V4Real and OPV2V datasets, thanks to DP-Net, SiCP surpasses state-of-the-art cooperative perception solutions while preserving the performance of standalone perception solutions. The source code can be found at https://github.com/DarrenQu/SiCP.
Deyuan Qu, Qi Chen 0018, Tianyu Bai, Hongsheng Lu, Heng Fan 0001, Song Fu, Qing Yang 0003
IROS7
2024 Channel attention & temporal attention based temporal convolutional network: A dual attention framework for remaining useful life prediction of the aircraft engines
Lin Lin 0014, Jinlei Wu, Song Fu, Sihao Zhang, Changsheng Tong, Lizheng Zu
Adv. Eng. Informatics3
2024 Integrating adversarial training strategies into deep autoencoders: A novel aeroengine anomaly detection framework
Lin Lin 0014, Lizheng Zu, Song Fu, Sihao Zhang, Shiwei Suo, Changsheng Tong
Eng. Appl. Artif. Intell.3
2024 SRSCL: A strong-relatedness-sequence-based fine-grained collective entity linking method for heterogeneous information networks
Lizheng Zu, Lin Lin 0014, Song Fu, Changsheng Tong
Expert Syst. Appl.4
2024 PathEL: A novel collective entity linking method based on relationship paths in heterogeneous information networks
Lizheng Zu, Lin Lin 0014, Song Fu, Shiwei Suo, Wenhui He, Jinlei Wu, Yancheng Lv
Inf. Syst.3
2024 SelectE: Multi-scale adaptive selection network for knowledge graph representation learning
Lizheng Zu, Lin Lin 0014, Song Fu, Jinlei Wu
Knowl. Based Syst.3
2023 $\mathrm{P}^{3}$: A Privacy-Preserving Perception Framework for Building Vehicle-Edge Perception Networks Protecting Data Privacy
abstract
With the wider adoption of edge computing services, intelligent edge devices, and high-speed V2X communication, compute-intensive tasks for autonomous vehicles, such as perception using camera, LiDAR, and/or radar data, can be partially offloaded to road-side edge units. However, data privacy becomes a major concern for vehicular edge computing, as sensor data with sensitive information from vehicles can be observed and used by edge servers. We aim to address the privacy problem by protecting both vehicles' sensor data and the detection results. In this paper, we present a privacy preserving perception$(\mathbf{P}^{3})$framework which provides a secure version of every commonly used layers in various perception CNN networks. They server as the building blocks to facilitate the construction of a privacy preserving CNN for any existing or future network.$\mathbf{P}^{3}$leverages the additive secret sharing theory to develop secure functions for perception networks. A vehicle's sensor data is split and encrypted into multiple secret shares, each of which is processed on an edge server by going through the secure layers of a detection network. The detection results can only be obtained by combining the partial results from the participating edge servers. We present two use cases where the secure layers in$\mathbf{P}^{3}$are used to build privacy preserving both single-stage and two-stage object detection CNNs. Experimental results indicate data privacy for vehicles is protected without comprising the detection accuracy and with a reasonable amount of performance degradation. To the best of our knowledge, this is the first work that provides a generic framework to ease the development of vehicle-edge perception networks protecting data privacy.
Tianyu Bai, Danyang Shao, Song Fu, Qing Yang 0003
ICCCN4
2023 Output-Directed Dynamic Quantization for DNN Acceleration
abstract
Quantization is an effective technique for reducing the number of computations and improving the performance of deep neural networks (DNNs). Weight quantization is popular because weights can be trained beforehand. However, weight quantization only targets the kernel weights and ignores the sensitivity of input features, which can lead to reduced accuracy. Fine-grained input quantization has gained attention as a way to speed up DNNs while maintaining accuracy. Existing approaches determine computation precision based on input sensitivity but do not effectively reduce computations for insensitive outputs or retain the precision of sensitive outputs. These limitations motivate us to develop an output-directed dynamic quantization method named ODQ in this paper. ODQ is a two-stage DNN quantization scheme designed to improve performance, reduce energy consumption, and maintain and often improve accuracy, compared with existing quantization methods. Specifically, inputs and weights go through sensitivity prediction and result generation. The high-order 2 bits of input and weight are used to predict output sensitivity. Result generation is performed only for predicted sensitive outputs. We designed an FPGA accelerator to optimize ODQ quantization performance for DNNs. We implement a prototype of ODQ and evaluate its performance using several state-of-the-art DNNs. Compared with a state-of-the-art input-directed quantization approach, ODQ achieves a 67.6% performance speedup and a 66.9% energy saving, with minimal accuracy degradation (≤ 0.6%).
Beilei Jiang, Xianwei Cheng, Yuan Li 0054, Jocelyn Zhang, Song Fu, Qing Yang 0003, Mingxiong Liu, Alejandro Olvera
ICPP5
2023 User-Defined Privacy Preserving Data Sharing for Connected Autonomous Vehicles Utilizing Edge Computing
abstract
In this paper, we present PRECISE, a novel privacy preserving data sharing framework for connected autonomous vehicles (CAVs). PRECISE allows users to define the objects or parts that they wish to protect privacy before sharing data with other vehicles. It leverages secure segmentation and inpainting technologies to protect sensitive data of vehicles. PRECISE explores the edges to offload resource-intensive deep learning workloads. To ensure data privacy in the processing on edge, PRECISE leverages additive secret sharing theory to define secure functions for deep neural networks (DNNs). Two secure DNN models, Secure SegNet and Secure Context Encoder, are introduced, along with detailed explanations of how to develop secure CNN layers and the secure functions used in building these layers. We have implemented a prototype of PRECISE and evaluated its performance. The experimental results demonstrate that PRECISE is lightweight, achieving secure segmentation in 3.47 seconds and secure inpainting in 0.99 seconds. The inference outputs from PRECISE remain the same as those from the original DNNs, while data privacy is protected. To the best of our knowledge, PRECISE is the first of its kind to provide user-defined privacy protection for sensor data sharing among CAVs.
Tianyu Bai, Qing Yang 0003, Song Fu
SEC3
2023 A novel method for aeroengine performance model reconstruction based on CDAE model
Lin Lin 0014, Wenhui He, Song Fu, Changsheng Tong, Lizheng Zu
Adv. Eng. Informatics4
2023 Novel aeroengine fault diagnosis method based on feature amplification
Lin Lin 0014, Wenhui He, Song Fu, Changsheng Tong, Lizheng Zu
Eng. Appl. Artif. Intell.3
2023 Using combinatorial optimization to solve entity alignment: An efficient unsupervised model
Lin Lin 0014, Lizheng Zu, Song Fu, Yancheng Lv
Neurocomputing4
2022 Distributed Data-Sharing Consensus in Cooperative Perception of Autonomous Vehicles
abstract
To enable self-driving without a human driver, an autonomous vehicle needs to perceive its surrounding obstacles using onboard sensors, of which the perception accuracy might be limited by their own sensing range. An effective way to improve vehicles’ perception accuracy is to let nearby vehicles exchange their sensor data so that vehicles can detect obstacles beyond their own sensing ranges, called cooperative perception. The shared sensor data, however, might disclose the sensitive information of vehicles’ passengers, raising privacy and safety concerns (e.g. stalking or sensitive location leakage).In this paper, we propose a new data-sharing policy for the cooperative perception of autonomous vehicles, of which the objective is to minimize vehicles’ information disclosure without compromising their perception accuracy. Considering vehicles usually have different desires for data-sharing under different traffic environments, our policy provides vehicles autonomy to determine what types of sensor data to share based on their own needs. Moreover, given the dynamics of vehicles’ data-sharing decisions, the policy can be adjusted to incentivize vehicles’ decisions to converge to the desired decision field, such that a healthy cooperation environment can be maintained in a long term. To achieve such objectives, we analyze the dynamics of vehicles’ data-sharing decisions by resorting to the game theory model, and optimize the data-sharing ratio in the policy based on the analytic results. Finally, we carry out an extensive trace-driven simulation to test the performance of the proposed data-sharing policy. The experimental results demonstrate that our policy can help incentivize vehicles’ data-sharing decisions to the desired decision fields efficiently and effectively.
Chenxi Qiu, Anna Cinzia Squicciarini, Qing Yang 0003, Song Fu, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001
ICDCS5
2022 MLCNN: Cross-Layer Cooperative Optimization and Accelerator Architecture for Speeding Up Deep Learning Applications
abstract
The ever-increasing number of layers, millions of parameters, and large data volume make deep learning workloads resource-intensive and power-hungry. In this paper, we develop a convolutional neural network (CNN) acceleration framework, named MLCNN, which explores algorithm-hardware co-design to achieve cross-layer cooperative optimization and acceleration. MLCNN dramatically reduces computation and on-off chip communication, improving CNN's performance. To achieve this, MLCNN reorders the position of nonlinear activation layers and pooling layers, which we prove results in a negligible accuracy loss; then the convolutional layer and pooling layer are co-optimized by means of redundant multiplication elimination, local addition reuse, and global addition reuse. To the best of our knowledge, MLCNN is the first of its kind that incorporates cooperative optimization across convolutional, activation, and pooling layers. We further customize the MLCNN accelerator to take full advantage of cross-layer CNN optimization to reduce both computation and on-off chip communication. Our analysis shows that MLCNN can significantly reduce (up to 98%) multiplications and additions. We have implemented a prototype of MLCNN and evaluated its performance on several widely used CNN models using both an accelerator-level cycle and energy model and RTL implementation. Experimental results show that MLCNN achieves 3.2x speedup and 2.9x energy efficiency compared with dense CNNs. MLCNN's optimization methods are orthogonal to other CNN acceleration techniques, such as quantization and pruning. Combined with quantization, our quantized MLCNN gains a 12.8x speedup and 11.3x energy efficiency compared with DCNN.
Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Song Fu, Qing Yang 0003, Mingxiong Liu
IPDPS6
2022 Slim-FCP: Lightweight-Feature-Based Cooperative Perception for Connected Automated Vehicles
abstract
Cooperative perception provides a novel way to conquer the sensing limitation on a single automated vehicle and potentially improves driving safety. To reduce the transmission data volume, existing solutions use the intermediate data generated by convolutional neural network (CNN) models, namely, feature maps, to achieve cooperative perception. The feature maps are however too large to be transmitted by the current V2X technology. We propose a novel approach, called Slim-FCP, to significantly reduce the transmission data size. It enables a channelwise feature encoder to remove irrelevant features for a better compression ratio. In addition, it adopts an intelligent channel selection strategy through which only representative channels of feature maps are selected for transmission. To evaluate the effectiveness of Slim-FCP, we further define a recall-to-bandwidth (RB) ratio metric to quantitatively measure how the recall of object detection changes with respect to the available network bandwidth. Experiment results show that Slim-FCP reduces the transmission data size by 75%, compared with the best state-of-the-art solution, with a subtle loss on object detection’s recall.
Jingda Guo, Dominic Carrillo, Qi Chen 0018, Qing Yang 0003, Song Fu, Hongsheng Lu
IEEE Internet Things J.5
2022 A storage computing architecture with multiple NDP devices for accelerating compaction performance in LSM-tree based KV stores
Hui Sun 0002, Yinliang Yue, Song Fu
J. Syst. Archit.5
2022 A Novel Time-Series Memory Auto-Encoder With Sequentially Updated Reconstructions for Remaining Useful Life Prediction
abstract
One of the significant tasks in remaining useful life (RUL) prediction is to find a good health indicator (HI) that can effectively represent the degradation process of a system. However, it is difficult for traditional data-driven methods to construct accurate HIs due to their incomprehensive consideration of temporal dependencies within the monitoring data, especially for aeroengines working under nonstationary operating conditions (OCs). Aiming at this problem, this article develops a novel unsupervised deep neural network, the so-called times series memory auto-encoder with sequentially updated reconstructions (SUR-TSMAE) to improve the accuracy of extracted HIs, which directly takes the multidimensional time series as input to simultaneously achieve feature extraction from both feature-dimension and time-dimension. Further, to make full use of the temporal dependencies, a novel long-short time memory with sequentially updated reconstructions (SUR-LSTM), which uses the errors not only from the current memory cell but also from subsequent memory cells to update the output layer's weight of the current memory cell, is developed to act as the reconstructed layer in the SUR-TSMAE. The use of SUR-LSTM can help the SUR-TSMAE rapidly reconstruct the input time series with higher precision. Experimental results on a public dataset demonstrate the outstanding performance of SUR-TSMAE in comparison with some existing methods.
Song Fu, Lin Lin 0014, Minghang Zhao
IEEE Trans. Neural Networks Learn. Syst.1
2021 APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGA
abstract
In this paper, we introduce APCNN, which explores algorithm-hardware co-design and provides a CNN acceleration framework with multi-layer cooperative optimization and customized design on FPGA. In terms of the algorithm design, the pooling layer is moved before the non-linear activation function and normalization in APCNN, which we prove causes negligible accuracy loss; the pooling layer is then co-optimized with the convolutional layer by means of redundant multiplication elimination, local addition reuse, and global addition reuse. We further design a dedicated accelerator to take full advantage of convolutional-pooling cross-layer optimization to not only accelerate computation but also reduce on-off chip data communication on FPGA. We demonstrate that our novel APCNN can achieve 75% multiplication and 75% addition reduction in the best case. For on-off chip data communication, a max{Row,Col} /(Row x Col) percent of memory footprint can be eliminated, where Row and Col are the number of rows and columns in the activation feature map respectively. We have implemented a prototype of APCNN and evaluated its performance on LeNet-5 and VGG16 using both an accelerator-level cycle and energy model and an RTL implementation. Our experimental results show that APCNN achieves a 2.5× speedup and 4.7× energy efficiency compared with the dense CNN. (This research was supported in part by NSF grants CCF-1563750, OAC-2017564, and CNS-2037982.)
Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Hui Zhao 0013, Song Fu
FPGA7
2021 Learning Connected Attentions for Convolutional Neural Networks
abstract
While self-attention mechanism has shown promising results for many vision tasks, it only considers the current features at a time. We show that such a manner cannot take full advantage of the attention mechanism. In this paper, we present Deep Connected Attention Network (DCANet), a novel design that boosts attention modules in a CNN model without any modification of the internal structure. To achieve this, we interconnect adjacent attention blocks, making information flow among attention blocks possible. With DCANet, all attention blocks in a CNN model are trained jointly, which improves the ability of attention learning. Our DCANet is generic. It is not limited to a specific attention module or base network architecture. Experimental results on ImageNet and MS COCO benchmarks show that DCANet consistently outperforms the state-of-the-art attention modules with a minimal additional computational overhead in all test cases. The code is available at: https://github.com/13952522076/DCANet.
Xu Ma 0005, Jingda Guo, Sihai Tang, Zhinan Qiao, Qi Chen 0018, Qing Yang 0003, Song Fu, Paparao Palacharla, Nannan Wang 0003, Xi Wang 0001
ICME7
2021 CoConv: Learning Dynamic Cooperative Convolution for Image Recognition
abstract
In this paper, we present a conceptually simple, yet powerful method for image recognition. The method, called Cooperative Dynamic Convolution (CoConv), introduces a cooperative learning of dynamic convolution from multiple convolutional experts. CoConv can be used as a substitute for the traditional static convolution, and can be seamlessly integrated in various visual models. Moreover, CoConv is easy to train with only a minimal computational overhead introduced in the inference phase. CoConv is trained by using multiple convolutional experts simultaneously, and the convolutional weights are merged by a weighted summation before convolutional operations for efficiency during inference. Results from extensive experiments show that CoConv leads to consistent improvement for image classification on various datasets, independent of the choice of the base convolutional network. Remarkably, CoConv improves the top-1 classification accuracy of ResNet18 by 3.06% on ImageNet. The code is available at: https://github.com/Nyquixt/CoConv.
Kien X. Nguyen 0002, Tiffany Ryu, Jocelyn Zhang, Xu Ma 0005, Qing Yang 0003, Song Fu, Paparao Palacharla, Nannan Wang 0003, Xi Wang 0001
ICME6
2021 A re-optimized deep auto-encoder for gas turbine unsupervised anomaly detection
Song Fu, Lin Lin 0014, Minghang Zhao
Eng. Appl. Artif. Intell.1
2021 CoFF: Cooperative Spatial Feature Fusion for 3-D Object Detection on Autonomous Vehicles
abstract
To reduce the amount of transmitted data, feature map-based fusion is recently proposed as a practical solution to cooperative 3-D object detection by autonomous vehicles (AVs). The precision of object detection, however, may require significant improvement, especially for objects that are far away or occluded. To address this critical issue for the safety of AVs and human beings, we propose a cooperative spatial feature fusion (CoFF) method for AVs to effectively fuse feature maps for achieving a higher 3-D object detection performance. Especially, CoFF differentiates weights among feature maps for a more guided fusion, based on how much new semantic information is provided by the received feature maps. It also enhances the inconspicuous features corresponding to far/occluded objects to improve their detection precision. The experimental results show that CoFF achieves a significant improvement in terms of both detection precision and effective detection range for AVs, compared to previous feature fusion solutions.
Jingda Guo, Dominic Carrillo, Sihai Tang, Qi Chen 0018, Qing Yang 0003, Song Fu, Xi Wang 0001, Nannan Wang 0003, Paparao Palacharla
IEEE Internet Things J.6
2021 Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
J. Parallel Distributed Comput.6
2021 Spatial Pyramid Attention for Deep Convolutional Neural Networks
abstract
Attention mechanisms have shown great success in computer vision. However, the commonly used global average pooling in some implementations aggregates a three-dimensional feature map to a one-dimensional attention map, leading a significant loss of structural information in the attention learning. In this article, we present a novel Spatial Pyramid Attention Network (SPANet), which exploits the structural information and channel relationships for better feature representation. SPANet enhances a base network by adding Spatial Pyramid Attention (SPA) blocks laterally. By rethinking the self-attention mechanism design, we further present three topology structures of attention path connection for our SPANet. They can be flexibly applied to various CNN architectures. SPANet is conceptually simple but practically powerful. It uses both structural regularization and structural information to achieve better learning capability. We have comprehensively evaluated the performance of SPANet on four benchmark datasets for different visual tasks. The experimental results show that SPANet significantly improves the recognition accuracy without adding much computation overhead. Using SPANet, we achieve an improvement of 1.6% top-1 classification accuracy on the ImageNet 2012 benchmark based on ResNet50, and SPANet outperforms SENet and other attention methods. SPANet also significantly improves the object detection performance by a clear margin with negligible additional computation overhead. When applying SPANet to RetinaNet based on the ResNet50 backbone, we improve the performance of the baseline model by 2.3 mAP and the enhanced model outperforms SENet and GCNet by 1.1 mAP and 1.7 mAP respectively. The code of SPANet is made publicly available.11[Online]. Available:https://github.com/13952522076/SPANet_TMM
Xu Ma 0005, Jingda Guo, Andrew Sansom, Mara McGuire, Andrew Kalaani, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu
IEEE Trans. Multim.9
2020 Gait Recognition for Co-Existing Multiple People Using Millimeter Wave Sensing
abstract
Gait recognition, i.e., recognizing persons from their walking postures, has found versatile applications in security check, health monitoring, and novel human-computer interaction. The millimeter-wave (mmWave) based gait recognition represents the most recent advance. Compared with traditional camera-based solutions, mmWave based gait recognition bears unique advantages of being still effective under non-line-of-sight scenarios, such as in black, weak light, or blockage conditions. Moreover, they are able to accomplish person identification while preserving privacy. Currently, there are only few works in mmWave gait recognition, since no public data set is available. In this paper, we build a first-of-its-kind mmWave gait data set, in which we collect gait of 95 volunteers 'seen' from two mmWave radars in two different scenarios, which together lasts about 30 hours. Using the data set, we propose a novel deep-learning driven mmWave gait recognition method called mmGaitNet, and compare it with five state-of-the-art algorithms. We find that mmGaitNet is able to achieve 90% accuracy for single-person scenarios, 88% accuracy for five co-existing persons, while the existing methods achieve less than 66% accuracy for both scenarios.
Song Fu, Hongyuan Liang, Anfu Zhou, Shilin Zhu, Huadong Ma, Jianhua Liu 0004, Ning Yang 0010
AAAI2
2020 A middle-ware approach to leverage the distributed data de-duplication capability on HPC and Cloud storage systems
abstract
The unprecedented growth in the volume and diversity of the data in today's HPC and Enterprise computing environment has posted challenging problems on data management and data space reduction. More than 71% of enterprise and HPC communities are seeking de-duplication technologies to reduce cost and increase the storage efficiency. The importance of applying data de-duplication techniques is critical for active research and development. Current implementations of data de-duplication systems are mainly hardware dependent, system dependent, and platform dependent. Also, most of these implementations are proprietary software and not in open source domains. In this paper, we present a new middle-ware design and implementation approach, named D3M, to support distributed data de-duplication feature on existing file and object storage systems. We also incorporate this proposed D3M middle-ware with the Redhat's Linux device layer de-duplication and compression driver, called VDO (Virtual Data Optimizer). With these two layers of data de-duplication support, we accommodate both client side and server-side data de-duplication features. Finally, we conduct various testing cases on HPC data sets and Enterprise data sets to illustrate the benefits and advantages of applying our bilayer data de-duplication middle-ware solution.
Hsing-bung Chen, Sihai Tang, Song Fu
IEEE BigData3
2020 Cascaded Context Dependency: An Extremely Lightweight Module For Deep Convolutional Neural Networks
abstract
In this paper, we present a cascaded context dependency module, which is a highly lightweight module that can improve the performance of deep convolutional neural networks for various visual tasks. Inspired by the feature pyramid work in object detection and the context dependency work in image recognition, we consider to cascade the contexts of multiscale feature maps to aggregate the locality and globality in a local region. We further extract the dependency between original input and cascaded contexts for feature recalibration. Without employing learnable layers, our method introduces almost no additional parameters and computations. Furthermore, Our module can be seamlessly plugged into many existing CNN architectures to improve the performance. Experiments on ImageNet and MS COCO benchmarks indicate that our method can achieve results on par with or better than related work. Qualitatively, we achieve an absolute 1.42% (77.3137% vs. 75.8974%) top-1 classification accuracy improvement based on ResNet50 on ImageNet 2012 validation set with negligible computational overhead. Besides, our method yields significant gains on the MS COCO benchmark for the object detection task. All codes and models are made publicly available1.1We submit all the codes, pre-trained models, and training log files to https://github.com/13952522076/ParameterFree.
Xu Ma 0005, Zhinan Qiao, Jingda Guo, Sihai Tang, Qi Chen 0018, Qing Yang 0003, Song Fu
ICIP7
2020 Spanet: Spatial Pyramid Attention Network for Enhanced Image Recognition
abstract
Attention mechanism has shown great success in computer vision. In this paper, we introduce Spatial Pyramid Attention Network (SPANet) to investigate the role of attention block for image recognition. Our SPANet is conceptually simple but practically powerful. It enhances the base network by adding Spatial Pyramid Attention (SPA) Blocks laterally. In contrast to other attention based networks that leverage global average pooling, our proposed SPANet considers both structural regularization and structural information. Furthermore, we investigate the topology structure of attention path connection and present three SPANet structures. SPA block is flexible to be deployed to various convolutional neural network (CNN) architectures. The experimental results show that our SPANet significantly improves the recognition accuracy without introducing much computation overhead compared with other CNN models. Codes are made publicly available11https://github.com/13952522076/SPANet.
Jingda Guo, Xu Ma 0005, Andrew Sansom, Mara McGuire, Andrew Kalaani, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu
ICME9
2020 Attention Meets Normalization and Beyond
abstract
To make Convolutional Neural Networks (CNNs) more efficient and accurate, various lightweight self-attention modules have been proposed. In this paper, we systematically study state-of-the-art attention modules in CNNs and discover that self-attention mechanism can be closely related to normalization. Based on this observation, we propose a novel attention module, named Normalization-Attention module (NA module in short), which is almost parameter-free. The NA module calculates the mean and standard deviation of intermediate feature maps and processes the feature context with normalization, which makes a CNN model easier to be trained and more responsive to informative features. Our proposed Normalization-Attention module can be integrated into various base CNN architectures, and used for many computer vision tasks, including image recognition, object detection, and more. Experimental results on ImageNet and MS COCO benchmarks show that our method outperforms state-of-the-art works using fewer parameters. Codes are made publicly available.
Xu Ma 0005, Jingda Guo, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu
ICME6
2020 Cooperative Mixed Reality Leveraging Edge Computing and Communication
abstract
Traditional Mixed Reality(MR) and Augmented Reality(AR) devices support a wide gamut of sensors, but the limited computational resource onboard such devices make advanced tasks difficult. With the introduction of Edge computing and more advanced Edge hardware, we are no longer bound to just the onboard processor or the Cloud as our only source of computational power. In our work, we introduce the use of Edge with MR devices to provide a cooperative perception capability to the MR device. We base our approach on the portability and low latency of the Edge. Through our prototype system, we demonstrate the potential of devices and evaluate the performance and feasibility through real world trials. Our evaluation proves that the system is capable of supporting cooperative perception tasks.
Sihai Tang, Bruce Haidi Chen, Jacob Hochstetler, Jason Hirsch, Song Fu
SEC5
2020 Position-Aware Recalibration Module: Learning From Feature Semantics and Feature Position
abstract
We present a new method to improve the representational power of the features in Convolutional Neural Networks (CNNs). By studying traditional image processing methods and recent CNN architectures, we propose to use positional information in CNNs for effective exploration of feature dependencies. Rather than considering feature semantics alone, we incorporate spatial positions as an augmentation for feature semantics in our design. From this vantage, we present a Position-Aware Recalibration Module (PRM in short) which recalibrates features leveraging both feature semantics and position. Furthermore, inspired by multi-head attention, our module is capable of performing multiple recalibrations where results are concatenated as the output. As PRM is efficient and easy to implement, it can be seamlessly integrated into various base networks and applied to many position-aware visual tasks. Compared to original CNNs, our PRM introduces a negligible number of parameters and FLOPs, while yielding better performance. Experimental results on ImageNet and MS COCO benchmarks show that our approach surpasses related methods by a clear margin with less computational overhead. For example, we improve the ResNet50 by absolute 1.75% (77.65% vs. 75.90%) on ImageNet 2012 validation dataset, and 1.5%~1.9% mAP on MS COCO validation dataset with almost no computational overhead. Codes are made publicly available.
Xu Ma 0005, Song Fu
IJCAI2
2019 Applying SDN based data network on HPC Big Data Computing - Design, Implementation, and Evaluation
abstract
Large scale storage data networks are difficult to conFigure and tricky to maintain [1][2][3]. As storage data volumes grow and the pace of change accelerates, it can be a struggle to keep up the integrity of a large scale storage network. Software Defined Networking (SDN) [5][6][7][8][10][11] provides a method to centrally conFigure and manage physical and virtual network devices such as routers, switches, and gateways in HPC datacenter. Current HPC computer cluster and data centers are using homogeneous data network technologies such as the Infiniband network (QDR, FDR, EDR, and HDR). Due to the connector, cable, and bandwidth backward compatibility issues, once an aged HPC computing system was retired, we have to demolish all established computer clusters, data network, and data storage. Eventually we lost all costly investment on the data network and storage. Those current deployment approaches do not provide a feasible and compatible growing path and cannot meet the need for future extreme scale HPC computing systems [4][12][13].
Hsing-bung Chen, Zhi Qiao 0001, Song Fu
IEEE BigData3
2019 Accelerating RNN on FPGA with Efficient Conversion of High-Level Designs to RTL
abstract
Recurrent Neural Network (RNN) is a powerful Deep Learning algorithm which has been widely used for speech recognition, handwriting recognition, context clustering, etc. Deep learning involves a large number of floating-point computations and requires a large amount of computing resource. Currently, software-based RNN implementations running on general-purpose processors like CPU and GPU take a long execution time and an excessive amount of energy. FPGA provides a low-power and highly parallel platform for accelerating RNN training and inference. With massive reconfigurable logics, FPGA can be customized to achieve dramatic speedup and energy efficiency for various deep learning applications. However, implementing RNN algorithms at the Register Transistor Level (RTL) is time-consuming, while existing tools for converting RNN code written in high-level languages to RTL designs are not efficient. In this paper, we present a design flow by which high-level RNN implementations (for example, in Python) can be converted to RTL designs efficiently and automatically. Experimental results show the generated RTL code that is deployed on an FPGA device is 7.87 times faster than the Python code run on CPU while achieving the same accuracy.
Zongze Li 0001, Song Fu
IEEE BigData2
2019 An Empirical Study of Quad-Level Cell (QLC) NAND Flash SSDs for Big Data Applications
abstract
As the SSD technology develops, quad-level cell (QLC) NAND based SSD is gradually being introduced to the market. And as such, we evaluate the QLC technologys impact on the landscape of modern datacenters. Since a large number of applications and workloads in the modern datacenters have far more read requests than they write requests, QLC SSD provides a promising solution. For example, real-time analytics and big data, machine and deep learning, and read-intensive AI applications are all read hungry perfectly suited for the QLC. Its favorable performance (especially in reads), high capacity and density greatly help modern datacenters to provide more efficient services to their customers. At the same time, the low cost of QLC SSD also helps to lower the cost of operation for datacenters. Additionally, we explore the state-of-art QLC SSD from the system architecture point of view to shows its key advancements from previous technologies. By conducting a comprehensive performance evaluation of QLC SSD, we are able to compare it with other types of SSD and analyze factors that impact its performance.
Shuwen Liang, Zhi Qiao 0001, Sihai Tang, Jacob Hochstetler, Song Fu, Weisong Shi, Hsing-bung Chen
IEEE BigData5
2019 Smart Home IoT Anomaly Detection based on Ensemble Model Learning From Heterogeneous Data
abstract
Nowadays, internet based home automation is made possible with the advent of intelligent device control. These electronic sensing devices transfer an enormous amount of data into the cloud. It is a challenge to discover hidden information from the massive amount of stored data in the cloud. In addition, privacy, security, and stability could also be a concern for users. Due to these issues becoming ever more prevalent in today’s society, the need to have access to readily anomaly detection becomes crucial for the modern smart home user. In this paper, we design, test and evaluate an ensemble model anomaly detection method. Our method targets the data anomalies present in general smart Internet of Things (IoT) devices, allowing for easy detection of anomalous events based on stored data. We make our method robust through ensemble machine learning model training. We aim to simulate different types of anomaly situations on publicly available smart home data sets, thereby exposing our models to likely real world phenomenons and events that may cause anomalies. Experiments are conducted on the processed data and evaluated for accuracy through validation and testing against independent and identically distributed labeled data.
Sihai Tang, Zhaochen Gu, Qing Yang 0003, Song Fu
IEEE BigData4
2019 Building Reliable High-Performance Storage Systems: An Empirical and Analytical Study
abstract
Due to the vast storage needs of high performance computing (HPC), the scale and complexity of storage systems in HPC data centers continue growing. Disk failures have become the norm. With the ever-increasing disk capacity, RAID recovery based on disk rebuild becomes more and more expensive, which causes significant performance degradation and even unavailability of storage systems. Declustered redundant array of independent disks shuffle data and parity blocks among all drives in a RAID group, which aims to accelerate RAID reconstruction and improve performance. With the popularity of ZFS file system and software RAID used in production systems, in this paper, we extensively evaluate and analyze declustered RAID with regard to the RAID I/O performance and recovery time on an high performance storage platform at Los Alamos National Laboratory. Our empirical study reveals that the speedup of declustered RAID over traditional RAID is sub-linear to the parallelism of recovery I/O. Furthermore, we formally model and analyze the reliability of declustered RAID using the mean-time-to-data-loss and discover that the improved recovery performance leads to a higher storage reliability compared with the traditional RAID.
Zhi Qiao 0001, Song Fu, Hsing-bung Chen, Bradley W. Settlemyer
CLUSTER2
2019 Cooper: Cooperative Perception for Connected Autonomous Vehicles Based on 3D Point Clouds
abstract
Autonomous vehicles may make wrong decisions due to inaccurate detection and recognition. Therefore, an intelligent vehicle can combine its own data with that of other vehicles to enhance perceptive ability, and thus improve detection accuracy and driving safety. However, multi-vehicle cooperative perception requires the integration of real world scenes and the traffic of raw sensor data exchange far exceeds the bandwidth of existing vehicular networks. To the best our knowledge, we are the first to conduct a study on raw-data level cooperative perception for enhancing the detection ability of self-driving systems. In this work, relying on LiDAR 3D point clouds, we fuse the sensor data collected from different positions and angles of connected vehicles. A point cloud based 3D object detection method is proposed to work on a diversity of aligned point clouds. Experimental results on KITTI and our collected dataset show that the proposed system outperforms perception by extending sensing area, improving detection accuracy and promoting augmented results. Most importantly, we demonstrate it is possible to transmit point clouds data for cooperative perception via existing vehicular network technologies.
Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu
ICDCS4
2019 Near-Data Processing-Enabled and Time-Aware Compaction Optimization for LSM-tree-based Key-Value Stores
abstract
With the growing volume of storage systems, the traditional relational databases cannot reach the high performance required by big-data applications. As high-throughput alternatives to relational databases, LSM-tree-based key-value stores (KV stores in short) are confronted with degraded write performance during compaction under update-intensive workloads. To address this issue, we design and implement a time-aware compaction optimization framework for KV stores called TStore. TStore explores the near-data processing (i.e., NDP) model. It dynamically partitions compaction tasks into both host and NDP-enabled device to minimize the total time of compaction. The partitioned compaction tasks are conducted by the host and the device in parallel. The NDP-based devices exhibit low-latency, high-performance and high-bandwidth capability, thus facilitating key-value stores. TStore can not only accomplish compaction for KV stores, but also improve overall performance by removing bottleneck in compaction. Results show that the TStore with an NDP framework can achieve 3.8x and 1.9x performance improvement over LevelDB and Co-KV under the db_bench workload. In addition, the TStore-enabled KV store outperforms LevelDB and Co-KV by a factor of 3.6x and 1.9x in throughput and 72.0% and 48.9% in latency, respectively, under realistic workloads generated by YCSB.
Hui Sun 0002, Jianzhong Huang 0001, Song Fu, Zhi Qiao 0001, Weisong Shi
ICPP4
2019 Exploring Declustered Software RAID for Enhanced Reliability and Recovery Performance in HPC Storage Systems
abstract
Redundant array of independent disks (RAID) has been widely used to address the reliability and performance issues of storage systems. As the scale of modern storage systems continues growing, disk failure becomes the norm. With the ever-increasing disk capacity, RAID recovery based on disk rebuild becomes more and more costly, which causes significant performance degradation and even unavailability of storage systems. Declustered data layout enables parallel RAID reconstruction by shuffling data and parity blocks among all drives (including spares) in a RAID group. However, the reliability and performance of declustered RAID in real-world storage environments have not been thoroughly studied. With the popularity of ZFS file system and software RAID used in production data centers, in this paper, we extensively evaluate declustered RAID with regard to the RAID recovery time and I/O performance on a high-performance storage platform at Los Alamos National Laboratory. Our empirical study reveals the advantages and disadvantages of declustered RAID technology. We qualitatively characterize the recovery performance of declustered RAID and compare with that of ZFS RAIDZ under various I/O workloads and access patterns. The experimental results show that the speedup of declustered RAID over traditional RAID is sub-linear to the parallelism of recovery I/O. Furthermore, we formally model and analyze the reliability of declustered RAID in terms of the mean-time-to-data-loss (MTTDL) and discover that the improved recovery performance leads to higher storage reliability compared with the traditional RAID.
Zhi Qiao 0001, Shuwen Liang, Hsing-bung Chen, Song Fu, Bradley W. Settlemyer
SRDS4
2018 Converting Unstructured System Logs into Structured Event List for Anomaly Detection
abstract
System logs provide invaluable resources for understanding system behavior and detecting anomalies on high performance computing (HPC) systems. As HPC systems continue to grow in both scale and complexity, the sheer volume of system logs and the complex interaction among system components make the traditional manual problem diagnosis and even automated line-by-line log analysis infeasible or ineffective. In this paper, we present a System Log Event Block Detection (SLEBD) framework that identifies groups of log messages that follow certain sequence but with variations, and explore these event blocks for event-based system behavior analysis and anomaly detection. Compared with the existing approaches that analyze system logs line by line, SLEBD is capable of characterizing system behavior and identifying intricate anomalies at a higher (i.e., event) level. We evaluate the performance of SLEBD by using syslogs collected from production supercomputers. Experimental results show that our framework and mechanisms can process streaming log messages, efficiently extract event blocks and effectively detect anomalies, which enables system administrators and monitoring tools to understand and process system events in real time. Additionally, we use the identified event blocks and explore deep learning algorithms to model and classify event sequences.
Zongze Li 0001, Matthew Davidson, Song Fu, Sean Blanchard, Michael Lang 0003
ARES3
2018 Reliability Characterization of Solid State Drives in a Scalable Production Datacenter
abstract
In recent years, NAND flash-based solid state drives (SSD) have been widely used in datacenters due to their better performance compared with the traditional hard disk drives. However, little is known about the reliability characteristics of SSDs in production systems. Existing works study the statistical distributions of SSD failures in the field. However, they do not go deep into SSD drives and investigate the unique error types and health dynamics that distinguish SSDs from hard disk drives. In this paper, we explore the SSD-specific SMART (Self-Monitoring, Analysis, and Reporting Technology) attributes to conduct an in-depth analysis of SSD reliability in a production environment. Data is collected from a scalable production system having several physical locations. Our dataset contains over a million records with more than twenty attributes. We leverage machine learning technologies, specifically data clustering and correlation analysis methods, to discover groups of SSDs which have different health status and relations among SSD-specific SMART attributes. Our results show that 1) Media wear affects the reliability of SSDs more than any other factors, and 2) SSDs transit from one health group to another which infers the reliability degradation of those drives. To the best of our knowledge, this is the first study that investigates SSD-specific SMART data to characterize SSD reliability in a production environment.
Shuwen Liang, Zhi Qiao 0001, Jacob Hochstetler, Song Fu, Weisong Shi, Devesh Tiwari, Hsing-bung Chen, Bradley W. Settlemyer, David Richard Montoya
IEEE BigData5
2018 ACTOR: Active Cloud Storage with Energy-Efficient On-Drive Data Processing
abstract
Storage systems are indispensable for big data processing and cloud computing services today. The ever-growing size of computation and data analytic results demands larger storage capacity, which challenges data processing and storage scalability. Moreover, the increasing complexity of storage hierarchy and "passive" storage devices make todays storage systems inefficient, which necessitates the adoption of new storage technologies. In this paper, we explore new Ethernet connected drives with on-drive embedded CPU and DRAM to develop an active cloud storage system where data can be processed on disk drives without data movement. These drives are micro-storage servers that can support software-defined storage. In addition to I/O operations, we test and evaluate on-drive data processing, including data compression, aggregation and erasure encoding, which provide natural support for data-intensive applications. Our experimental results show that Open Ethernet Drive can significantly lower the energy consumption while maintaining the data processing throughput simultaneously by ensuring data availability and storage scalability. Results and findings from this work will facilitate scheduling of on-drive compute resource for building active and scalable cloud storage systems.
Zhi Qiao 0001, Shuwen Liang, Nandini Damera, Song Fu, Hsing-bung Chen, Michael Lang 0003
IEEE BigData4
2018 Understanding and Analyzing Interconnect Errors and Network Congestion on a Large Scale HPC System
abstract
Today's High Performance Computing (HPC) systems are capable of delivering performance in the order of petaflops due to the fast computing devices, network interconnect, and back-end storage systems. In particular, interconnect resilience and congestion resolution methods have a major impact on the overall interconnect and application performance. This is especially true for scientific applications running multiple processes on different compute nodes as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks state-of-practice experience reports that detail how different interconnect errors and congestion events occur on large-scale HPC systems. Therefore, in this paper, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors and congestion events. We also study the interaction between interconnect, errors, network congestion and application characteristics.
Saurabh Gupta 0002, Tirthak Patel, Michael Wilder, Weisong Shi, Song Fu, Christian Engelmann, Devesh Tiwari
DSN6
2016 Improving Coding Performance and Energy Efficiency of Erasure Coding Process for Storage Systems - A Parallel and Scalable Approach
abstract
Erasure code based object storage systems are becoming popular choices for archive storage systems due to cost-effective storage space saving schemes and higher fault-resilience capabilities. Both erasure code encoding and decoding procedures involve heavy array, matrix, and table-lookup compute intensive operations. With today's advanced CPU design technologies such as multi-core, many-core, and streaming SIMD instruction sets we can effectively and efficiently adapt the erasure code technology in cloud storage systems and apply it to handle very large-scale date sets. Current solutions of the erasure coding process are based on single process approach which is not capable of processing very large data sets efficient and effectively. To prevent the bottleneck of a single process erasure encoding process, we utilize the task parallelism property from a multicore computing system and improve erasure coding process with parallel processing capability. We have leveraged open source erasure coding software and implemented a concurrent and parallel erasure coding software, called parEC. The proposed parEC process is realized through MPI run time parallel I/O environment and then data placement process is applied to distribute encoded data blocks to their destination storage devices. In this paper, we present the software architecture of parEC. We conduct various performance testing cases on parEC's software components. We present our early experience of using parEC, and address parEC's current status and future development works.
Hsing-bung Chen, Song Fu
CLOUD2
2016 Relational Synthesis of Text and Numeric Data for Anomaly Detection on Computing System Logs
abstract
Monitoring high performance computing systems has become increasingly difficult as researchers and system analysts face the challenge of synthesizing a wide range of monitoring information in order to detect system problems on ever larger machines. We present a method for anomaly detection on syslog data, one of the most important data streams for determining system health. Syslog messages pose a difficult question for analysis because they include a mix of structured natural language text as well as numeric values. We present an anomaly detection framework that combines graph analysis, relational learning, and kernel density estimation to detect unusual syslog messages. We design an event block detector, which finds groups of related syslog messages, to retrieve the entire section of syslog messages associated with a single anomalous line. Our novel approach successfully retrieves anomalous behaviors inserted into syslog files from a virtual machine, including messages indicating serious system problems. We also test our approach on syslog messages from the Trinity supercomputer and find that our methods do not generate significant false positives.
Elisabeth Baseman, Sean Blanchard, Zongze Li 0001, Song Fu
ICMLA4
2016 Quantifying entity criticality for fault impact analysis and dependability enhancement in software-defined networks
abstract
Software-defined networking (SDN) empowers network operators with more flexibility to program their networks. With SDN, network management moves from codifying functionality in terms of low-level device configurations to building virtualized software entities that facilitate network management and debugging. By separating the complexity of state distribution from network specification, SDN provides new ways to solve existing routing problems while allowing the use of dependability techniques. However, the dependability of SDN itself is still an open issue. In this paper, we study the criticality of various entities in a virtualized network, which is important for analyzing the impact of component failures. We propose a network entity criticality framework with a set of models to quantify the importance of different entities for network dependability. We incorporate the topology information of entities in the calculation of their criticality. We validate the proposed models and evaluate their performance on two virtualized networks used in production environments. Our experimental results show that the proposed criticality models can achieve a high accuracy for quantifying the criticality of virtualized entities. In addition, the graph entropy criticality model and the PageRank criticality model scale well and can work for large-scale virtualized networks.
Zhiang Deng, Song Fu
IPCCC3
2016 Parallel Erasure Coding: Exploring Task Parallelism in Erasure Coding for Enhanced Bandwidth and Energy Efficiency
abstract
Very large data sets within the range of megabytes to terabytes generated daily from checkpoint-and- restart processes are seen in today's scientific simulations. Reliability and durability are two important factors to build an archive storage system. Erasure code based object storage systems are becoming popular choices for archive storage systems due to cost-effective storage space saving schemes and higher fault-resilience capabilities. Both erasure code encoding and decoding procedures involve heavy array, matrix, and table-lookup compute intensive operations. Current solutions of the erasure coding process are based on single process approach which is not capable of processing very large data sets efficient and effectively. In this paper, we address the bottleneck problem of single process erasure encoding by leveraging task parallelism offered by multi-core computers. We add parallel processing capability to the erasure coding process. More specifically, we develop a parallel erasure coding software, called parEC. It explores the MPI run time parallel I/O environment and integrates data placement process for distributing encoded data blocks to destination storage devices. We evaluate the performance of parEC in terms of both encoding throughput and energy efficiency. We also compare the performance of two task scheduling algorithms for parEC. Our experimental results show parEC can significantly reduce the encoding time (i.e., by 74.06%-96.86%) and energy consumption (i.e., by 73.57%-96.86%), and Demand-based Workload Assignment (DBWA) algorithm can a high system utilization (i.e., 95.23%).
Hsing-bung Chen, Song Fu
NAS2
2016 TracSim: Simulating and scheduling trapped power capacity to maximize machine room throughput
Michael Lang 0003, Scott Pakin, Song Fu
Parallel Comput.4
2015 MR-Graph: A Customizable GPU MapReduce
abstract
The MapReduce programming model has been widely used in Big Data and Cloud applications. Criticism on its inflexibility when being applied to complicated scientific applications recently emerges. Several techniques have been proposed to enhance its flexibility. However, some of them exert special requirements on applications, while others fail to support the increasingly popular coprocessors, such as Graphics Processing Unit (GPU). In this paper, we propose MR-Graph, a customizable and unified framework for GPU-based MapReduce, which aims to improve the flexibility, scalability and performance of MapReduce. MR-Graph addresses the limitations and restrictions of the traditional MapReduce execution paradigm. The three execution modes integrated in MR-Graph facilitates users to write their applications in a more flexible fashion by defining a Map and Reduce function call graph. MR-Graph efficiently explores the memory hierarchy in GPUs to reduce the data transfer overhead between execution stages and accommodate big data applications. We have implemented a prototype of MR-Graph and experimental results show the effectiveness of using MR-Graph for flexible and scalable GPU-based MapReduce computing.
Zhi Qiao 0001, Shuwen Liang, Hai Jiang 0003, Song Fu
CSCloud4
2015 PASSI: A Parallel, Reliable and Scalable Storage Software Infrastructure for active storage system and I/O environments
abstract
Storage systems are a foundational component of computational, experimental, and observational science today. The ever-growing size of computation and simulation results demands huge storage capacity, which challenges the storage scalability and causes data corruption and disk failure to be commonplace in exascale storage environments. Moreover, the increasing complexity of storage hierarchy and passive storage devices make today's storage systems inefficient, which force the adoption of new storage technologies. In this position paper, we propose a Parallel, Reliable and Scalable Storage Software Infrastructure (PASSI) to support the design and prototyping of next-generation active storage environment. The goal is to meet the scaling and resilience need of extreme scale science by ensuring that storage systems are pervasively intelligent, always available, never lose or damage data and energy-efficient.
Hsing-bung Chen, Song Fu
IPCCC2
2015 A customizable MapReduce framework for complex data-intensive workflows on GPUs
abstract
The MapReduce programming model has been widely used in big data and cloud applications. Criticism on its inflexibility when being applied to complicated scientific applications recently emerges. Several techniques have been proposed to enhance its flexibility. However, some of them exert special requirements on applications, while others fail to support the increasingly popular coprocessors, such as Graphics Processing Unit (GPU). In this paper, we propose MR-Graph, a customizable and unified framework for GPU-based MapReduce, which aims to improve the flexibility and performance of MapReduce. MR-Graph addresses the limitations and restrictions of the traditional MapReduce execution paradigm. The three execution modes integrated in MR-Graph facilitates users to write their applications in a more flexible fashion by defining a Map and Reduce function call graph. MR-Graph efficiently explores the memory hierarchy in GPUs to reduce the data transfer overhead between execution stages and accommodate big data applications.We have implemented a prototype of MR-Graph and experimental results show the effectiveness of using MR-Graph for flexible and scalable GPU-based MapReduce computing.
Zhi Qiao 0001, Shuwen Liang, Hai Jiang 0003, Song Fu
IPCCC4
2015 Differentiated Failure Remediation with Action Selection for Resilient Computing
abstract
As the fault frequency is increasing with the component count in modern and future computer systems, resilience becomes increasingly critical. Existing work on anomaly detection and fault prediction enables failure avoidance techniques to circumvent fault effects proactively. In addition, traditional fault tolerance techniques can be applied to handle faults reactively. Different types of faults may affect different components of a system and have various manifestations. They need to be treated differently. However, the existing fault handling techniques uniformly treat all faults without considering their types and distinct properties. In this paper, we present a differentiated fault remediation framework with action selection (DFRAS) which integrates both preventive and reactive remediation actions differentiated for different types of faults with their urgency requirements. We investigate four major types of faults and identify candidate remediation actions. We apply the urgency requirements as constraints for action selection. We propose formal performance models to quantify the wasted time of the candidate actions, and develop a decision making method to select the best actions that minimize the overall remediation cost. We have implemented a prototype of DFRAS and evaluated its performance by simulations and experiments. Simulation and experimental results show that the integrated fault remediation strategies can significantly reduce the remediation overhead. The developed DFRAS system is lightweight, making it feasible for online fault management in large-scale systems.
Song Fu, Nathan DeBardeleben, Qiang Guan, Cheng-Zhong Xu 0001
PRDC2
2014 F-SEFI: A Fine-Grained Soft Error Fault Injection Tool for Profiling Application Vulnerability
abstract
As the high performance computing (HPC) community continues to push towards exascale computing, resilience remains a serious challenge. With the expected decrease of both feature size and operating voltage, we expect a significant increase in hardware soft errors. HPC applications of today are only affected by soft errors to a small degree but we expect that this will become a more serious issue as HPC systems grow. We propose F-SEFI, a Fine-grained Soft Error Fault Injector, as a tool for profiling software robustness against soft errors. In this paper we utilize soft error injection to mimic the impact of errors on logic circuit behavior. Leveraging the open source virtual machine hypervisor QEMU, F-SEFI enables users to modify emulated machine instructions to introduce soft errors. F-SEFI can control what application, which sub-function, when and how to inject soft errors with different granularities, without interference to other applications that share the same environment. F-SEFI does this without requiring revisions to the application source code, compilers or operating systems. We discuss the design constraints for F-SEFI and the specifics of our implementation. We demonstrate use cases of F-SEFI on several benchmark applications to show how data corruption can propagate to incorrect results.
Qiang Guan, Nathan DeBardeleben, Sean Blanchard, Song Fu
IPDPS4
2013 Wavelet-based multi-scale anomaly identification in cloud computing systems
abstract
Modern cloud computing systems contain thousands of computing and storage servers. Such a scale combined with ever-growing system complexity of their components and interactions, introduces a key challenge to failure and resource management for highly dependable cloud computing. Automated anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system level dependability assurance. In this paper, we present a wavelet-based multi-scale anomaly identification mechanism, that can analyze profiled cloud performance metrics in both time and frequency domains and identify anomalous cloud behaviors. Learning technologies are exploited to adapt the selection of mother wavelets and a sliding detection window is employed handle cloud dynamicity and improve anomaly detection accuracy. We test a prototype implementation of our cloud anomaly detection mechanism on an institute-wide cloud system. Experimental results show our approach can identify cloud failures accurately.
Qiang Guan, Song Fu
GLOBECOM2
2013 Exploring Time and Frequency Domains for Accurate and Automated Anomaly Detection in Cloud Computing Systems
abstract
Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is crucial for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud behaviors, we need to monitor the cloud execution and collect runtime cloud performance data. For different types of failures, the data display different correlations with the performance metrics. In this paper, we present a wavelet-based multi-scale anomaly identification mechanism, that can analyze profiled cloud performance metrics in both time and frequency domains and identify anomalous cloud behaviors. Learning technologies are exploited to adapt the selection of mother wavelets and a sliding detection window is employed to handle cloud dynamicity and improve anomaly detection accuracy. We have implemented a prototype of the anomaly identification system and conducted experiments on an on-campus cloud computing environment. Experimental results show the proposed mechanism can achieve 93.3% detection sensitivity while keeping the false positive rate as low as 6.1% while outperforming other tested anomaly detection schemes.
Qiang Guan, Song Fu, Nathan DeBardeleben, Sean Blanchard
PRDC2
2013 Adaptive Anomaly Identification by Exploring Metric Subspace in Cloud Computing Infrastructures
abstract
Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud behaviors, we need to monitor the cloud execution and collect runtime cloud performance data. These data consist of values of performance metrics for different types of failures, which display different correlations with the performance metrics. In this paper, we present an adaptive anomaly identification mechanism that explores the most relevant principal components of different failure types in cloud computing infrastructures. It integrates the cloud performance metric analysis with filtering techniques to achieve automated, efficient, and accurate anomaly identification. The proposed mechanism adapts itself by recursively learning from the newly verified detection results to refine future detections. We have implemented a prototype of the anomaly identification system and conducted experiments in an on-campus cloud computing environment and by using the Google data center traces. Our experimental results show that our mechanism can achieve more efficient and accurate anomaly detection than other existing schemes.
Qiang Guan, Song Fu
SRDS2
2012 A Hybrid Anomaly Detection Framework in Cloud Computing Using One-Class and Two-Class Support Vector Machines
Song Fu, Jianguo Liu 0001, Husanbir Singh Pannu
ADMA1
2012 A self-evolving anomaly detection framework for developing highly dependable utility clouds
abstract
Utility clouds continue to grow in scale and in the complexity of their components and interactions, which introduces a key challenge to failure and resource management for highly dependable cloud computing. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To identify anomalies, we need to monitor the system execution and collect health-related runtime performance data. These data are usually unlabeled and a prior failure history is not always available in production systems, especially for newly deployed or managed utility clouds. In this paper, we present a self-evolving anomaly detection framework with mechanisms for dependability assurance in utility clouds. No prior failure history is required. The detector self-evolves by recursively exploring newly generated verified detection results for future anomaly identification. Statistical learning technologies are exploited in detector determination and working dataset selection. Experimental results in an institute-wide cloud computing system show that the detection accuracy improves as it evolves. With self-evolvement, the detector can achieve 92.1% detection sensitivity and 83.8% detection specificity, which makes it well suitable for building highly dependable utility clouds.
Husanbir Singh Pannu, Jianguo Liu 0001, Song Fu
GLOBECOM3
2012 AFD: Adaptive failure detection system for cloud computing infrastructures
abstract
Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Autonomic failure detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled, and thus a prior failure history is not always available in production clouds, especially for newly managed or deployed systems. In this paper, we present an Adaptive Failure Detection (AFD) framework for cloud dependability assurance. AFD employs data description using hypersphere for adaptive failure detection. Based on the cloud performance data, AFD detects possible failures, which are verified by the cloud operators. They are confirmed as either true failures with failure types or normal states. AFD adapts itself by recursively learning from these newly verified detection results to refine future detections. Meanwhile, AFD exploits the observed but undetected failure records reported by the cloud operators to identify new types of failures. We have implemented a prototype of the AFD system and conducted experiments in an on-campus cloud computing environment. Our experimental results show that AFD can achieve more efficient and accurate failure detection than other existing schemes.
Husanbir Singh Pannu, Jianguo Liu 0001, Qiang Guan, Song Fu
IPCCC4
2012 An adaptive power management framework for autonomic resource configuration in cloud computing infrastructures
abstract
Power is becoming an increasingly important concern for large-scale cloud computing systems. Meanwhile, cloud service providers leverage virtualization technologies to facilitate service consolidation and enhance resource utilization. However, the introduction of virtualization makes the cloud infrastructure more complex, and thus challenges cloud power management. In a virtualized environment, resource needs to be configured at runtime at the cloud, server and virtual machine levels to achieve high power efficiency. In addition, cloud power management should guarantee high users' SLA (service level agreement) satisfaction. In this paper, we present an adaptive power management framework in the cloud to achieve autonomic resource configuration. We propose a software and lightweight approach to accurately estimate the power usage of virtual machines and cloud servers. It explores hypervisor-observable performance metrics to build the power usage model. To configure cloud resources, we consider both the system power usage and the SLA requirements, and leverage learning techniques to achieve autonomic resource allocation and optimal power efficiency. We implement a prototype of the proposed power management system and test it on a cloud testbed. Experimental results show the high accuracy (over 90%) of our power usage estimation mechanism and our resource configuration approach achieves the lowest energy usage among the compared four approaches.
Qiang Guan, Song Fu
IPCCC3
2012 Efficient and Accurate Anomaly Identification Using Reduced Metric Space in Utility Clouds
abstract
The online detection of anomalies is a vital element of operations in utility clouds. Detection should function for different levels of abstraction including hardware and software, and for the various metrics used in cloud computing systems. Given ever-increasing cloud sizes coupled with the complexity of system components, continuous monitoring leads to the overwhelming volume of data collected by health monitoring tools. High metric dimensionality and existence of interacting metrics compromise the detection accuracy and lead to high detection complexity. In this paper, we present a metric selection framework and propose systematic approaches to effectively identify and select the most essential metrics for online anomaly detection in utility clouds. Specifically, a mutual information based approach selects metrics with the maximized mutual relevance and the minimized redundancy. Then metric space combination and separation are explored to reduce the metric dimensionality further. Experimental results on utility cloud scenarios demonstrate the viability and efficiency of this framework. The selected metrics contribute to a high efficiency and accuracy in anomaly detection.
Qiang Guan, Chi-Chen Chiu, Song Fu
NAS4
2012 CDA: A Cloud Dependability Analysis Framework for Characterizing System Dependability in Cloud Computing Infrastructures
abstract
Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Dependability assurance is crucial for building sustainable cloud computing services. Although many techniques have been proposed to analyze and enhance reliability of distributed systems, there is little work on understanding the dependability of cloud computing environments. As virtualization has been an enabling technology for the cloud, it is imperative to investigate the impact of virtualization on the cloud dependability, which is the focus of this work. In this paper, we present a cloud dependability analysis (CDA) framework with mechanisms to characterize failure behavior in cloud computing infrastructures. We design the failure-metric DAGs (directed a cyclic graph) to analyze the correlation of various performance metrics with failure events in virtualized and non-virtualized systems. We study multiple types of failures. By comparing the generated DAGs in the two environments, we gain insight into the impact of virtualization on the cloud dependability. This paper is the first attempt to study this crucial issue. In addition, we exploit the identified metrics for failure detection. Experimental results from an on-campus cloud computing test bed show that our approach can achieve high detection accuracy while using a small number of performance metrics.
Qiang Guan, Chi-Chen Chiu, Song Fu
PRDC3
2012 AAD: Adaptive Anomaly Detection System for Cloud Computing Infrastructures
abstract
Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Autonomic failure detection is a crucial technique for understanding emergent, cloudwide phenomena and self-managing cloud resources for system-level dependability assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled, and thus a prior failure history is not always available in production clouds, especially for newly managed or deployed systems. In this paper, we present an Adaptive Anomaly Detection (AAD) framework for cloud dependability assurance. It employs data description using hypersphere for adaptive failure detection. Based on the cloud performance data, AAD detects possible failures, which are verified by the cloud operators. They are confirmed as either true failures with failure types or normal states. The algorithm adapts itself by recursively learning from these newly verified detection results to refine future detections. Meanwhile, it exploits the observed but undetected failure records reported by the cloud operators to identify new types of failures. We have implemented a prototype of the algorithm and conducted experiments in an on-campus cloud computing environment. Our experimental results show that AAD can achieve more efficient and accurate failure detection than other existing scheme.
Husanbir Singh Pannu, Jianguo Liu 0001, Song Fu
SRDS3
2011 Proactive Failure Management by Integrated Unsupervised and Semi-Supervised Learning for Dependable Cloud Systems
abstract
Cloud computing systems continue to grow in their scale and complexity. They are changing dynamically as well due to the addition and removal of system components, changing execution environments, frequent updates and upgrades, online repairs and more. In such large-scale complex and dynamic systems, failures are common. In this paper, we present a failure prediction mechanism exploiting both unsupervised and semi-supervised learning techniques for building dependable cloud computing systems. The unsupervised failure detection method uses an ensemble of Bayesian models. It characterizes normal execution states of the system and detects anomalous behaviors. After the anomalies are verified by system administrators, labeled data are available. Then, we apply supervised learning based on decision tree classier to predict future failure occurrences in the cloud. Experimental results in an institute-wide cloud computing system show that our proposed method can forecast failure dynamics with high accuracy.
Qiang Guan, Song Fu
ARES3
2011 Characterizing Power and Energy Usage in Cloud Computing Systems
abstract
Power and energy are primary concerns in the design and management of modern cloud computing systems and data centers. Operational costs for powering and cooling large-scale cloud systems will soon exceed acquisition costs. To improve the energy effciency of cloud computing systems and applications, it is critical to profile the power usage of real systems and applications. Many factors influence power and energy usage in cloud systems, including each components electrical specification, the system usage characteristics of the applications, and system software. In this work, we present the power profiling results on a cloud test bed. We combine hardware and software that achieves power and energy profiling at server granularity. We collect the power and energy usage data with varying server/cloud configurations, and quantify their correlation. Our experiments reveal conclusively how different system configurations affect the server/cloud power and energy usage.
Song Fu
CloudCom2
2011 Performance Metric Selection for Autonomic Anomaly Detection on Cloud Computing Systems
abstract
With ever-growing complexity and dynamicity of cloud computing systems, dependability assurance has become a major concern in system design and management. In this paper, we propose a framework for autonomic anomaly detection in the cloud. Mutual information is exploited to quantify the relevance and redundancy among the large number of performance metrics. An incremental search algorithm is presented for metric selection. We apply principal component analysis to further reduce the metric dimension, while keeping the variance in the health- related data as much as possible. A detection mechanism with semi-supervised decision tree classifiers works on the reduce metric dimensionality and identifies anomalies. We have implemented a prototype of our autonomic anomaly detection framework and evaluated its performance on an institute-wide cloud computing system.
Song Fu
GLOBECOM1
2011 Ensemble of Bayesian Predictors for Autonomic Failure Management in Cloud Computing
abstract
In modern cloud computing systems, hundreds and even thousands of cloud servers are interconnected by multi-layer networks. In such large-scale and complex systems, failures are common. Proactive failure management is a crucial technology to characterize system behaviors and forecast failure dynamics in the cloud. To make failure predictions, we need to monitor the system execution and collect health-related runtime performance data. However, in newly deployed or managed cloud systems, these data are usually unlabeled. Supervised learning based approaches are not suitable in this case. In this paper, we present an unsupervised failure detection method using an ensemble of Bayesian models. It estimates the probability distribution of runtime performance data collected by health monitoring tools when cloud servers perform normally. It characterizes normal execution states of the system and detects anomalous behaviors. Experimental results in an institute-wide cloud computing system show that our methods can achieve high true positive rate and low false positive rate for proactive failure management.
Qiang Guan, Song Fu
ICCCN3
2011 Macropower: A coarse-grain power profiling framework for energy-efficient cloud computing
abstract
Power and energy consumption has become a major concern in modern data centers and cloud systems. In order to develop efficient power management mechanisms for green clouds, we need a deep understanding of the influence of system configurations on the power consumption in real cloud systems. Power profiling provides such a vehicle. Existing fine-grain profiling approaches require special hardwired connections to the pins of individual hardware devices, which is not practical for large-scale production clouds. Moreover, they cannot provide a macroscopic view of the cloud-wide power dynamics. In this paper, we present macropower, a coarse-grain power and energy profiling framework. It provides a combination of hardware and software tools that achieves power/energy profiling at server granularity. It uses direct or derived measurements to isolate and combine influences from system components in cloud power profiles. It also generates the correlations between system activities and server/cloud-wide power/energy usage. We implement a prototype of macropower and test it in a cloud testbed. The profiled data are analyzed and the impact of system configurations on the server/cloud power usage is quantified, which is valuable for autonomic and energy-efficient management of cloud resources.
Song Fu
IPCCC2
2011 Randomized load balancing strategies with churn resilience in peer-to-peer networks
Song Fu, Cheng-Zhong Xu 0001, Haiying Shen
J. Netw. Comput. Appl.1
2010 Anomaly detection in large-scale coalition clusters for dependability assurance
abstract
In large-scale high-performance computing systems, component failures become norms instead of exceptions. Failure occurrence as well as its impact on system performance and operation costs are becoming an increasingly important concern to system designers and administrators. When a compute node fails to function properly, health-related data are valuable for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. Manual detection is time-consuming and error-prone. It does not scale well. In this paper, we present an autonomic mechanism for anomaly detection in coalition clusters. It is composed of a set of techniques that facilitates automatic analysis of system health data. We apply data transformation to format health data in a uniform manner. Then principal variables are chosen by feature selection, which reduces the data size. Clustering and outlier detection are explored to identify nodes with anomalous behavior. We evaluate our prototype implementation on a production institution-wide computational grid. The results show that our mechanism can effectively detect faulty nodes with high accuracy and low computation overhead.
Qiang Guan, Derek Smith, Song Fu
HiPC3
2010 auto-AID: A data mining framework for autonomic anomaly identification in networked computer systems
abstract
Networked computer systems continue to grow in scale and in the complexity of their components and interactions. Component failures become norms instead of exceptions in these environments. A failure will cause one or multiple computer(s) to be unavailable, which affects the resource utilization and system throughput. When a computer fails to function properly, health-related data are valuable for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. In this paper, we present auto-AID, an autonomic mechanism for anomaly identification in networked computer systems. It is composed of a set of data mining techniques that facilitates automatic analysis of system health data. The identification results are very valuable for the system administrators to manage systems and schedule the available resources. We implement a prototype of auto-AID and evaluate it on a production institution-wide compute grid. The results show that auto-AID can effectively identify anomalies with little human intervention.
Qiang Guan, Song Fu
IPCCC2
2010 Profiling and analysis of power consumption for virtualized systems and applications
abstract
In this paper, we present a power/energy profiling framework for collecting and analyzing energy consumption data in compute clouds. We implement a prototype of our profiling framework and test it in a cloud computing environment. By measuring the power/ energy consumption with various combinations of system configuration and settings, we build knowledge of the extent to which each factor influences the system power dynamics. As a future work, we plan to explore the collected profiling data and design resource management mechanisms that consider both the performance requirements of applications and energy budget of a cloud system for green computing.
Song Fu
IPCCC2
2010 Dependability enhancement for coalition clusters with autonomic failure management
abstract
In large-scale compute clusters, failures become norms instead of exceptions. Autonomic management of failures and resources in such systems is becoming more and more important. In this paper, we propose a failure management mechanism with prediction functionality for large coalition systems. It analyzes failure behaviors in a system and forecasts the prospective failure occurrences based on characterized failure dynamics. A prototype of our failure management system is implemented and deployed in a coalition cluster environment. Prediction results in the experiments show our proposed mechanism can accurately capture the failure trend in the coalition cluster.
Song Fu
ISCC1
2010 Failure-aware resource management for high-availability computing clusters with distributed virtual machines
Song Fu
J. Parallel Distributed Comput.1
2010 Quantifying event correlations for proactive failure management in networked computing systems
Song Fu, Cheng-Zhong Xu 0001
J. Parallel Distributed Comput.1
2009 Proactive Resource Management for Failure Resilient High Performance Computing Clusters
abstract
Virtual machine (VM) technology provides an additional layer of abstraction for resource management in high-performance computing (HPC) systems. In large-scale computing clusters, component failures become norms instead of exceptions, caused by the ever-increasing system complexity. VM construction and reconfiguration is a potent tool for efficient online system maintenance and failure resilience. In this paper, we study how VM-based HPC clusters benefits from failure prediction in resource management for dependable computing. We consider both the reliability and performance status of compute nodes in making selection decisions. We define a capacity-reliability metric to combine the effects of both factors, and propose the Best-fit algorithm to find the best qualified nodes on which to instantiate VMs to run user jobs. We have conducted experiments using failure traces from the Los Alamos National Laboratory (LANL) HPC clusters. The results show the enhancement of system dependability by using our proposed strategy with practically achievable accuracy of failure prediction. With the Best-fit strategies, the job completion rate is increased by 10.5% compared with that achieved in the current LANL HPC cluster. The task completion rate reaches 82.5% with improved utilization of relatively unreliable nodes.
Song Fu, Cheng-Zhong Xu 0001
ARES1
2009 Failure-Aware Construction and Reconfiguration of Distributed Virtual Machines for High Availability Computing
abstract
In large-scale clusters and computational grids, component failures become norms instead of exceptions. Failure occurrence as well as its impact on system performance and operation costs have become an increasingly important concern to system designers and administrators. In this paper, we study how to efficiently utilize system resources for high-availability clusters with the support of the virtual machine (VM) technology. We design a reconfigurable distributed virtual machine (RDVM) infrastructure for clusters computing. We propose failure-aware node selection strategies for the construction and reconfiguration of RDVMs. We leverage the proactive failure management techniques in calculating nodes' reliability status. We consider both the performance and reliability status of compute nodes in making selection decisions. We define a capacity-reliability metric to combine the effects of both factors in node selection, and propose best-fit algorithms to find the best qualified nodes on which to instantiate VMs to run parallel jobs. We have conducted experiments using failure traces from production clusters and the NAS parallel benchmark programs on a real cluster. The results show the enhancement of system productivity and dependability by using the proposed strategies. With the best-fit strategies, the job completion rate is increased by 17.6% compared with that achieved in the current LANL HPC cluster, and the task completion rate reaches 91.7%.
Song Fu
CCGRID1
2008 Random choices for churn resilient load balancing in peer-to-peer networks
abstract
Peer-to-peer (P2P) networks based on consistent hashing functions have an inherent load uneven distribution problem. Things are even worse in unstructured P2P systems. The objective of load balancing in P2P networks is to balance the workload of the network nodes in proportion to their capacity so as to eliminate traffic bottleneck. It is challenging because of the dynamic nature of overlay networks and time-varying load characteristics. Random choices schemes can balance load effectively while incurring only a small overhead, making such schemes appealing for practical systems. Existing theoretical work analyzing properties of random choices algorithms can not be applied in the highly dynamic and heterogeneous P2P systems. In this paper, we characterize the behaviors of randomized search schemes in the general P2P environment. We extend the supermarket model by investigating the impact of node heterogeneity and churn to the load distribution in P2P networks. We prove that by using d-way random choices schemes, the length of the longest queue in P2P systems with heterogeneous nodal capacity and node churn for d ≥ 2 is clog logn/logd + O(1) with high probability, where c is a constant.
Song Fu, Cheng-Zhong Xu 0001, Haiying Shen
IPDPS1
2007 Exploring event correlation for failure prediction in coalitions of clusters
abstract
In large-scale networked computing systems, component failures become norms instead of exceptions. Failure prediction is a crucial technique for self-managing resource burdens. Failure events in coalition systems exhibit strong correlations in time and space domain. In this paper, we develop a spherical covariance model with an adjustable timescale parameter to quantify the temporal correlation and a stochastic model to describe spatial correlation. We further utilize the information of application allocation to discover more correlations among failure instances. We cluster failure events based on their correlations and predict their future occurrences. We implemented a failure prediction framework, called PREdictor of Failure Events Correlated Temporal-Spatially (hPREFECTs), which explores correlations among failures and forecasts the time-between-failure of future instances. We evaluate the performance of hPREFECTs in both offline prediction of failure by using the Los Alamos HPC traces and online prediction in an institute-wide clusters coalition environment. Experimental results show the system achieves more than 76% accuracy in offline prediction and more than 70% accuracy in online prediction during the time from May 2006 to April 2007.
Song Fu, Cheng-Zhong Xu 0001
SC1
2007 Quantifying Temporal and Spatial Correlation of Failure Events for Proactive Management
abstract
Networked computing systems continue to grow in scale and in the complexity of their components and interactions. Component failures become norms instead of exceptions in these environments. Moreover, failure events exhibit strong correlations in time and space domain. In this paper, we develop a spherical covariance model with an adjustable timescale parameter to quantify the temporal correlation and a stochastic model to characterize spatial correlation. The models are further extended to take into account the information of application allocation to discover more correlations among failure instances. We cluster failure events based on their correlations and predict their future occurrences. Experimental results on a production coalition system, the Wayne State Grid, show the offline and online predictions by our predicting system can forecast 72.7% to 85.3% of the failure occurrences and capture failure correlations in cluster coalition environment.
Song Fu, Cheng-Zhong Xu 0001
SRDS1
2007 Coordinated access control with temporal and spatial constraints on mobile execution in coalition environments
Song Fu, Cheng-Zhong Xu 0001
Future Gener. Comput. Syst.1
2006 Stochastic modeling and analysis of hybrid mobility in reconfigurable distributed virtual machines
Song Fu, Cheng-Zhong Xu 0001
J. Parallel Distributed Comput.1
2005 Service Migration in Distributed Virtual Machines for Adaptive Grid Computing
abstract
Computational grids can integrate geographically distributed resources into a seamless environment. To facilitate managing these heterogeneous resources, the virtual machine technology provides a powerful layer of abstraction and allows multiple applications to multiplex the resources of a grid computer. On the other hand, the grid dynamics requires the virtual machine system be distributed and reconfigurable. However, the existing migration approaches only move the execution entities, such as processes, threads, and mobile agents, among servers and leave the runtime services behind. They are not potent to achieve service reconfiguration in face of server overload or failures. In this paper, we propose a service migration mechanism, which moves the computational services of a virtual server, for instance a shared array runtime support system, to available servers for adaptive grid computing. In this way, parallel jobs can resume computation on a remote server without requiring service preinstallation. As an illustration of the service migration mechanism, we incorporated it into a Java-compliant distributed virtual machine, DSA, and formed a Mobile DSA (M-DSA) to accommodate adaptive parallel applications in grids. We measured the performance of M-DSA in the execution of applications from the SPLASH-2 benchmark suite on a campus grid. Experimental results show that service migration can achieve system adaptivity effectively.
Song Fu, Cheng-Zhong Xu 0001
ICPP1
2005 Distributed Shared Arrays: An Integration of Message Passing and Multithreading on SMP Clusters
Ramzi Basharahil, Brian Wims, Cheng-Zhong Xu 0001, Song Fu
J. Supercomput.4
2004 Migration Decision for Hybrid Mobility in Reconfigurable Distributed Virtual Machines
abstract
Virtual machine (VM) is an important mechanism to multiplex computer resources. The increasing popularity of network computing has renewed research interests in the adaptive and distributed virtual machines. Service migration is a vital technique to construct reconfigurable VMs. By incorporating mobile agent technology, VM systems can improve their resource utilization, load-balancing and fault-tolerance significantly. This work focuses on the decision problem of hybrid mobility for load-balancing in reconfigurable distributed VMs. We tackle this problem from three aspects: migration candidate determination, migration timing and destination server selection. The service migration timing and destination server selection are formulated as two optimization models. We derive the optimal migration policy for distributed and heterogeneous systems based on stochastic optimization theories. Renewal processes are applied to model the dynamics of migration. We solve the agent migration problem by dynamic programming and extend the optimal service migration decision by considering the interplay of the hybrid mobility. Our decision policy is complementary to the existing service and agent migration techniques. Its accuracy is verified by simulations.
Song Fu, Cheng-Zhong Xu 0001
ICPP1
2003 A Rate-Based Multicast Protocol for Large-Scale Reliable Transport
abstract
This paper presents a RAte-based Multicast Protocol for lArge-scale Reliable Transport (RAMPART), which is designed to provide scalable, TCP-friendly and responsive reliable multicast transport service for applications of bulk-data transfer. We proposed a dual-bitmap mechanism to solve the problem of duplicate packets and eliminate unnecessary retransmissions. A new scalable indirect RTT measurement algorithm is introduced. And Specific Receiver selection process is distributed among all group members with records of Local/spl I.bar/SR/spl I.bar/rate and Global/spl I.bar/SR/spl I.bar/rate assuring correct SR selection and avoiding frequent SR changes. Control tree is constructed with the aid of multicast routers using a label-based sorting algorithm. RAMPART protocol was implemented on NS. Simulations were conducted with respect to RTT measurement, throughput analysis and TCP-friendliness. The results show RAMPART is effective.
Song Fu, Zhiquan Jin
AINA1