Song Fu

dblp:57/2335 · DBLP profile ↗
← Back
16ranked-venue papers in the field
1as first author
8since 2021 · last 2026
ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 7Other / Interdisciplinary · 7Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1 (1 first)
YearPublicationVenuePosition
2026 Performance evaluation method for modules based on organic fusion of data-driven methods and mechanistic knowledge
Wenhui He, Lin Lin 0014, Song Fu
Adv. Eng. Informatics3
2026 Gene recombination-guided convolution neural network for early fault diagnosis of aero-engines
Jinlei Wu, Lin Lin 0014, Song Fu, Lingyu Yue, Sihao Zhang
Adv. Eng. Informatics4
2026 Engine-specific degradation prediction of aviation engines via transferable snippet augmentation: A trend-grouped fine-tuning perspective
Minghang Zhao, Song Fu
Adv. Eng. Informatics6
2025 FD-LLM: Large language model for fault diagnosis of complex equipment
Sihao Zhang, Song Fu
Adv. Eng. Informatics3
2025 Continual contrastive reinforcement learning: Towards stronger agent for environment-aware fault diagnosis of aero-engines through long-term optimization under highly imbalance scenarios
Minghang Zhao, Song Fu
Adv. Eng. Informatics6
2024 Channel attention & temporal attention based temporal convolutional network: A dual attention framework for remaining useful life prediction of the aircraft engines
Lin Lin 0014, Jinlei Wu, Song Fu, Sihao Zhang, Changsheng Tong, Lizheng Zu
Adv. Eng. Informatics3
2024 PathEL: A novel collective entity linking method based on relationship paths in heterogeneous information networks
Lizheng Zu, Lin Lin 0014, Song Fu, Shiwei Suo, Wenhui He, Jinlei Wu, Yancheng Lv
Inf. Syst.3
2023 A novel method for aeroengine performance model reconstruction based on CDAE model
Lin Lin 0014, Wenhui He, Song Fu, Changsheng Tong, Lizheng Zu
Adv. Eng. Informatics4
2020 A middle-ware approach to leverage the distributed data de-duplication capability on HPC and Cloud storage systems
abstract
The unprecedented growth in the volume and diversity of the data in today's HPC and Enterprise computing environment has posted challenging problems on data management and data space reduction. More than 71% of enterprise and HPC communities are seeking de-duplication technologies to reduce cost and increase the storage efficiency. The importance of applying data de-duplication techniques is critical for active research and development. Current implementations of data de-duplication systems are mainly hardware dependent, system dependent, and platform dependent. Also, most of these implementations are proprietary software and not in open source domains. In this paper, we present a new middle-ware design and implementation approach, named D3M, to support distributed data de-duplication feature on existing file and object storage systems. We also incorporate this proposed D3M middle-ware with the Redhat's Linux device layer de-duplication and compression driver, called VDO (Virtual Data Optimizer). With these two layers of data de-duplication support, we accommodate both client side and server-side data de-duplication features. Finally, we conduct various testing cases on HPC data sets and Enterprise data sets to illustrate the benefits and advantages of applying our bilayer data de-duplication middle-ware solution.
Hsing-bung Chen, Sihai Tang, Song Fu
IEEE BigData3
2019 Applying SDN based data network on HPC Big Data Computing - Design, Implementation, and Evaluation
abstract
Large scale storage data networks are difficult to conFigure and tricky to maintain [1][2][3]. As storage data volumes grow and the pace of change accelerates, it can be a struggle to keep up the integrity of a large scale storage network. Software Defined Networking (SDN) [5][6][7][8][10][11] provides a method to centrally conFigure and manage physical and virtual network devices such as routers, switches, and gateways in HPC datacenter. Current HPC computer cluster and data centers are using homogeneous data network technologies such as the Infiniband network (QDR, FDR, EDR, and HDR). Due to the connector, cable, and bandwidth backward compatibility issues, once an aged HPC computing system was retired, we have to demolish all established computer clusters, data network, and data storage. Eventually we lost all costly investment on the data network and storage. Those current deployment approaches do not provide a feasible and compatible growing path and cannot meet the need for future extreme scale HPC computing systems [4][12][13].
Hsing-bung Chen, Zhi Qiao 0001, Song Fu
IEEE BigData3
2019 Accelerating RNN on FPGA with Efficient Conversion of High-Level Designs to RTL
abstract
Recurrent Neural Network (RNN) is a powerful Deep Learning algorithm which has been widely used for speech recognition, handwriting recognition, context clustering, etc. Deep learning involves a large number of floating-point computations and requires a large amount of computing resource. Currently, software-based RNN implementations running on general-purpose processors like CPU and GPU take a long execution time and an excessive amount of energy. FPGA provides a low-power and highly parallel platform for accelerating RNN training and inference. With massive reconfigurable logics, FPGA can be customized to achieve dramatic speedup and energy efficiency for various deep learning applications. However, implementing RNN algorithms at the Register Transistor Level (RTL) is time-consuming, while existing tools for converting RNN code written in high-level languages to RTL designs are not efficient. In this paper, we present a design flow by which high-level RNN implementations (for example, in Python) can be converted to RTL designs efficiently and automatically. Experimental results show the generated RTL code that is deployed on an FPGA device is 7.87 times faster than the Python code run on CPU while achieving the same accuracy.
Zongze Li 0001, Song Fu
IEEE BigData2
2019 An Empirical Study of Quad-Level Cell (QLC) NAND Flash SSDs for Big Data Applications
abstract
As the SSD technology develops, quad-level cell (QLC) NAND based SSD is gradually being introduced to the market. And as such, we evaluate the QLC technologys impact on the landscape of modern datacenters. Since a large number of applications and workloads in the modern datacenters have far more read requests than they write requests, QLC SSD provides a promising solution. For example, real-time analytics and big data, machine and deep learning, and read-intensive AI applications are all read hungry perfectly suited for the QLC. Its favorable performance (especially in reads), high capacity and density greatly help modern datacenters to provide more efficient services to their customers. At the same time, the low cost of QLC SSD also helps to lower the cost of operation for datacenters. Additionally, we explore the state-of-art QLC SSD from the system architecture point of view to shows its key advancements from previous technologies. By conducting a comprehensive performance evaluation of QLC SSD, we are able to compare it with other types of SSD and analyze factors that impact its performance.
Shuwen Liang, Zhi Qiao 0001, Sihai Tang, Jacob Hochstetler, Song Fu, Weisong Shi, Hsing-bung Chen
IEEE BigData5
2019 Smart Home IoT Anomaly Detection based on Ensemble Model Learning From Heterogeneous Data
abstract
Nowadays, internet based home automation is made possible with the advent of intelligent device control. These electronic sensing devices transfer an enormous amount of data into the cloud. It is a challenge to discover hidden information from the massive amount of stored data in the cloud. In addition, privacy, security, and stability could also be a concern for users. Due to these issues becoming ever more prevalent in today’s society, the need to have access to readily anomaly detection becomes crucial for the modern smart home user. In this paper, we design, test and evaluate an ensemble model anomaly detection method. Our method targets the data anomalies present in general smart Internet of Things (IoT) devices, allowing for easy detection of anomalous events based on stored data. We make our method robust through ensemble machine learning model training. We aim to simulate different types of anomaly situations on publicly available smart home data sets, thereby exposing our models to likely real world phenomenons and events that may cause anomalies. Experiments are conducted on the processed data and evaluated for accuracy through validation and testing against independent and identically distributed labeled data.
Sihai Tang, Zhaochen Gu, Qing Yang 0003, Song Fu
IEEE BigData4
2018 Reliability Characterization of Solid State Drives in a Scalable Production Datacenter
abstract
In recent years, NAND flash-based solid state drives (SSD) have been widely used in datacenters due to their better performance compared with the traditional hard disk drives. However, little is known about the reliability characteristics of SSDs in production systems. Existing works study the statistical distributions of SSD failures in the field. However, they do not go deep into SSD drives and investigate the unique error types and health dynamics that distinguish SSDs from hard disk drives. In this paper, we explore the SSD-specific SMART (Self-Monitoring, Analysis, and Reporting Technology) attributes to conduct an in-depth analysis of SSD reliability in a production environment. Data is collected from a scalable production system having several physical locations. Our dataset contains over a million records with more than twenty attributes. We leverage machine learning technologies, specifically data clustering and correlation analysis methods, to discover groups of SSDs which have different health status and relations among SSD-specific SMART attributes. Our results show that 1) Media wear affects the reliability of SSDs more than any other factors, and 2) SSDs transit from one health group to another which infers the reliability degradation of those drives. To the best of our knowledge, this is the first study that investigates SSD-specific SMART data to characterize SSD reliability in a production environment.
Shuwen Liang, Zhi Qiao 0001, Jacob Hochstetler, Song Fu, Weisong Shi, Devesh Tiwari, Hsing-bung Chen, Bradley W. Settlemyer, David Richard Montoya
IEEE BigData5
2018 ACTOR: Active Cloud Storage with Energy-Efficient On-Drive Data Processing
abstract
Storage systems are indispensable for big data processing and cloud computing services today. The ever-growing size of computation and data analytic results demands larger storage capacity, which challenges data processing and storage scalability. Moreover, the increasing complexity of storage hierarchy and "passive" storage devices make todays storage systems inefficient, which necessitates the adoption of new storage technologies. In this paper, we explore new Ethernet connected drives with on-drive embedded CPU and DRAM to develop an active cloud storage system where data can be processed on disk drives without data movement. These drives are micro-storage servers that can support software-defined storage. In addition to I/O operations, we test and evaluate on-drive data processing, including data compression, aggregation and erasure encoding, which provide natural support for data-intensive applications. Our experimental results show that Open Ethernet Drive can significantly lower the energy consumption while maintaining the data processing throughput simultaneously by ensuring data availability and storage scalability. Results and findings from this work will facilitate scheduling of on-drive compute resource for building active and scalable cloud storage systems.
Zhi Qiao 0001, Shuwen Liang, Nandini Damera, Song Fu, Hsing-bung Chen, Michael Lang 0003
IEEE BigData4
2012 A Hybrid Anomaly Detection Framework in Cloud Computing Using One-Class and Two-Class Support Vector Machines
Song Fu, Jianguo Liu 0001, Husanbir Singh Pannu
ADMA1