Xiaohua Shi

dblp:82/3923 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Automatic Compiler Tuning for Inlining with Machine Learning
abstract
Automatic compiler tuning has become one of the hot research areas attracting extensive attention in recent years. By using machine learning models to learn from large-scale experimental samples, it can efficiently search through the massive combinations of compiler optimization parameters within limited time, identifying parameter sets that better suit the current microprocessor architecture and application programs, thereby achieving higher execution efficiency of target programs than default compilation options. This paper designs and implements an automatic compiler tuning mechanism, employing an XGBoost performance model built through a “learn-while-searching” approach to guide the search process of simulated annealing algorithm (SA). The XGBoost model can predict the performance deltas of different GCC optimization parameters, and these deltas are used to bias the annealing acceptance probability. This mechanism combines the global modeling capability of machine learning with the local search advantages of the SA, enhancing both search efficiency and effectiveness. Based on this tuning mechanism, the paper conducts in-depth research on the inlining optimization process inside the GCC compiler, proposing and implementing two automatic tuning schemes: (1) automatic tuning of inlining compilation options, which influences the number of inlined functions through different values of these options, thereby affecting program performance; (2) automatic tuning of function inlining at call sites, which uses GCC plugins technology to replace the ipa-inline pass in GCC optimization, enabling control over whether to inline functions at each call site and thus impacting program performance. Extensive experiments are conducted on benchmark suites such as SPEC CPU2006, cBench, and WRF. The results show that compared with the Critical Flag Selection-based Compiler Auto-tuning (CFSCA) method published in ASE2023, the Bayesian Optimization-based Compiler Auto-tuning (BOCA) method published in ICSE2021, and other traditional search algorithms, the proposed tuning mechanism achieves better optimization effects on most benchmarks.
Xiaohua Shi, Yuchen Feng 0001, Changhai Zhao, Jiamin Wen, Minqiang Shang
Int. J. Pattern Recognit. Artif. Intell.2
2025 Optimizing AutoTVM by Parallel Genetic Algorithms
abstract
To make deep neural networks automatically achieve the same or better performance compared with those in hand-optimized libraries, Tensor Virtual Machine (TVM) has combined a genetic algorithm (GA) with its AutoTVM auto-tuning process. The genetic algorithm of TVM has a primary and classic design, with restrictions in terms of searching scope, ability, and efficiency. Meanwhile, the current AutoTVM process is time-consuming. The whole process may last hours on GPUs. As such, we propose a new auto-tuning method that is based on a parallel GA and takes advantage of the strengths of the Roofline model-based cost models and machine learning classification models to widen the search scope and improve search efficiency. The new auto-tuning method achieves double optimization on both tuning results and tuning time. A series of experiments show that the new way improves the inference time of typical deep networks by about 8–14% and speeds up the time consumption of the auto-tuning process up to 1.2–[Formula: see text] on GPUs compared with the original GA process of AutoTVM.
Changhai Zhao, Jiamin Wen, Minqiang Shang, Yuchen Feng 0001, Xiaohua Shi
Int. J. Pattern Recognit. Artif. Intell.5
2022 Deep Multi-mode Learning for Book Spine Recognition
Wanru Yang, Xiaohua Shi
WISA2
2022 Optimizing an AltaRica simulator
abstract
Summary Since the AltaRica language has, in fact, become the European industry standard, the AltaRica‐based simulation method is one of the important research contents in the field of safety analysis methods. At the same time, in practical applications, simulations are mostly executed tens of thousands of times. Therefore, it is practical to optimize the execution time of the simulation method. We first analyze the hotspots in the execution of simulation methods and then propose two optimization strategies for two hot issues. The optimization method is designed and implemented by the method of space for time. Finally, the effectiveness and application scenarios of the optimization method are illustrated by experiments with different scale models. Moreover, the experiment data shows that the optimization method is effective in most cases.
Wenru Wang, Xinghai Lu, Xiaohua Shi
Concurr. Comput. Pract. Exp.5
2021 DATAM: A model-based tool for dependability analysis
abstract
Summary Dependability analysis is the main method to evaluate the design of safety‐critical systems, which is able to analyze the source of the faults and find them as early as possible. With the increasing system scale and complexity, Model‐Based Dependability Analysis (MDBA) has become the mainstream, so that it is crucial to provide powerful models that accurately reflect the real systems and easy to be built. However, modeling a system is error‐prone and it is difficult to verify the correctness of the model being built. Therefore, this paper proposes our dependability modeling tool called DATAM (Dependability Analysis Tool of AltaRica Model), which is based on AltaRica, a dataFlow language. We present a method for converting key elements of AltaRica into model/GUI components, thereby ensuring the consistency of the model with the modeling language. GUI‐based operations and rich custom components ensure ease of using DATAM. Besides, some components of the model can be edited directly through the AltaRica script and can be interchanged with GUI components. Finally, DATAM supplies varieties of reliability calculation functions. We demonstrate the ability of DATAM through a case study and compare the results with SimFia education version.
Xinghai Lu, Xiaohua Shi, Wenru Wang
Concurr. Comput. Pract. Exp.2
2021 A safety simulation analysis algorithm for Altarica language
abstract
Summary Altarica is a modeling language for safety analysis and supports simulation analysis. Although Altarica is widely used in the industry, research studies on simulation algorithm are rarely found. Therefore, we design and implement a simulation algorithm. We first briefly introduce the syntax and characteristics of Altarica, and then describe the design and implementation of the algorithm in detail, and finally accurately simulate the most probable sequence of events. We use Reverse Polish Notation to deal with complex event triggered conditions, and support three kinds of synchronization, including Synchronization, Broadcasting and Common Cause Failure, and support multiple probability distribution types. At the end of this paper, through case study and comparison with SIMFIA, the correctness of the algorithm is proved.
Wenru Wang, Xiaohua Shi, Xinghai Lu
Concurr. Comput. Pract. Exp.2
2019 Community detection in scientific collaborative network with bayesian matrix learning
Xiaohua Shi, Hongtao Lu 0001
Frontiers Comput. Sci.1
2017 Design and study of leg automatic leveling control system of military bridge
abstract
Against the shortages as low precision and long operating time of traditional leg leveling method of military bridge, design automatic leveling control system for bridge with four legs. First establish the leg leveling space model of the bridge, analyze the relationship between the leg regulated quantity and the horizontal regulated angle, and design automatic leveling control system based on “high point chase” leveling scheme. Then establish SIMULINK simulation model of leg automatic leveling control system based on adaptive fuzzy PID algorithm. The simulation experiment results demonstrate that the leveling algorithm and automatic leveling control scheme are correct and feasible, and could improve the leg leveling speed, accuracy and automation level of military bridge equipment.
Xiaohua Shi, XiaoQing Zhu
ICIS2
2017 Adaptive Overlapping Community Detection with Bayesian NonNegative Matrix Factorization
Xiaohua Shi, Hongtao Lu 0001, Guanbo Jia
DASFAA (2)1
2016 Community Inference with Bayesian Non-negative Matrix Factorization
Xiaohua Shi, Hongtao Lu 0001
APWeb (1)1
2016 Development of electrostatic decay time intelligent test instrument and software design
abstract
The problems come from ESD become more and more serious. To evaluate the capability of ESD protection of materials entirely, design a multi-function electrostatic charge decay intelligent test instrument based on single chip compute. It adopts a new kind of efficient electrostatic decay time test scheme, so as to satisfy test condition that different materials have the same initial and terminal potential, and make test results not influenced by Cross-Over effect. Then introduce the system composition and operating principle of intelligent test instrument, and describe the specific software design method from man-machine interaction function, electrostatic potential measurement function and decay time calculation function, finally perform charging method test research with designed intelligent test instrument and analyze the test results.
Xiaohua Shi, HongYu Yu
SNPD2
2015 Community Detection in Social Network with Pairwisely Constrained Symmetric Non-Negative Matrix Factorization
abstract
Non-negative Matrix Factorization (NMF) aims to find two non-negative matrices whose product approximates the original matrix well, and is widely used in clustering condition with good physical interpretability and universal applicability. Detecting communities with NMF can keep non-negative network physical definition and effectively capture communities-based structure in the low dimensional data space. However some NMF methods in community detection did not concern with more network inner structures or existing ground-truth community information.
Xiaohua Shi, Hongtao Lu 0001, Yangcheng He
ASONAM1
2015 LeakTracer: Tracing leaks along the way
abstract
Unnecessary references in managed languages, such as Java and C#, often cause memory leaks without any immediate symptoms. These leaks become manifest when the program has been running for a long time (usually several hours, days or even weeks). Garbage collectors cannot handle this situation, since it only reclaims objects that have no external references to them. Consequently, when the number of leaked objects becomes large, garbage collection frequency increases and program performance degrades. Ultimately, the program will crash. This paper introduces LeakTracer, a tool that helps diagnose memory leaks in managed languages. The core of LeakTracer is the use of a novel leak predictor, which not only considers object size and staleness as a whole to predict leaked objects, but also carefully adjusts their contributions to the leak possibility of an object, according to the careful observation of activities of common objects during their lifetimes. We have implemented LeakTracer in two parts: (1) an online object events tracker in the Apache Harmony DRL virtual machine, and (2) an offline analyzer embedding our predictor. We have successfully used LeakTracer to find leaks in several real-world programs, and our case studies show that leak predictor can pinpoint leaked objects with high accuracy.
Hengyang Yu, Xiaohua Shi
SCAM2
2015 Non-negative Matrix Factorization with Pairwise Constraints and Graph Laplacian
Yangcheng He, Hongtao Lu 0001, Lei Huang 0005, Xiaohua Shi
Neural Process. Lett.4
2015 Parallelization of a color-entropy preprocessed Chan-Vese model for face contour detection on multi-core CPU and GPU
Xiaohua Shi, Fredrick Park, Jack Xin, Yingyong Qi
Parallel Comput.1
2014 Testing system for CAN bus-oriented embedded software
abstract
Based on the analysis on the characteristics of embedded system based on CAN bus, this paper proposed one CAN bus-oriented test system scheme. The scheme consists of a non-real time host and a real time target; the non-real time host constructs test data, while the real time target drives the data transmission, processes the response of tested system, and returns the results to the non-real time host for analyzing. The system can test, analyze and evaluate the CAN bus-oriented embedded application in a real time, closed-loop and nonintrusive way.
Shunkun Yang, Dongxiao Tang, Xiaohua Shi
ICIS3
2014 Address Chain: Profiling Java Objects without Overhead in Java Heaps
Xiaohua Shi, Junru Xie, Hengyang Yu
APLAS1
2014 An OpenCL micro-benchmark suite for GPUs and CPUs
Xiaohua Shi
J. Supercomput.2
2012 An OpenCL Approach of Prestack Kirchhoff Time Migration Algorithm on General Purpose GPU
abstract
OpenCL is an open standard for portable, parallel programming across heterogeneous platforms. In this paper, we presented how to implement and optimize Prestack Kirchhoff Time Migration algorithm, which is one of the most widely adopted imaging methods for seismic data processing, on OpenCL and GPGPU. We introduced how to port the original CUDA program to OpenCL, and how to optimize the OpenCL program to get the competitive performance comparing with the original CUDA version. Our OpenCL version of Kirchhoff Migration algorithm on NVidia 8800GT is 8.9 times faster than its original CPU version on AMD245 2.9HGZ, and almost as fast as its CUDA version.
Peiyuan Sun, Xiaohua Shi
PDCAT2
2012 Profiling Object Life Ranges for Detecting Memory Leaks in Java Virtual Machine
abstract
Java Virtual Machine (JVM) has its garbage collector (GC) to manage all the java objects and the heap memory. However, in many Java applications, some Java objects will not be visited by the programs, but could not be automatically reclaimed by the garbage collector. These Java objects would cause the memory leaks. Obtaining the life ranges of all Java objects in the heap memory could help programmers to find the kind of memory leaks. This paper presents two approaches to scan the garbage-collector-managed heap memory to record the living status of all Java object at runtime. One approach is called All-Scanning mechanism, and the other is called Incremental-Scanning mechanism. The All-Scanning mechanism scans heap objects after GC reclaiming the objects in a static way. The Incremental-Scanning mechanism could incrementally collect the status of heap objects when GC reclaims the objects. We use SpecJVM98 and SpecJBB2000 to evaluate the runtime performance and analyze the profiling results for both mechanisms. The profiling results of the two mechanisms are consistent, while the Incremental-Scanning mechanism is at least 1.82 times faster than the All-Scanning one on average. From the profiling results, we could figure out the life cycles of every object, the movements of the object allocations, and the potentially leaked objects. We also use three workloads, namely List Leak, Swap Leak and MySql, to demonstrate how to find out the leaked objects from the profiling results.
Qingyue Sun, Xiaohua Shi, Junru Xie
PDCAT2
2012 An OpenCL Micro-Benchmark Suite for GPUs and CPUs
abstract
OpenCL (Open Computing Language) is the first open, royalty-free standard for cross-platform, parallel programming of modern processors in personal computers, servers and handheld/embedded devices. OpenCL is vendor-independent and hence not specialized for any particular compute device. In order to develop efficient OpenCL applications for the particular platform, we still need a more profound understanding of the architecture features on the OpenCL model and computing devices. For this purpose, we design and implement an OpenCL micro-benchmark suite for GPUs and CPUs. We introduce the implementations of our OpenCL micro benchmarks and present the performance results of hardware and software features like the bus bandwidth, memory architectures, branch architectures and thread hierarchy, etc., evaluated by our micro benchmarks on multi-core X86 CPU and NVIDIA's GPU.
Xiaohua Shi, Qingyue Sun
PDCAT2
2009 A Practical Approach of Curved Ray Prestack Kirchhoff Time Migration on GPGPU
Xiaohua Shi
APPT1
2001 An RNN-based algorithm to detect prosodic phrase for Chinese TTS
abstract
The goal of the work presented here is to automatically predict the prosodic phrase boundaries from the text for Chinese TTS (text-to-speech) by using the trigram of the POS (part-of-speech) with information of the breaks between the prior two word-pairs by using a RNN (recurrent neural network). Prosodic phrase boundaries are very important to a Chinese TTS system because they will influence the prosodic model for speech synthesis. In this paper, the algorithm tries to use RNN to find some mapping relationship between the POS sequence and prosodic phrase boundaries, and hopes to improve the naturalness of synthesized speech.
Zhiwei Ying, Xiaohua Shi
ICASSP2