Xin Huo

dblp:65/9705 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Predictive Model-Assisted Iterative Learning Control for Suppressing Unknown Periodic Disturbances
abstract
This paper presents a feedforward online predictive model-assisted iterative learning control (PMA-ILC) method, addressing the critical challenge of achieving higher tracking accuracy along with strong task flexibility and disturbance suppression ability. The PMA-ILC framework comprises two key components: 1) a space-dependent oblique projection-based iterative learning control (SOBP-ILC) term for reducing tracking errors effectively without affecting feedback stability, 2) an online model predictive compensation term to dynamically generate an optimal feedforward signal at each sampling instant through iterative computation, significantly enhancing tracking precision and disturbance suppression. By integrating these strategies, the proposed method ensures superior performance and robustness, making it suitable for high-precision tracking applications. Simulation results demonstrate the effectiveness of the proposed method.
Aijing Wu, Xin Huo, Linrui Wang
IECON2
2025 Multi-Metric Fusion-Based Ensemble Clustering for Time Series via Diverse Elastic Distance Functions
abstract
The growing volume of unlabeled industrial process data necessitates the development of advanced time series clustering (TSCL) algorithms for exploratory analysis. To overcome the inherent limitations of conventional single-clustering algorithms in handling complex temporal patterns and enhance robustness, a novel ensemble TSCL framework, integrating multi-metric fusion and diverse elastic distance functions, is proposed in this article. Initially, to extract multi-domain patterns, diverse base clusterers are fit for augmented time series sequences via diverse elastic distance measures, with the performance evaluated on the training subsets through multiple internal clustering validation metrics. Then, a multi-metric fusion scheme, leveraging entropy to balance the contributions of metrics, is introduced to assign adaptive weights to each clusterer, which are then incorporated into the construction of a weighted co-association matrix. To enhance the discriminative power of the similarity structure and generate robust consensus partitions, an agglomerative hierarchical clustering algorithm with average linkage is applied to the refined distance matrix, which is transformed from the co-association matrix through a nonlinear mapping. Finally, the competitiveness of the proposed algorithm is demonstrated via the comparative and ablative experiments on datasets from the UCR archive.
Baohan Mi, Xin Huo, Changchun He, Chentao Liu
INDIN2
2025 Conflict-Averse Multi-Objective Extremum Seeking and Its Application in Wastewater Treatment Processes
abstract
This paper addresses multi-objective optimization problems using conflict-averse multi-objective extremum seeking (CAMOES) for unknown static mapping. As for the traditional multi-objective extremum seeking (MOES) methods, the initial conditions affect the optimal solution, which may result in over-optimizing some objectives. To overcome this issue, by minimizing the average loss function with the 2-norm of the worst individual objective, conflict-averse gradient estimator is investigated to determine the weighting factors and regularization parameter adaptively, extenuating the adverse impact of the prior parameters and initial conditions on the performance of extremum seeking. To cope with the integral windup problem for constrained inputs, penalty-function anti-windup mechanism is explored with smoothing saturation function for multiple-input multiple-output (MIMO) extremum seeking. The stability of CAMOES is theoretically analyzed. Some bi-objective optimization problems are conducted, including a numerical example and a practical problem for wastewater treatment processes (WWTPs). The results demonstrate that the proposed CAMOES exhibits competitive performance. Note to Practitioners—The initial conditions may lead to different Pareto solutions, which brings adverse impact on the multi-objective optimization performance in WWTPs. How to mitigate the influence of the initial parameters on the optimization performance and deal with the changing operation condition is the key to balancing the trade-off between the effluent quality (EQ) and energy cost (EC). This paper presents a control optimization scheme with intelligent modeling, extremum seeking optimization and the control part. The ErrCor-RBF neural network is built to predict the EC and EQ in WWTPs. For constraint conditions, the penalty-function anti-windup mechanism is exploited to avoid integral windup problems. The optimal set-points are dynamically adjusted by the proposed CAMOES to cope with the varying influent flow rate. The experimental results prove the proposed scheme has better operation and control performance. This paper proposes conflict-averse multi-objective extremum seeking (CAMOES), which is applied to the bi-objective optimization of WWTPs but can also be applied to other systems with multi-objective optimization tasks. In future research, we will address the time delay problem in WWTPs.
Minghui Chu 0001, Xin Huo, Kemao Ma
IEEE Trans Autom. Sci. Eng.2
2025 A Self-Learning Approach to Heterogeneous Multi-Robot Coalition Formation Under Uncertainty
abstract
Coalition structure is an effective cooperation architecture for task implementation in the field of multi-robot systems. Nevertheless, the presence of uncertainty inherently complicates the decision-making process for robots and may potentially result in suboptimal coordination. To tackle this challenge, this paper considers an uncertain multi-robot coalition formation scenario in which task information (the types of all tasks) is incompletely known to robots. Given the local beliefs of robots, the problem of multi-robot coalition formation under uncertainty is formulated as a coalition formation game. In this game, each robot is a rational and self-interested player and tends to join a coalition according to its preference. A polynomial-time coalition formation algorithm is proposed to identify the social agreement, i.e., Nash stable partition, among the robots. The convergence of the proposed algorithm is strictly guaranteed as long as the communication topology of the considered system is strongly connected. The coalition formation game is then extended to a dynamic game, and we propose a belief updating algorithm that enables robots to update their beliefs as the game is played repeatedly. Simulation results demonstrate the effectiveness of our proposed algorithms, and the robots will eventually learn the true type of each task.Note to Practitioners—The work reported in this article will be beneficial for deploying multi-robot systems to support cooperative surveillance applications. In these scenarios, robots can form stable coalitions to perform surveillance tasks, even when the task information is incompletely known due to sensor noise or limited sensor range. The problem of multi-robot coalition formation is known to be NP-hard, which becomes even more challenging when uncertainty is taken into account. This paper proposes game-based algorithms to solve the problem of multi-robot coalition formation under uncertainty. Some practical schemes are introduced as benchmarks to further illustrate the performance improvement brought by our proposed algorithms. The proposed algorithms are further evaluated through a real-world experiment, demonstrating their effectiveness for practical application.
Xin Huo, Hao Zhang 0008, Zhuping Wang, Chao Huang 0018, Huaicheng Yan 0001
IEEE Trans Autom. Sci. Eng.1
2025 Mecha: Multiview Enhanced Characteristics via Series Shuffling for Time Series Classification and Its Application to Turntable Circuit
abstract
In various electronic circuits with privacy and security constraints, only a single status monitoring signal is obtained, which renders domain knowledge inapplicable. To address this challenge, the data-driven univariate time series classification (TSC) method is a suitable and effective solution for diagnosing the system’s fault type. The current feature-based TSC method is interpretable by extracting a descriptive statistics-based feature collection, but its TSC performance is limited due to poor extensibility and feature redundancy. Therefore, in this paper, a novel and extensible feature-based TSC algorithm with ensemble structure and enhancement framework Multiview Enhanced Characteristics (Mecha) is proposed, which consists of three components. In the diverse feature extractor, the global and local patterns are enhanced via shuffling mapping with dilation and interleaving mechanisms, improving the feature diversity and expressiveness. In the ensemble feature selector, diverse and stable multiview features are adaptively generated by multiple filters and intersections based on the feature stability and diversity scores. In the heterogeneous ensemble classifier, the ridge regression with cross-validation and extremely randomized trees classifiers are integrated via hard voting to enhance classifier diversity. Finally, the state-of-the-art feature-based TSC performance and application effectiveness of the proposed Mecha without domain knowledge is verified on public UCR datasets and a practical turntable current dataset.
Changchun He, Xin Huo, Baohan Mi
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 A DTW-Gaussian Spatiotemporal Self-Attention Network and Its Application in Industrial Fault Diagnosis With Unequal-Length Sensor Data
abstract
Equipment operating in intermittent industrial processes generates complex system coupling and unequal-length sensor data, which are considered significant challenges for fault diagnosis in industrial production. To address these challenges, a generic fault diagnosis approach, the dynamic time warping Gaussian spatiotemporal self-attention network (DGSSANet), is proposed. dynamic time warping combined with Gaussian blur is employed to handle sensor signal complexity, transforming unequal-length data into warping matrices of equal dimensions. The inclusion of spatiotemporal self-attention, with the multiscale temporal attention module and multielement spatial self-attention module, enhances the ability of DGSSANet to capture spatial and temporal relationships. DGSSANet is validated through benchmark experiments on bearing datasets and applied to excavator cases, with its interpretability analyzed via visualization. The results demonstrate that DGSSANet addresses the challenges of fault diagnosis in intermittent production processes effectively.
Hewei Gao, Xin Huo, Yuchen Jiang 0001, Changchun He
IEEE Trans. Ind. Informatics2
2025 Multichannel-Based Multiview Shallow Fusion for Time Series Classification and Its Application in Fault Diagnosis
abstract
In the current time series classification (TSC) field, shallow concatenation, deep fusion, and hybrid ensemble multichannel frameworks (MCF) represented by convolution-based, deep learning, and hybrid methods have achieved competitive TSC performance. However, the massive kernels, deep fusion, and heterogeneous ensemble mechanisms, which are the core of the three frameworks, respectively, lead to overfitting risks. Therefore, in this article, a novel convolution-based TSC algorithm multichannel-based multiview shallow fusion (MC-MSF) within a new shallow fusion ensemble-based MCF is proposed. MC-MSF enhances feature diversity, quality, and classifier diversity while suppressing the overfitting risks via three shallow components. For feature diversity, the original series is mapped to the connected multichannel series spaces, and then diverse pooling features are extracted via a single-layer convolution with fewer kernels. For feature quality, the power of proportion of positive values (PPPV) features with adaptive powers are extracted based on alternating gradient descent, and the multiview shallow feature fusion is implemented to generate fused features. For classifier diversity, diverse linear classifiers are trained on the combined multiview feature vectors to ensemble homogeneously. The state-of-the-art TSC accuracy is achieved by MC-MSF via the sequential operation of three effective shallow components, as verified by comparative experiments on the public UCR and real excavator fault diagnosis application datasets.
Changchun He, Xin Huo, Yuchen Jiang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Tracking Differentiator-based Multiview Dilated Characteristics for Time Series Classification
abstract
In industrial production, obtaining the discrete status of equipment from massive time series data has become an urgent demand, which places higher requirements on the time series classification (TSC) algorithm. The feature-based TSC method achieves interpretable and potential classification capability by extracting meaningful descriptive statistical features. Limitations exist in current feature-based TSC algorithms, hindering the acquisition of diverse statistical features due to a lack of knowledge and insufficient research on processing redundant features, thereby limiting performance enhancements. Therefore, in this article, a novel feature-based TSC algorithm tracking differentiator-based multiview dilated characteristics (TD-MVDC) is proposed. The innovative introduction of a tracking differentiator combined with dilation mapping as a preprocessor into the feature-based TSC method is proposed to improve feature diversity efficiently. Ensemble feature selection based on filter feature selectors with different store ratios is designed to generate multiview features to enhance feature stability quickly. Linear classifiers and hard voting to fastly classify and integrate multiview features to increase classification performance robustly. Finally, comparative and ablative experiments are conducted between TD-MVDC and the representative feature-based TSC algorithms on the extensively compared UCR archive to verify the effectiveness of the proposed algorithm.
Changchun He, Xin Huo
INDIN2
2024 An Efficient Matching Game Approach to Association Formation in UAV-Enabled Hierarchical Distributed Learning
abstract
Distributed machine learning has emerged as a promising data processing technology for next-generation communication systems. It leverages the computational capabilities of local nodes to efficiently handle large datasets, creating highly accurate data-driven models for analysis and prediction purposes. However, the performance of distributed machine learning can be significantly hampered by communication bottlenecks and node dropouts. In this article, a novel unmanned aerial vehicle (UAV)-enabled hierarchical distributed learning architecture is proposed to support machine learning applications, e.g., regional monitoring. Multiple UAV receivers (URs) are introduced as wireless relays to improve the communication between the UAV transmitters (UTs) and the cloud server. Our objective is to identify the optimal UT-UR association to maximize the social welfare of the network, which is distinctly different from the existing works that focus on the unilateral profit-maximizing problem. We formulate a two-side many-to-one matching game to model the UT-UR association problem, and a two-phase many-to-one matching algorithm is designed to identify the stable matching. The validity of our proposed scheme is verified through in-depth numerical simulations.
Xin Huo, Hao Zhang 0008, Zhuping Wang, Huaicheng Yan 0001, Chun Liu 0003
IEEE Trans. Cybern.1
2023 ESO-Based Cyclic Iterative Learning Control for Continuous System with Varying Trials
abstract
This paper devises a cyclic iterative learning control (CILC) method based on extended state observer (ESO) to enhance the space-induced periodic disturbances rejection performance as well as uncertainties for continuous system with varying trials. Conventional ILC is based on the strict premise that the system executes the same length of trial for the subsequent trial, which is not appropriate in the majority of occasions. Besides, there exist aperiodic disturbances, uncertainties, and spatial-varying properties. To make the problem tractable, a data-driven ESO-based cyclic iterative learning control approach is proposed and synthesized. The cyclic memory whose data are stored with respect to the angular position rather than time are compressed and the filter compensation of ESO-based CILC is investigated. The reduction of error are guaranteed by iteration under rotating circles. Numerical simulations and comparisons are provided to demonstrate the efficacy and effectiveness of the proposed method.
Aijing Wu, Xin Huo
IECON2
2023 FT-FVC: fast transformation-based feature vector concatenation for time series classification
Changchun He, Xin Huo, Hewei Gao
Appl. Intell.2
2023 PID-Like Model Free Adaptive Control With Discrete Extended State Observer and Its Application on an Unmanned Helicopter
abstract
Aiming at the problem that how to design a controller without using model information for an unmanned helicopter (UH), a novel data-driven method based on model free adaptive control (MFAC) is proposed in this article. A new form of dynamic linearization (DL) equation composed of pseudo-partial derivative (PPD) and external disturbance is built in order to reduce the influence of disturbance on PPD. Since there are internal and external unknowns included in the DL equation, a nominal DL equation without external disturbance item is introduced to make the estimation of the two items decoupled. The estimation algorithm of PPD is designed based on the nominal DL equation, and a discrete extended state observer is applied specifically into the estimation of external disturbance based on the original DL equation. Besides, with the nominal DL equation, a modified MFAC scheme called PID-like MFAC, which considers the historical input and output (I/O) data is given to make a better performance and a more convenient application compared with the typical MFAC scheme. Furthermore, the stability of the control scheme is proved. Finally, comparative simulations are carried out to verify the effectiveness of the proposed control scheme. An experiment based on UH also verifies its practical performance and the theoretical findings.
Xin Huo, Kemao Ma, Ruihang Ji
IEEE Trans. Ind. Informatics2
2023 Improved Gradient Estimation for Fast Extremum Seeking: A Parametric Proportional-Integral Observer-Based Approach
abstract
In this article, a complete parametric proportional-integral observer (PPIO)-based approach is proposed to improve the performance of gradient estimation for the fast extremum seeking (ES) scheme acting on a Hammerstein plant. Unlike the prevailing gradient estimation approach of the fast ES which uses a Luenberger observer without an explicit way to obtain the observer gains, a systemic complete PPIO is established based on a complete parametric solution to a type of generalized Sylvester matrix equations. The proposed PPIO presents complete parameterization of all the gain matrices as well as the left eigenvectors in terms of some sets of design parameters that represent the degrees of design freedom. Then, the gradient estimator is constructed by multiplying the states of PPIO and the demodulation signal. Moreover, a synthetic objective function, which includes weighted performance indices of the transient error and the steady-state accuracy, is formulated. The performance of the gradient estimator is improved by minimizing the synthetic objective function through adjusting the degrees of freedom of the PPIO, and all explicit values of the parametric gain matrices are derived with the adjusted degrees of freedom. In turn, a faster and more accurate gradient estimation scheme can be obtained and significantly improve the convergence of the closed-loop system. Besides, the proposed PPIO-based estimator has excellent performance under the noise condition. Simulation examples and an application to the lean-burn combustion system are used to illustrate the effectiveness of the proposed gradient estimation scheme.
Weizhen Liu, Xin Huo, Kemao Ma, Weichao Sun
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Adaptive-Critic Design for Decentralized Event-Triggered Control of Constrained Nonlinear Interconnected Systems Within an Identifier-Critic Framework
abstract
This article studies the decentralized event-triggered control problem for a class of constrained nonlinear interconnected systems. By assigning a specific cost function for each constrained auxiliary subsystem, the original control problem is equivalently transformed into finding a series of optimal control policies updating in an aperiodic manner, and these optimal event-triggered control laws together constitute the desired decentralized controller. It is strictly proven that the system under consideration is stable in the sense of uniformly ultimate boundedness provided by the solutions of event-triggered Hamilton-Jacobi-Bellman equations. Different from the traditional adaptive critic design methods, we present an identifier-critic network architecture to relax the restrictions posed on the system dynamics, and the actor network commonly used to approximate the optimal control law is circumvented. The weights in the critic network are tuned on the basis of the gradient descent approach as well as the historical data, such that the persistence of excitation condition is no longer needed. The validity of our control scheme is demonstrated through a simulation example.
Xin Huo, Hamid Reza Karimi, Xudong Zhao 0001, Bohui Wang, Guangdeng Zong
IEEE Trans. Cybern.1
2022 Performance Improvement in Fast Extremum Seeking With Adaptive Phase Compensator and High-Gain Optimizer
abstract
In this article, an adaptive phase compensator and a high-gain optimizer are proposed for the fast extremum-seeking (ES) scheme acting on a Hammerstein plant. Unlike the widely applied three-time scale tuning, a time scale-independent tuning is achieved by restricting the plant to have a Hammerstein structure and elevating the adaptation of the control input which makes it as fast as the dither signal. Albeit the existed fast ES schemes utilize a high-frequency sinusoidal dither in order to achieve fast minimization of the plant’s static nonlinearity, the phase shift of the plant could not be neglected as the frequency is high and the ES may fail to reach the extreme values if the shift is large enough. Thus, an effective adaptive frequency varying phase compensation method is introduced to deal with the phase shift under the assumptions in this article. Moreover, a high-gain optimizer is presented to remove the restriction that the adaptation gain has to be small and could still guarantee the stability properties, which makes the proposed scheme more practical. The effectiveness of the proposed scheme is illustrated by numerical simulations.
Weizhen Liu, Xin Huo, Kemao Ma
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Single-network ADP for solving optimal event-triggered tracking control problem of completely unknown nonlinear systems
abstract
In this paper, we propose an optimal event-triggered tracking control scheme for completely unknown nonlinear systems under the adaptive dynamic programming (ADP) framework. A data-driven model based on recurrent neural networks (RNNs) is first constructed to model the system uncertainties including the drift dynamics and the input gain matrix, and the modeling error caused by NN approximation is well eliminated through adding a compensation term in the data-driven model such that the model state can asymptotically track the system state. Apart from the traditional construction of optimal tracking controllers, in this paper, an augmented system is developed and a discounted performance function is considered to achieve the optimality. By employing the Bellman optimal principle, an event-triggered tracking Hamilton–Jacobi–Bellman (HJB) equation is then formulated. The approximate solution of the HJB equation can be obtained by virtue of a critic NN, which significantly simplifies the implementation architecture of ADP. Both the historical state data and the current state data are incorporated into the updating of the weight vector in the critic NN, in this circumstance, the persistence of excitation assumption is not needed anymore. It is strictly proven via Lyapunov stability theory that the tracking error state and the critic NN weight are uniformly ultimately bounded. Simulation results examine the validity of the design scheme.
Ning Xu 0013, Ben Niu 0003, Huanqing Wang 0001, Xin Huo, Xudong Zhao 0001
Int. J. Intell. Syst.4
2021 Small-Gain Technique-Based Adaptive Neural Output-Feedback Fault-Tolerant Control of Switched Nonlinear Systems With Unmodeled Dynamics
abstract
In this article, the issue of adaptive neural fault-tolerant control (FTC) is addressed for a class of uncertain switched nonstrict-feedback nonlinear systems with unmodeled dynamics and unmeasurable states. In such a system, the uncertain nonlinear parts are identified by radial basis function (RBF) neural networks (NNs). Also, with the help of the structural characteristics of RBF NNs, the violation between the nontsrict-feedback form and backstepping method is tackled. Then, based on the small-gain technique, input-to-state practical stability (ISpS) theory, and common Lyapunov function (CLF) approach, an adaptive fault-tolerant tracking controller with only three adaptive laws is developed by designing an observer. It is shown that the designed controller can ensure that all the closed-loop signals are bounded under arbitrary switching, while the tracking error can converge to a small area of the origin. Finally, two simulation examples are provided to demonstrate the feasibility of the suggested control approach.
Li Ma 0008, Ning Xu 0013, Xudong Zhao 0001, Guangdeng Zong, Xin Huo
IEEE Trans. Syst. Man Cybern. Syst.5
2019 Adaptive neural control for switched nonlinear systems with unknown backlash-like hysteresis and output dead-zone
Li Ma 0008, Xin Huo, Xudong Zhao 0001, Ben Niu 0003, Guangdeng Zong
Neurocomputing2
2019 Observer-based adaptive fuzzy tracking control of MIMO switched nonlinear systems preceded by unknown backlash-like hysteresis
Xin Huo, Li Ma 0008, Xudong Zhao 0001, Ben Niu 0003, Guangdeng Zong
Inf. Sci.1
2015 A Pattern Specification and Optimizations Framework for Accelerating Scientific Computations on Heterogeneous Clusters
abstract
Clusters with accelerators at each node have emerged as the dominant high-end architecture in recent years. Such systems can be extremely hard to program because of the underlying heterogeneity and the need for exploiting parallelism at multiple levels. Thus, easing parallel programming today requires not only high-level programming models, but ones from which hybrid parallelism can be extracted. In this paper, we focus on the following question: "can simple APIs be developed for several classes of popular scientific applications, to ease application development and yet maintain parallel efficiency, on clusters with accelerators?". We approach this problem by individually considering popular patterns that arise in scientific computations. By developing APIsfor generalized reductions, irregular reductions, and stencil computations, we show that several complex scientific applications can be supported. We enable compact specification of these applications (40% of the code size of MPI), while also enabling parallelization across nodes and devices within a node, and with work distribution across CPU and GPU cores. We enable a number of optimizations that are normally implemented by hand by scientific programmers. We compare well against existing MPI applications while scaling across nodes, and against handwritten CUDA applications for executions on a single GPU, and yet can scale by using all parallelism simultaneously. On a cluster with 64GPUs, we achieve speedups between 600 and 1800 over sequential(single CPU core) versions.
Linchuan Chen, Xin Huo, Gagan Agrawal
IPDPS2
2015 Efficient and Simplified Parallel Graph Processing over CPU and MIC
abstract
Intel Xeon Phi (MIC architecture) is a relatively new accelerator chip, which combines large-scale shared memory parallelism with wide SIMD lanes. Mapping applications on anode with such an architecture to achieve high parallel efficiency's a major challenge. In this paper, we focus on developing system for heterogeneous graph processing, which is able to utilize both a many-core Xeon Phi and a multi-core CPU ozone node. We propose a simple programming API with unintuitive interface for expressing SIMD parallelism. We develop efficient techniques for supporting our high-level API, focusing on exploiting wide SIMD lanes, massive number of cores, and partitioning of the work across CPU and accelerator, while handling the irregularity of graph applications. The components of our runtime system include a condensed static memory buffer, which supports efficient message insertion and SIMD message reduction while keeping memory requirements low, and specifically formic, a pipelining scheme for efficient message generation by avoiding frequent locking operations. Besides, a hybrid graph partitioning module is able to effectively partition the workload between the CPU and the MIC, ensuring balanced workload and low communication overhead. The main observations from our experimental evaluation using five popular applications are: formic executions, pipelining scheme is up to 3.36x faster than naive approach using locking based message generation, and the speedup over OpenMP ranges from 1.17 to 4.15. Heterogeneous-MIC execution achieves a speedup of up to 1.41 over the better of the CPU-only and MIC-only executions.
Linchuan Chen, Xin Huo, Bin Ren 0002, Surabhi Jain, Gagan Agrawal
IPDPS2
2014 A programming system for xeon phis with runtime SIMD parallelization
abstract
The Intel Xeon Phi offers a promising solution to coprocessing, since it is based on the popular x86 instruction set. However, to fully utilize its potential, applications must be vectorized to leverage the wide SIMD lanes, in addition to effective large-scale shared memory parallelism. Compared to the SIMT execution model on GPGPUs with CUDA or OpenCL, SIMD parallelism with a SSE-like instruction set imposes many restrictions, and has generally not benefitted applications involving branches, irregular accesses, or even reductions in the past. In this paper, we consider the problem of accelerating applications involving different communication patterns on Xeon Phis, with an emphasis on effectively using available SIMD parallelism. We offer an API for both shared memory and SIMD parallelization, and demonstrate its implementation. We use implementations of overloaded functions as a mechanism for providing SIMD code, which is assisted by runtime data reordering and our methods to effectively manage control flow. Our extensive evaluation with 6 popular applications shows large gains over the SIMD parallelization achieved by the production (ICC) compiler, and we even outperform OpenMP for MIMD parallelism.
Xin Huo, Bin Ren 0002, Gagan Agrawal
ICS1
2014 A cooperation earth observation model of SAR satellite and optical remote sensing satellite
abstract
A cooperation earth observation model constituted by two SAR satellites and one optical remote sensing satellite is proposed. In this model, the optical remote sensing satellite orbit plane locates in the middle of two SAR satellites, and two SAR satellites observe the same ground area from two sides of the optical satellite respectively. However, because of the influence of earth curvature, not only orbits of three satellites must meet certain conditions to realize the cooperation goal that is to observing the same ground area simultaneously, but also SAR view-angle must be adjusted with optical satellite movement in its orbit. So here the cooperation model of two SAR satellites and one optical remote sensing satellite is discusses and the explicit expressions of simultaneous observable conditions are deduced. Then, the calculation method of SAR view-angles is presented. At last, computer simulation results are employed to confirm the mathematical analysis.
Hui Liu 0030, Lei Zhang 0104, Xin Huo, Jinying Luan, Kexin Lan, Changfeng Jing, Wei Li 0085
IGARSS4
2013 Efficient scheduling of recursive control flow on GPUs
abstract
Graphics processing units (GPUs) have rapidly emerged as a very significant player in high performance computing. Single instruction multiple thread (SIMT) pipelines are typically used in GPUs to exploit parallelism and maximize performance. Although support for unstructured control flow has been included in GPUs, efficiently managing thread divergence for arbitrary parallel programs remains a critical challenge. In this paper, we focus on the problem of supporting recursion in modern GPUs. We design and comparatively evaluate various algorithms to manage thread divergence encountered in recursive programs. The results improve upon traditional post-dominator based reconvergence mechanisms designed to handle thread divergence due to control flow within a procedure.
Xin Huo, Sriram Krishnamoorthy, Gagan Agrawal
ICS1
2012 Accelerating MapReduce on a coupled CPU-GPU architecture
abstract
The work presented here is driven by two observations. First, heterogeneous architectures that integrate a CPU and a GPU on the same chip are emerging, and hold much promise for supporting power-efficient and scalable high performance computing. Second, MapReduce has emerged as a suitable framework for simplified parallel application development for many classes of applications, including data mining and machine learning applications that benefit from accelerators. This paper focuses on the challenge of scaling a MapReduce application using the CPU and GPU together in an integrated architecture. We develop different methods for dividing the work, which are the map-dividing scheme, where map tasks are divided between both devices, and the pipelining scheme, which pipelines the map and the reduce stages on different devices. We develop dynamic work distribution schemes for both the approaches. To achieve high load balance while keeping scheduling costs low, we use a runtime tuning method to adjust task block sizes for the map-dividing scheme. Our implementation of MapReduce is based on a continuous reduction method, which avoids the memory overheads of storing key-value pairs. We have evaluated the different design decisions using 5 popular MapReduce applications. For 4 of the applications, our system achieves 1.21 to 2.1 speedup over the better of the CPU-only and GPU-only versions. The speedups over a single CPU core execution range from 3.25 to 28.68. The runtime tuning method we have developed achieves very low load imbalance, while keeping scheduling overheads low. Though our current work is specific to MapReduce, many underlying ideas are also applicable towards intra-node acceleration of other applications on integrated CPU-GPU nodes.
Linchuan Chen, Xin Huo, Gagan Agrawal
SC2
2011 Porting irregular reductions on heterogeneous CPU-GPU configurations
abstract
Heterogeneous architectures are playing a significant role in High Performance Computing (HPC) today, with the popularity of accelerators like the GPUs, and the new trend towards the integration of CPUs and GPUs. Developing applications that can effectively use these architectures is a major challenge. In this paper, we focus on one of the dwarfs in the Berkeley view on parallel computing, which are the irregular applications arising from unstructured grids. We consider the problem of executing these reductions on heterogeneous architectures comprising a multi-core CPU and a GPU. We have developed a Multi-level Partitioning Framework, which has the following features: (1) it supports GPU execution of irregular reductions even when the dataset size exceeds the size of the device memory, (2) it can enable pipelining of partitioning performed on the CPU, and the computations on the GPU, and (3) it supports dynamic distribution of work between the multi-core CPU and the GPU. Our extensive evaluation using two different irregular applications demonstrates the effectiveness of our approach.
Xin Huo, Vignesh T. Ravi, Gagan Agrawal
HiPC1
2011 An execution strategy and optimized runtime support for parallelizing irregular reductions on modern GPUs
abstract
GPUs have rapidly emerged as a very significant player in high performance computing. However, despite the popularity of CUDA, there are significant challenges in porting different classes of HPC applications on modern GPUs. This paper focuses on the challenges of implementing irregular applications arising from unstructured grids on modern NVIDIA GPUs. Considering the importance of irregular reductions in scientific and engineering codes, substantial effort was made in developing compiler and runtime support for parallelization or optimization of these codes in the previous two decades, with different efforts targeting distributed memory machines, distributed shared memory machines, shared memory machines, or cache performance improvement on uniprocessor machines. However, there have not been any systematic studies on parallelizing these applications on modern GPUs. There are at least two significant challenges associated with porting this class of applications on modern GPUs. The first is related to correct and efficient parallelization while using a large number of threads. The second challenge is effective use of shared memory. Since data accesses cannot be determined statically, runtime partitioning methods are needed for effectively using the shared memory. This paper describes an execution methodology that can address the above two challenges. We have also developed optimized runtime modules to support our execution methodology. Our approach and runtime methods have been extensively evaluated using two indirection array based applications.
Xin Huo, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal
ICS1
2010 Approaches for parallelizing reductions on modern GPUs
abstract
GPU hardware and software has been evolving rapidly. CUDA versions 1.1 and higher started supporting atomic operations on device memory, and CUDA versions 1.2 and higher started supporting atomic operations on shared memory. This paper focuses on parallelizing applications involving reductions on GPUs. Prior to the availability of support for locking, these applications could only be parallelized using full replication, i.e., by creating a copy of the reduction object for each thread. However, CUDA 1.1 (1.2) onwards, use of atomic operations (on shared memory) is another option, though some effort is still required in supporting locking on floating point numbers and for supporting coarse-grained locking. Based on the tradeoffs between locking and full replication, we also introduce a hybrid approach, in which a group of threads use atomic operations to update one copy of the reduction object. Using three data mining algorithms that follow the reduction structure - k-means clustering, Principal Component Analysis (PCA) and k-nearest neighbor search (kNN), we evaluate the relative performance of these three approaches. We show how the relative performance of these techniques can vary depending upon the application and its parameters. The hybrid approach we have introduced clearly outperforms other approaches in several cases.
Xin Huo, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal
HiPC1
2010 Kinematics analysis of a novel all-attitude flight simulator
Kemao Ma, Xin Huo, Yu Yao 0004
Sci. China Inf. Sci.4