Chuanfu Xu

dblp:43/826 · DBLP profile ↗
← Back
45ranked-venue papers
3as first author
31since 2021 · last 2027
0000-0002-4876-2368ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2027 Collaborative-adversarial jailbreaking: A propagation-aware attack framework for multi-agent code generation systems
Zhaoyang Qu, Mingyang Geng, Yunxin Mao, Shanzhi Gu, Chuanfu Xu, Haotian Wang 0001
Neural Networks5
2026 Failure Localization in Multi-Agent Code Generation via Knowledge-Guided and Transferable Reasoning
abstract
Recent advances in multi-agent Large Language Model-based code generation enable collaborative software development through role-specialized agents. However, failure localization of code generation remains challenging due to inter-agent dependencies and solution-path multiplicity. Consequently, existing prompting-based localization methods exhibit vulnerability towards semantically valid but non-canonical strategies. To address this, we propose FLKR (Failure Localization via Knowledge-guided Reasoning), an self-supervised framework that combines behavior encoding, knowledge-strategy alignment, and consistency scoring for solution-path invariant localization. To evaluate, we also introduce COFL (Code Oriented Failure Localization), the first expert-annotated benchmark for fine-grained failure localization. Experiments show FLKR outperforms state-of-the-art prompting-based baselines by up to 14 points in Fault Localization Accuracy and 45 points in Top-1 accuracy, with strong performance in divergent, real-world, and refinement-critical cases. Such results demonstrate that our proposed FLKR generalizes well to real-world software development scenarios and opens up a new direction for failure-aware refinement recommendation by providing precise and interpretable responsibility signals.
Mingyang Geng, Shanzhi Gu, Chuanfu Xu, Zhaoyang Qu, Haotian Wang 0001
AAAI4
2026 DFA-SpTRSV: A Depth-First Asynchronous Algorithm for Sparse Triangular Solver
Zhimeng Han, Chuanfu Xu, Haozhong Qiu, Yue Ding 0001, Yonggang Che
IPDPS2
2026 BAAS: A Bidirectional Aggregation and Affinity-Aware Scheduling Framework for Parallelizing Sparse Matrix Computations
Qingyang Zhang 0009, Chuanfu Xu, Chun Huang 0006, Zhimeng Han, Jie Liu 0002
IPDPS3
2026 A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue Solvers
abstract
Krylov subspace methods are widely used in scientific computing to solve large sparse linear systems and eigenvalue problems. Their performance bottleneck is often dominated by high-order matrix-power kernels (MPK), especially in polynomial preconditioners that must scale to millions or billions of variables. We present Diagonal Block MPK (DBMPK), a lightweight and parallel-friendly optimization that partitions the input matrix into diagonal blocks and off-diagonal regions. This design enables efficient intra-block data reuse and eliminates inter-block dependencies. It improves cache locality, parallelism, and reduces preprocessing overheads, compared to existing techniques. Our evaluation on x86 and Arm HPC platforms shows that DBMPK improves MPK performance by 26.6%-38.4%. When applied to polynomial preconditioners for linear systems and eigenvalue problems, it achieves consistent end-to-end speedups of 18.6%-34.0%, including in weak scaling tests on 128 nodes, demonstrating strong scalability and practical impact.
Xiaojian Yang, Yuhui Ni, Shengguo Li, Dezun Dong, Chuanfu Xu, Haipeng Jia, Jie Liu 0002
PPoPP6
2026 Mitigating sensitive information leakage in LLMs4Code through machine unlearning
Shanzhi Gu, Zhaoyang Qu, Ruotong Geng, Mingyang Geng, Shangwen Wang, Chuanfu Xu, Haotian Wang 0001, Dezun Dong
Neural Networks6
2026 A Memory-Aware Sparse Matrix-Matrix Multiplication on Multicore Architectures
abstract
Sparse matrix–matrix multiplication (SpMM) is a fundamental operation in scientific computing with broad applications across numerous domains. Tiling is a key optimization technique for improving data locality and is widely adopted in high-performance computing. However, the irregular data access patterns inherent to SpMM make it challenging to exploit tiling effectively for data reuse. In this article, we propose MaSpMM , a memory-aware SpMM framework that integrates cache-aware tiling with a segment-oriented data layout. MaSpMM stores matrices as continuous segments to enhance data locality within each tile. Moreover, since many sparse matrices in real-world applications exhibit symmetry, we further develop MaSpMM-Sym, an extension that recursively partitions symmetric matrices to eliminate write conflicts and further improve locality. To adapt to diverse scenarios, we finally introduce MaSpMM-Adap, which adaptively selects the most suitable approach for each input matrix. Comprehensive evaluations on both x86 and ARM CPUs demonstrate that MaSpMM-Adap achieves average speedups of up to 1.86× over Intel oneMKL, 1.84× over ASpT, and 1.75× over J-Stream.
Deshun Bi, Shengguo Li, Haozhong Qiu, Chuanfu Xu, Xiaojian Yang, Dezun Dong, Tiaojie Xiao, Jie Liu 0002
ACM Trans. Archit. Code Optim.4
2026 Towards Efficient Symmetric Sparse Matrix-Vector Multiplication on Multi-Cores
abstract
Exploiting matrix symmetry to halve memory footprint offers a substantial opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently, i.e., either are non-scalable or yield poor performance for large high-bandwidth irregular matrices. This article extends DCS-SpMV , a D ivide-and- C onquer (DC) based shared-memory implementation of S ymmetric SpMV. The key idea of DCS-SpMV is to recursively divide and reorder the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. The DC approach naturally transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution DCH-SpMV that executes these two parts using DCS-SpMV and the standard SpMV, respectively. We also develop a machine learning model for DCH-SpMV to predict the optimal number of DC recursions on a given matrix and architecture. In this work, we further optimize the hybrid DC implementation by reducing data conflicts before the DC preprocessing. First, we present a conflict-pruning strategy to decouple certain highly dense columns or rows from the conflict graph of a symmetric matrix. Second, we implement a heuristic to adaptively select the lower or upper triangular part of a symmetric matrix, leading to fewer data conflicts. Our optimizations not only facilitate the DC preprocessing, but also improve the performance of DCH-SpMV. We evaluate our work on both x86 and ARM multi-core CPUs using 298 symmetric sparse matrices from the SuiteSparse Matrix Collection. Our new optimizations improve the performance of previous version [ 42 ] by up to 4.89×, demonstrating significant speedup over the state-of-the-art approaches including the vendor-tuned Intel oneMKL library.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che
ACM Trans. Archit. Code Optim.2
2026 "Stones From Other Hills Can Polish Jade": Zero-Shot Anomaly Synthesis via Cross-Domain Anomaly Injection
abstract
Industrial image anomaly detection (IAD) is a pivotal topic with huge value. Due to the nature of anomalies, real anomalies in a specific modern industrial domain (i.e., domain-specific anomalies) are usually too rare to collect, which severely hinders IAD. Thus, zero-shot anomaly synthesis (ZSAS), which synthesizes pseudo anomaly images without any domain-specific anomaly, emerges as a vital technique for IAD. However, existing solutions are either unable to synthesize authentic pseudo anomalies, or require cumbersome training. Thus, we focus on ZSAS and propose a brand-new paradigm that can realize both authentic and training-free ZSAS. It is based on a chronically-ignored fact: Although domain-specific anomalies are rare, real anomalies from other domains (i.e., cross-domain anomalies) are actually abundant and directly applicable to ZSAS. Specifically, our new ZSAS paradigm makes three-fold contributions: First, we propose a novel method named Cross-domain Anomaly Injection (CAI), which directly exploits cross-domain anomalies to enable highly authentic ZSAS in a training-free manner. Second, to supply CAI with sufficient cross-domain anomalies, we build the first Domain-agnostic Anomaly Dataset (DAAD) within our best knowledge, which provides ZSAS with abundant real anomaly patterns. Third, we propose a CAI-guided Diffusion Mechanism, which can further break the quantity limit of real anomalies and enable unlimited anomaly synthesis. Our head-to-head comparison with existing ZSAS solutions justifies the superior performance of our paradigm for IAD and demonstrates it as an effective and pragmatic ZSAS solution.
Siqi Wang 0001, Yuanze Hu, Xinwang Liu 0002, Siwei Wang 0001, Guangpu Wang, Chuanfu Xu, Jie Liu 0002, Ping Chen 0004
IEEE Trans. Image Process.6
2025 Me-MPK: Accelerating Krylov Subspace Solvers via Memory-efficient Matrix-Power Kernel
abstract
This paper focuses on optimizing the Matrix-Power Kernel (MPK), which relies on a series of Sparse Matrix-Vector multiplications (SpMVs) using the same sparse matrix. MPK is a crucial component of Krylov subspace methods for solving large sparse linear systems in various fields, including circuit simulations. MPK offers a potential for matrix reuse in cache, which can accelerate memory-bound sparse solvers. Additionally, many sparse matrices encountered in applications are symmetric, allowing us to reduce the memory footprint for SpMVs by half. However, reusing the matrix introduces data dependencies between subsequent SpMVs, and symmetric SpMVs can result in data conflicts during shared-memory parallelization. Previous research has often focused on either matrix reuse or symmetry, failing to leverage both aspects effectively. This paper proposes a unified, memory-efficient approach called Me-MPK that takes advantage of both cache reuse and matrix symmetry for MPK on shared-memory multi-core systems. We first introduce a unified dependency graph for a sparse matrix, which represents all potential data dependencies and conflicts. Next, we perform architecture-aware recursive partitioning on this graph to create subgraphs and formulate a separating subgraph that decouples all dependencies and conflicts among the subgraphs. These independent subgraphs are then scheduled for parallel execution of SpMV or symmetric SpMV in a specified order to optimize cache reuse. We apply Me-MPK in two s-Step Krylov subspace solvers, and our evaluations show that Me-MPK significantly outperforms the current state-of-the-art solutions, delivering an average speedup of up to 2.00X and 1.86X on X86 and ARM CPUs, respectively. As a result, we achieve overall speedup in the sparse solvers of up to $\mathbf{1. 6 5 X}$ and $\mathbf{1. 5 8 X}$.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Shengguo Li, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002
DAC2
2025 CRAMG: A Communication-Reduced Algebraic Multigrid Method
abstract
Algebraic multigrid (AMG) is widely used to accelerate largescale sparse linear solvers.In distributed environments, neighboring communication overhead in AMG significantly impacts overall solution time.We propose Communication-Reduced Algebraic Multigrid (CRAMG) methods to minimize inter-process data exchange and message count by fusing interpolation/restriction operators with residual computations.This reduces communication frequency from four per level to as few as two.Experiments show up to 45% reduction in data exchange and 35% fewer messages.Performance evaluations on an Intel platform demonstrate significant improvements
Xiaojian Yang, Yunqing Huang, Dezun Dong, Chuanfu Xu, Jie Liu 0002, Xiaoqiang Yue, Shengguo Li
ICS5
2025 DAS-ILU: A Distributed Asynchronous Parallel ILU Factorization Based on Domain Decomposition
abstract
This paper presents DAS-ILU, a Distributed Asynchronous parallel Incomplete LU factorization method based on domain decomposition. DAS-ILU partitions the computational domain into independently processed interior nodes and asynchronously updated separator nodes, thereby reducing cross-processor dependencies and halving the separator size compared to conventional methods. To further improve performance, it employs optimized data exchange patterns to minimize communication overhead and extends support to block-structured sparse matrices via exact block inversions. Comprehensive evaluations on a range of problem types—including structural mechanics, computational fluid dynamics, and reservoir simulation demonstrate the superior performance of DAS-ILU. Compared to state-of-the-art ILU implementations, DAS-ILU achieves solve time speedups of up to 2.07 × over Chow-Patel’s fine-grained parallel ILU and up to 4.11 × over HYPRE’s ILU. Moreover, DAS-ILU exhibits strong robustness when applied to challenging nonsymmetric and indefinite systems.
Shengguo Li, Xiaojian Yang, Yunqing Huang, Chuanfu Xu, Dezun Dong, Jianchun Wang, Jie Liu 0002
SC6
2025 DCSolver: Accelerating Sparse Iterative Solvers via Divide-and-Conquer on GPUs
abstract
Sparse iterative solvers are commonly used in various fields. However, certain essential kernels of these solvers, such as sparse triangular solves (SpTRSV), present significant challenges for efficient parallelization due to data dependencies . Previous methods, like level-scheduling or multi-coloring, typically involve creating a Task Dependency Graph (TDG) to represent data dependencies and identify independent sets from the TDG for parallel execution. However, these approaches often result in limited parallelism with substantial synchronization overheads or negatively impact the solver convergence rate. This article introduces DCSolver , a Divide-and-Conquer (DC) framework designed to efficiently parallelize sparse solvers with data dependencies on GPUs. To achieve this, we break down the solver TDG into independent subgraphs, allowing us to exploit both coarse-grained and fine-grained parallelism. To efficiently allocate GPU threads for subgraphs with varying degrees of parallelism, we have developed an adaptive in-warp scheduling strategy. Additionally, we propose a hybrid parallelization scheme in DCSolver, which involves employing different parallel approaches for different DC recursions to achieve a more optimal balance between parallelism and convergence for solvers. To evaluate the effectiveness of DCSolver, we apply it to two preconditioned Krylov subspace solvers and an unstructured mesh Computational Fluid Dynamics (CFD) solver. Our results show that when compared with the state-of-the-art methods, DCSolver accelerates the time-to-solution of solvers by an average speedup of up to 26.19X.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002
ACM Trans. Archit. Code Optim.2
2025 Automatic GPU memory access optimization for AoSoA-based application in OP2 framework
Zongjing Chen, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
J. Supercomput.4
2024 MARO: Enabling Full MPI Automatic Refactoring in DSL-Based Programming Framework
Zongjing Chen, Yonggang Che, Chuanfu Xu
ICA3PP (4)4
2024 Towards Scalable Unstructured Mesh Computations on Shared Memory Many-Cores
abstract
Due to data conflicts or data dependences, exploiting shared memory parallelism on unstructured mesh applications is highly challenging. The prior approaches are neither general nor scalable on emerging many-core processors. This paper presents a general and scalable shared memory approach for unstructured mesh computations. We recursively divide and reorder an unstructured mesh to construct a task dependency tree (TDT), where massive parallelism is exposed and data conflicts as well as data dependences are respected. We propose two recursion strategies to support popular programming models on both CPUs and GPUs for TDT. We evaluate our approach by applying it to an industrial unstructured Computational Fluid Dynamics (CFD) software. Experimental results show that our approach significantly outperforms the prior shared memory approaches, delivering up to 8.1× performance improvement over the engineer-tuned implementations.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Yonggang Che, Shizhao Chen, Jie Liu 0002
PPoPP2
2024 A Conflict-aware Divide-and-Conquer Algorithm for Symmetric Sparse Matrix-Vector Multiplication
abstract
Exploiting matrix symmetry to halve memory footprint offers an opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently. This paper proposes DCS-SpMV, a Divide-and-Conquer (DC) algorithm for efficient Symmetric SpMV. The key idea is to recursively divide the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. Our DC algorithm transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution that executes these two parts using DCS-SpMV and traditional SpMV respectively. We develop a machine learning model to predict an optimal hybrid implementation for a given matrix and architecture. We evaluate our work on both X86 and ARM CPUs, demonstrating significant performance improvement over the state-of-the-art.
Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Shizhao Chen, Yonggang Che, Jie Liu 0002
SC2
2024 Extending OP2 framework to support portable parallel programming of complex applications
Zongjing Chen, Kangjin Huang, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
CCF Trans. High Perform. Comput.4
2024 Discriminative object tracking by domain contrast
Huayue Cai, Xiang Zhang 0008, Long Lan, Changcheng Xiao, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
Comput. Vis. Image Underst.5
2024 Improving CUDA performance of an unstructured high-order CFD application under OP2 framework
Kangjin Huang, Yonggang Che, Chuanfu Xu, Jian Zhang 0115
J. Supercomput.3
2023 Developing a proxy application for an industrial unstructured CFD software: preliminary results
abstract
As programming models and architectures evolve in the exa-scale era, porting large-scale HPC applications are becoming increasingly difficult and expensive. In HPC community, mini-apps are often developed to mimic real-world applications, and it offers an easy way to benchmark new HPC platforms. In this paper, we design and implement a mini-app MiniFS as a proxy for an industry-level unstructured Computational Fluid Dynamics (CFD) software FlowStar. The main purpose of MiniFS is to evaluate different shared memory approaches on emerging multi/many-core architectures, because data conflicts and data dependencies in unstructured CFD pose tough challenges for shared memory parallelization. Results show that our mini-app can represent the performance characteristic of the original application. However, the existing approaches are unscalable on modern multi/many-cores. It is imperative to develop novel scalable shared memory approaches for unstructured CFD.
Chuanfu Xu, Jian Zhang 0115, Liang Deng, Haozhong Qiu, Weixi Dai, Yongzhen Lin, Yue Ding 0001, Yonggang Che
ICPADS2
2023 wrBench: Comparing Cache Architectures and Coherency Protocols on ARMv8 Many-Core Systems
Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001
J. Comput. Sci. Technol.4
2022 Deep Anomaly Discovery from Unlabeled Videos via Normality Advantage and Self-Paced Refinement
abstract
While classic video anomaly detection (VAD) requires labeled normal videos for training, emerging unsupervised VAD (UVAD) aims to discover anomalies directly from fully unlabeled videos. However, existing UVAD methods still rely on shallow models to perform detection or initialization, and they are evidently inferior to classic VAD methods. This paper proposes a full deep neural network (DNN) based solution that can realize highly effective UVAD. First, we, for the first time, point out that deep reconstruction can be surprisingly effective for UVAD, which inspires us to unveil a property named “normality advantage”, i.e., normal events will enjoy lower reconstruction loss when DNN learns to reconstruct unlabeled videos. With this property, we propose Localization based Reconstruction (LBR) as a strong UVAD baseline and a solid foundation of our solution. Second, we propose a novel self-paced refinement (SPR) scheme, which is synthesized into LBR to conduct UVAD. Unlike ordinary self-paced learning that injects more samples in an easy-to-hard manner, the proposed SPR scheme gradually drops samples so that suspicious anomalies can be removed from the learning process. In this way, SPR consolidates normality advantage and enables better UVAD in a more proactive way. Finally, we further design a variant solution that explicitly takes the motion cues into account. The solution evidently enhances the UVAD performance, and it sometimes even surpasses the best classic VAD methods. Experiments show that our solution not only significantly outperforms existing UVAD methods by a wide margin (5% to 9% AUROC), but also enables UVAD to catch up with the mainstream performance of classic VAD.
Siqi Wang 0001, Zhiping Cai, Xinwang Liu 0002, Chuanfu Xu, Chengkun Wu
CVPR5
2022 An Unsupervised Short- and Long-Term Mask Representation for Multivariate Time Series Anomaly Detection
Qiucheng Miao, Chuanfu Xu, Chengkun Wu
ICONIP (6)2
2022 Parallelizing and Balancing Coupled DSMC/PIC for Large-scale Particle Simulations
abstract
In high-performance and parallel computing, an important application class is particle simulation. Due to massive particle migration among distributed simulation workers across simulation iterations, achieving balanced runtime work distribution is vital for accelerating large-scale realistic particle simulations. This paper proposes a novel approach to enable dynamic load balance for distributed numerical particle simulations, specifically targeting the latest coupled DSMC/PI C method. Unlike prior work, our approach adopts a dual, nested unstructured grid organization to facilitate coupled DSMC/PIC computation and runtime grid distribution. Our implementation leverages both centralized and distributed communication strategies to dynamically migrate particles among arbitrary parallel processes. It then employs a load balancer - driven by a carefully designed analytical model and a grid remapping mechanism - to dynamically redistribute the simulation workloads among parallel simulation workers. By constantly monitoring and redis-tributing the simulation work across workers, our approach can adapt to the change of particle distribution across simulation iterations, avoiding a few workers becoming the performance bottleneck of the entire simulation process. We integrate our techniques into a coupled DSMC/PIC solver and apply them to simulate the plasma plume with hydrogen atoms and ions. Experimental results show that our approach can scale well up to 1500+ processes with billions of particles, exhibiting the state-of-the-art parallel simulation scalability and efficiency for plasma plume simulation.
Haozhong Qiu, Chuanfu Xu, Dali Li, Zheng Wang 0001
IPDPS2
2022 FlowDNN: a physics-informed deep neural network for fast and accurate flow prediction
abstract
for flow-related design optimization problems, e.g., aircraft and automobile aerodynamic design, computational fluid dynamics (CFD) simulations are commonly used to predict flow fields and analyze performance. While important, CFD simulations are a resource-demanding and time-consuming iterative process. The expensive simulation overhead limits the opportunities for large design space exploration and prevents interactive design. In this paper, we propose FlowDNN, a novel deep neural network (DNN) to efficiently learn flow representations from CFD results. FlowDNN saves computational time by directly predicting the expected flow fields based on given flow conditions and geometry shapes. FlowDNN is the first DNN that incorporates the underlying physical conservation laws of fluid dynamics with a carefully designed attention mechanism for steady flow prediction. This approach not only improves the prediction accuracy, but also preserves the physical consistency of the predicted flow fields, which is essential for CFD. Various metrics are derived to evaluate FlowDNN with respect to the whole flow fields or regions of interest (RoIs) (e.g., boundary layers where flow quantities change rapidly). Experiments show that FlowDNN significantly outperforms alternative methods with faster inference and more accurate results. It speeds up a graphics processing unit (GPU) accelerated CFD solver by more than 14 000×, while keeping the prediction error under 5%.
Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Siqi Wang 0001, Shizhao Chen, Jianbin Fang, Zheng Wang 0001
Frontiers Inf. Technol. Electron. Eng.3
2022 Label Propagated Nonnegative Matrix Factorization for Clustering
abstract
Semi-supervised learning (SSL) that utilizes plenty of unlabeled examples to boost the performance of learning from limited labeled examples is a powerful learning paradigm with widely real-world applications such as information retrieval and document clustering. Label propagation (LP) is a popular SSL method which propagates labels through the dataset along high density areas defined by unlabeled examples, but it is fragile to bridge examples. Semi-supervised K-Means uses labeled examples to initialize clustering centers to separate different examples, however, semi-supervised K-Means fails in the situation of imbalanced issues, that is, the example size of each class varies significantly. This paper proposes a novel label propagated nonnegative matrix factorization method (LPNMF) to handle clean labeled but biased data and its extension LPNMF-E to handle noisy labeled data based on the framework of NMF. LPNMF decomposes the whole dataset into the product of a basis matrix and a coefficient matrix. To propagate labels to unlabeled examples, LPNMF regards the class indicators of labeled examples as their coefficients and iteratively updates both basis matrix and coefficients of unlabeled examples. LPNMF absorbs the merits from both semi-supervised K-Means and label propagation to handle their respective shortages. Specifically, on the one hand, LPNMF learns representative clustering centers based on the distribution of the dataset, similar to semi-supervised K-means, and thus is robust to the bridge examples. On the other hand, LPNMF pushes labels according to the affinity between examples, similar to label propagation, and thus relieves the biased problem. Moreover, we introduce a LPNMF extension to handle the noisy label case. LPNMF-E relaxes the constraint of labeled examples. Since the label of each labeled example also obtains label information from the global distribution of the whole dataset and local manifold of its neighbors, LPNMF-E outputs reliable class indicators even if a portion of examples are incorrectly labeled. Theoretical analyses for the generalization ability of our proposed models are also provided. Experimental results on both clean and noisy labeled datasets confirm the effectiveness of LPNMF and LPNMF-E compared with both LP and the representative semi-supervised K-Means algorithms.
Long Lan, Tongliang Liu, Xiang Zhang 0008, Chuanfu Xu, Zhigang Luo
IEEE Trans. Knowl. Data Eng.4
2021 Optimizing Barrier Synchronization on ARMv8 Many-Core Architectures
abstract
Synchronization operations are commonly seen in OpenMP programs where a parallel construct often works with an explicit or implicit barrier operation. While OpenMP synchronization has been extensively studied on the traditional x86 CPU architectures, there is little work on understanding OpenMP barrier synchronization operations on ARMv8 high-performance many-cores. This paper presents the first comprehensive performance study on OpenMP barrier implementations on emerging ARMvS-based many-cores. We evaluate seven representative barrier algorithms on three distinct ARMv8 architectures: Phytium 2000+, ThunderX2, and Kunpeng920. We empirically show that the existing synchronization implementations exhibit poor scalability on ARMv8 architectures compared to the x86 counterpart. We then propose various optimization strategies for improving these widely used synchronization algorithms on each platform. We showcase that our optimizations yield 12.6x performance improvement over the GCC implementation and 4.7x improvement over the LLVM implementation, translating to 1.6x improvement over the state-of-the-art best-performing algorithm. We share our experience and practical insights on optimizing OpenMP synchronization operations on emerging ARMv8 multi-core CPU architectures.
Wanrong Gao, Jianbin Fang, Chun Huang 0006, Chuanfu Xu, Zheng Wang 0001
CLUSTER4
2021 Joint motion context and clip augmentation for spatio-temporal action detection
abstract
This paper endeavors to leverage spatio-temporal visual cues to improve video-based action detection. As a result, a NOn-Local Action detector based on anchor-free called NOLA is proposed, which is built off a recent moving center detector (MOC) and further extends it by efficiently aggregating long-range spatio-temporal information. In detail, a significantly efficient spatio-temporal motion-aware non-local block is explored to provide global motion contexts for the entire predictive branches of MOC. This byproduct can make the large batch samples run on a resource limited device. Besides, a light-weighted data augmentation method termed clip augmentation designed for video-based tasks is proposed, which serves to improve the generalization ability of the detector with economical scale-and-addition operation. NOLA works with two above schemes in real-time as well. Experiments on two benchmark datasets show that NOLA significantly exceeds MOC. Compared to other existing methods,, NOLA reaches the state-of-the-art, in terms of video-level mean of average precision (video mAP).
Xurui Ma, Xiang Zhang 0008, Chengkun Wu, Chuanfu Xu, Jie Liu 0002, Zhigang Luo
ICMV4
2021 AA-LSTM: An Adversarial Autoencoder Joint Model for Prediction of Equipment Remaining Useful Life
Chengkun Wu, Chuanfu Xu, Zhenghua Wang
PAKDD (1)3
2021 Learning deep discriminative embeddings via joint rescaled features and log-probability centers
Huayue Cai, Xiang Zhang 0008, Long Lan, Guohua Dong, Chuanfu Xu, Xinwang Liu 0002, Zhigang Luo
Pattern Recognit.5
2020 FlowGAN: A Conditional Generative Adversarial Network for Flow Prediction in Various Conditions
abstract
Many flow-related design optimization problems like aircraft and automobile aerodynamic design are solved via computational fluid dynamics (CFD) simulations. However, CFD simulations are known to be resource-demanding and time-consuming. Deep learning (DL) is emerging as a viable means to accelerate CFD simulations by directly predicting the outcomes of multiple simulation iterations. While promising, existing DL-based models have to be re-trained whenever the flow condition changes, which incurs significant training overhead for real-life scenarios with a wide range of flow conditions. This paper presents FLOWGAN, a novel conditional generative adversarial network for accurate prediction of flow fields in various conditions. FlowGAN is designed to directly obtain the generation of solutions to flow fields in various conditions based on observations rather than re-training. As FlowGAN does not rely on knowledge of the underlying governing equations, it can quickly adapt to various flow conditions and avoid the need for expensive re-training. We evaluate FlowGAN by applying it to scenarios of simulating both the whole flow field and selected regions of interest (RoI). Compared to the state-of-the-art DL based methods, FlowGAN significantly reduces the prediction errors by 2.27% while exhibiting a better generalization ability.
Donglin Chen, Xiang Gao 0020, Chuanfu Xu, Shizhao Chen, Jianbin Fang, Zhenghua Wang, Zheng Wang 0001
ICTAI3
2020 Cloze Test Helps: Effective Video Anomaly Detection via Learning to Complete Video Events
abstract
As a vital topic in media content interpretation, video anomaly detection (VAD) has made fruitful progress via deep neural network (DNN). However, existing methods usually follow a reconstruction or frame prediction routine. They suffer from two gaps: (1) They cannot localize video activities in a both precise and comprehensive manner. (2) They lack sufficient abilities to utilize high-level semantics and temporal context information. Inspired by frequently-used cloze test in language study, we propose a brand-new VAD solution named Video Event Completion (VEC) to bridge gaps above: First, we propose a novel pipeline to achieve both precise and comprehensive enclosure of video activities. Appearance and motion are exploited as mutually complimentary cues to localize regions of interest (RoIs). A normalized spatio-temporal cube (STC) is built from each RoI as a video event, which lays the foundation of VEC and serves as a basic processing unit. Second, we encourage DNN to capture high-level semantics by solving a visual cloze test. To build such a visual cloze test, a certain patch of STC is erased to yield an incomplete event (IE). The DNN learns to restore the original video event from the IE by inferring the missing patch. Third, to incorporate richer motion dynamics, another DNN is trained to infer erased patches' optical flow. Finally, two ensemble strategies using different types of IE and modalities are proposed to boost VAD performance, so as to fully exploit the temporal context and modality information for VAD. VEC can consistently outperform state-of-the-art methods by a notable margin (typically 1.5%-5% AUROC) on commonly-used VAD benchmarks. Our codes and results can be verified at github.com/yuguangnudt/VEC_VAD
Siqi Wang 0001, Zhiping Cai, En Zhu, Chuanfu Xu, Jianping Yin, Marius Kloft
ACM Multimedia5
2019 Effective End-to-end Unsupervised Outlier Detection via Inlier Priority of Discriminative Network
abstract
Despite the wide success of deep neural networks (DNN), little progress has been made on end-to-end unsupervised outlier detection (UOD) from high dimensional data like raw images. In this paper, we propose a framework named E^3Outlier, which can perform UOD in a both effective and end-to-end manner: First, instead of the commonly-used autoencoders in previous end-to-end UOD methods, E^3Outlier for the first time leverages a discriminative DNN for better representation learning, by using surrogate supervision to create multiple pseudo classes from original unlabelled data. Next, unlike classic UOD that utilizes data characteristics like density or proximity, we exploit a novel property named inlier priority to enable end-to-end UOD by discriminative DNN. We demonstrate theoretically and empirically that the intrinsic class imbalance of inliers/outliers will make the network prioritize minimizing inliers' loss when inliers/outliers are indiscriminately fed into the network for training, which enables us to differentiate outliers directly from DNN's outputs. Finally, based on inlier priority, we propose the negative entropy based score as a simple and effective outlierness measure. Extensive evaluations show that E^3Outlier significantly advances UOD performance by up to 30% AUROC against state-of-the-art counterparts, especially on relatively difficult benchmarks.
Siqi Wang 0001, Yijie Zeng, Xinwang Liu 0002, En Zhu, Jianping Yin, Chuanfu Xu, Marius Kloft
NeurIPS6
2019 Collaborating CPUs and MICs for Large-Scale LBM Multiphase Flow Simulations
Chuanfu Xu, Dali Li, Yonggang Che, Zhenghua Wang
NPC1
2018 Flexible ranking extreme learning machine based on matrix-centering transformation
abstract
Existing ranking ELM algorithms bias to imbalanced queries since they equally treat each pairwise error. In this study we propose a flexible ranking ELM method based on matrix-centering transformation to replace the traditional graph Laplacian matrix based methods. Specifically, we introduce a useful query-level normalized loss function and enforce the matrix-centering transformation to it to avoid training a bias model. Fortunately, by this setting, we can also greatly simplify the learning process of ELM because of the symmetry and idempotence of the centering matrix. Based on the proposed framework, three different ranking ELM variants are implemented: (a) a regularized ranking ELM model; (b) an enhanced incremental ranking ELM model; and (c) an online sequential ranking ELM model. Experimental results demonstrate that our proposed ranking ELM algorithms can obtain comparable or better performances than the state-of-the-art ranking algorithms.
Shizhao Chen, Kai Chen 0020, Chuanfu Xu, Long Lan
IJCNN3
2018 Petascale scramjet combustion simulation on the Tianhe-2 heterogeneous supercomputer
Yonggang Che, Meifang Yang, Chuanfu Xu, Yutong Lu
Parallel Comput.3
2017 Performance modeling and optimization of parallel LU-SGS on many-core processors for 3D high-order CFD simulations
Dali Li, Chuanfu Xu, Xiang Gao 0020, Xiaogang Deng
J. Supercomput.2
2016 Parallelizing and optimizing large-scale 3D multi-phase flow simulations on the Tianhe-2 supercomputer
abstract
Summary The lattice Boltzmann method (LBM) is a widely used computational fluid dynamics method for flow problems with complex geometries and various boundary conditions. Large‐scale LBM simulations with increasing resolution and extending temporal range require massive high‐performance computing (HPC) resources, thus motivating us to port it onto modern many‐core heterogeneous supercomputers like Tianhe‐2. Although many‐core accelerators such as graphics processing unit and Intel MIC have a dramatic advantage of floating‐point performance and power efficiency over CPUs, they also pose a tough challenge to parallelize and optimize computational fluid dynamics codes on large‐scale heterogeneous system. In this paper, we parallelize and optimize the open source 3D multi‐phase LBM code openlbmflow on the Intel Xeon Phi (MIC) accelerated Tianhe‐2 supercomputer using a hybrid and heterogeneous MPI+OpenMP+Offload+single instruction, mulitple data (SIMD) programming model. With cache blocking and SIMD‐friendly data structure transformation, we dramatically improve the SIMD and cache efficiency for the single‐thread performance on both CPU and Phi, achieving a speedup of 7.9X and 8.8X, respectively, compared with the baseline code. To collaborate CPUs and Phi processors efficiently, we propose a load‐balance scheme to distribute workloads among intra‐node two CPUs and three Phi processors and use an asynchronous model to overlap the collaborative computation and communication as far as possible. The collaborative approach with two CPUs and three Phi processors improves the performance by around 3.2X compared with the CPU‐only approach. Scalability tests show that openlbmflow can achieve a parallel efficiency of about 60% on 2048 nodes, with about 400K cores in total. To the best of our knowledge, this is the largest scale CPU‐MIC collaborative LBM simulation for 3D multi‐phase flow problems. Copyright © 2015 John Wiley & Sons, Ltd.
Dali Li, Chuanfu Xu, Yongxian Wang, Zhifang Song, Xiang Gao 0020, Xiaogang Deng
Concurr. Comput. Pract. Exp.2
2015 Realistic Performance Characterization of CFD Applications on Intel Many Integrated Core Architecture
abstract
This paper studies the performance characteristics of computational fluid dynamics (CFD) applications on Intel Many Integrated Core (MIC) architecture. Three CFD applications, BT-MZ, LM3D and HOSTA, are evaluated on Intel Knights Corner (KNC) coprocessor, the first public MIC product. The results show that the pure OpenMP scalability of these applications is not sufficient to utilize the potential of a KNC coprocessor. While utilizing the hybrid MPI/OpenMP programming model helps to improve the parallel scalability, the maximum parallel speedup relative to a single thread is still not satisfactory. The OpenCL version of BT-MZ performs better than the OpenMP version but is not comparable to the MPI version and the hybrid MPI/OpenMP version. At the micro-architecture level, while the three CFD applications achieve reasonable instruction execution rates and L1 data cache hit rates, use a large percent of vector instructions, they have low arithmetic density, incur very high branch misprediction rates and do not utilize the Vector Processing Unit efficiently. As a result, they achieve very low single thread floating-point efficiency. For these applications to attain competitive performance on the MIC architecture as on the Xeon processors, both the parallel scalability and the single thread performance should be improved, which is a difficult task.
Yonggang Che, Chuanfu Xu, Jianbin Fang, Yongxian Wang, Zhenghua Wang
Comput. J.2
2014 Balancing CPU-GPU Collaborative High-Order CFD Simulations on the Tianhe-1A Supercomputer
abstract
HOSTA is an in-house high-order CFD software that can simulate complex flows with complex geometries. Large scale high-order CFD simulations using HOSTA require massive HPC resources, thus motivating us to port it onto modern GPU accelerated supercomputers like Tianhe-1A. To achieve a greater speedup and fully tap the potential of Tianhe-1A, we collaborate CPU and GPU for HOSTA instead of using a naive GPU-only approach. We present multiple novel techniques to balance the loads between the store-poor GPU and the store-rich CPU, and overlap the collaborative computation and communication as far as possible. Taking CPU and GPU load balance into account, we improve the maximum simulation problem size per Tianhe-1A node for HOSTA by 2.3X, meanwhile the collaborative approach can improve the performance by around 45% compared to the GPU-only approach. Scalability tests show that HOSTA can achieve a parallel efficiency of above 60% on 1024 Tianhe-1A nodes. With our method, we have successfully simulated China's large civil airplane configuration C919 containing 150M grid cells. To our best knowledge, this is the first paper that reports a CPUGPU collaborative high-order accurate aerodynamic simulation result with such a complex grid geometry.
Chuanfu Xu, Lilun Zhang, Xiaogang Deng, Jianbin Fang, Guangxue Wang, Yonggang Che, Yongxian Wang, Wei Liu 0013
IPDPS1
2014 Test-driving Intel Xeon Phi
abstract
Based on Intel's Many Integrated Core (MIC) architecture, Intel Xeon Phi is one of the few truly many-core CPUs - featuring around 60 fairly powerful cores, two levels of caches, and graphic memory, all interconnected by a very fast ring. Given its promised ease-of-use and high performance, we took Xeon Phi out for a test drive. In this paper, we present this experience at two different levels: (1) the microbenchmark level, where we stress "each nut and bolt" of Phi in the lab, and (2) the application level, where we study Phi's performance response in a real-life environment. At the microbenchmarking level, we show the high performance of five components of the architecture, focusing on their maximum achieved performance and the prerequisites to achieve it. Next, we choose a medical imaging application (Leukocyte Tracking) as a case study. We observed that it is rather easy to get functional code and start benchmarking, but the first performance numbers can be far from satisfying. Our experience indicates that a simple data structure and massive parallelism are critical for Xeon Phi to perform well. When compiler-driven parallelization and/or vectorization fails, programming Xeon Phi for performance can become very challenging.
Jianbin Fang, Henk J. Sips, Lilun Zhang, Chuanfu Xu, Yonggang Che, Ana Lucia Varbanescu
ICPE4
2014 Microarchitectural performance comparison of Intel Knights Corner and Intel Sandy Bridge with CFD applications
Yonggang Che, Lilun Zhang, Yongxian Wang, Chuanfu Xu, Wei Liu 0013, Zhenghua Wang
J. Supercomput.4
2009 MPTD: A Scalable and Flexible Performance Prediction Framework for Parallel Systems
Chuanfu Xu, Yonggang Che, Zhenghua Wang
APPT1
2008 Analyzing the Efficiency and Bottleneck of Scientific Programs on Imagine Stream Processor by Simulation
abstract
Imagine stream processor has shown high performance and efficiency for media applications. Its potential for scientific applications is of great interest to the high performance computing community. This paper investigates this subject from a new angle. It roughly classifies the scientific programs into three classes based on their computation to memory access ratios. For each class, typical programs are programmed with StreamC/KernelC stream language and simulated based on the cycle-accurate simulator of Imagine. In-depth analysis is carried out for the performance data, with special attentions on the performance bottlenecks. The performance data obtained on Imagine are compared against data on two general-purpose x86 processors. The results show that programs with no DRAM accesses attain high floating point performance and efficiencies on Imagine. These programs' performance is only restricted by limited ILP (Instruction-Level Parallelism) and load imbalance across ALUs. Programs with computation to memory operation ratios O(n) attain absolute floating point performance on Imagine comparable to that obtained on general-purpose processors, but their floating-point efficiencies are not satisfactory. It is essential to optimize these programs for high SRF (Stream Register File) and LRF (Local Register File) reuse and high ILP on Imagine. Programs with lower computation to memory operation ratios attain much lower floating-point performance and efficiencies on Imagine, compared to those obtained on x86 processors.
Yonggang Che, Chuanfu Xu, Zhenghua Wang
ISPA2