Xinhai Chen 0001

dblp:231/1664-1 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
36since 2021 · last 2026
0000-0002-2931-4893ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 16 since 2021Systems, architecture and hardware · 16 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Learning to Generate Structured Meshes with In-Context: Toward Generalization in Mesh Generation
abstract
Structured mesh generation serves as a crucial preprocessing step in numerical simulations and can be formulated as a mapping problem from geometry to structured mesh. Existing approaches typically establish an isolated mapping for each geometry. This geometry-specific paradigm fails to capture and leverage commonalities across geometries, inevitably requiring recomputation or costly retraining for new geometries. To overcome this limitation, we propose ICL-Mesh, a meta-learning framework based on in-context learning (ICL) for structured mesh generation. It treats learning one mapping as one task and trains a single neural network to extract commonalities across tasks and learn from in-context examples within each task, enabling rapid generalization to unseen tasks without parameter updates. Experimental results demonstrate that ICL-Mesh effectively generalizes to diverse geometries with only a few context examples, and even without examples. It also exhibits robustness to in-context example order sensitivity and can be extended to various mesh generation scenarios, including mesh refinement and coarsening.
Xinhai Chen 0001, Jiaming Peng, Jie Liu 0002
AAAI2
2026 Block-Aware Adaptive State Management for Optimistic Parallel Discrete Event Simulation
Gencheng Liu, Chuhe Hong, Xinhai Chen 0001, Qingyang Zhang 0009, Jie Liu 0002
ICS7
2026 THAC: Unlocking Performance in Parallel HPC Applications via UQ-Aware Automated Approximation
abstract
While approximate computing offers a promising paradigm for performance gains, significant technical challenges obstruct its practical adoption in High Performance Computing (HPC). The process of manually identifying approximable regions, quantifying the risk of cascading errors, and navigating the vast combinatorial tuning space is intricate and error-prone. To overcome these obstacles, we present Tianhe Approximate Computing (THAC), a framework that establishes a systematic, automated methodology. THAC implements a principled, three-stage workflow: (1) a hybrid AI-driven synthesizer that automatically identifies potential approximation candidates; (2) a principled Uncertainty Quantification screen that rigorously validates and prunes high-risk options through a principled risk-aware screening stage; and (3) a novel hierarchical Bayesian optimizer that efficiently navigates the search space. Although demonstrated on OpenMP as a representative case study, the framework is designed with extensibility for generic parallel patterns. Evaluation on a diverse suite of scientific benchmarks (including computational fluid dynamics and hydrodynamics) demonstrates the effectiveness of THAC, achieving a geometric mean speedup of 2.1× under strict quality budgets. These results confirm THAC as a robust solution that makes approximate computing in HPC practical and reliable.
Zhenhao Zhao, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002, Binglin Wang
ICS4
2026 Accelerating High-Frequency Electromagnetic Scattering Prediction with KNN-Augmented Radial Basis Function Networks
Chao Li 0002, Xinhai Chen 0001, Tiaojie Xiao, Jie Liu 0002
IPDPS4
2026 Physics-informed residual learning with low-rank adaptation for unsupervised mesh generation
Jiaming Peng, Xinhai Chen 0001, Qingling Wang, Zhiquan Lai, Dongsheng Li 0001, Jie Liu 0002
Comput. Aided Geom. Des.2
2026 Survey of storage systems in high performance computing
abstract
Abstract As high performance computing (HPC) moves towards exascale, storage systems face core challenges such as data flooding, bandwidth bottlenecks, mixed load coordination, and performance cost balancing. This article systematically reviews the cutting-edge technologies of high performance storage systems, covering four aspects: storage architecture, hardware, software, and networking. At the architecture level, storage computing separation, distributed and hierarchical architectures decouple computing and storage resources, and optimize latency and scalability through high-speed networks. Typical cases include supercomputer systems such as Frontier and Fugaku. In terms of hardware, persistent memory, all flash array, and integrated storage and computing chips significantly improve throughput and reduce latency, while ZNS SSD and QLC technology optimize cost and lifespan. At the software level, distributed parallel file systems respond to massive small files and high concurrency access through burst buffering technology. In network communication, low latency protocols such as Slingshot, InfiniBand, and RoCE support TB level bandwidth, while CXL technology promotes storage resource pooling. In the future, photon interconnection, AI native architecture, and green energy-saving technologies will further promote the development of high performance storage towards efficiency and intelligence, to support ZB level storage requirements in scenarios such as Exascale computing and AI training.
Gen Zhang, Zhenlong Song, Xinhai Chen 0001, Yong Dong
CCF Trans. High Perform. Comput.4
2026 Cascaded spectral operator transformer with mixture-of-experts for urban wind field prediction
Jie Li 0002, Xinhai Chen 0001, Yonggang Che, Qingyang Zhang 0009
Eng. Appl. Artif. Intell.4
2026 Flux-conserved physics-informed neural networks for electromagnetic scattering computation
Chenyu Peng, Tiaojie Xiao, Sifan Wang, Xinhai Chen 0001, Chunye Gong
Eng. Appl. Artif. Intell.5
2026 LDNO: A low-power dynamic neural operator inspired by liquid state machines for solving partial differential equations
Chengxue Huang, Jie Liu 0002, Qingyang Zhang 0009, Xinhai Chen 0001, Bo Yang 0023
Neurocomputing6
2026 PI-MeshONet: A generalizable and self-supervised method for structured mesh generation
Xinhai Chen 0001, Jiaming Peng, Jie Liu 0002
Knowl. Based Syst.2
2026 GNNRL-smoothing: A prior-free reinforcement learning model for mesh optimization
Xinhai Chen 0001, Chunye Gong, Bo Yang 0023, Liang Deng, Yufei Pang, Xiang Zhang 0008, Jie Liu 0002
Neural Networks2
2026 MeshONet: A generalizable and efficient operator learning method for structured mesh generation
Xinhai Chen 0001, Jiaming Peng, Jie Liu 0002
Neural Networks2
2026 ST-FlowNet: A lightweight framework for long-term spatio-temporal flow field prediction
Qisong Xiao, Xinhai Chen 0001, Haijian Yang, Chunye Gong, Jie Liu 0002
Neural Networks2
2026 FastCC: A System-Algorithm Co-Design for Connected Components Computation on Large Power-Law Graphs
abstract
Connected Components (CC) computation is a fundamental graph analytics kernel. While BFS-sampling has emerged as the state-of-the-art approach for power-law graphs, its performance in existing implementations is severely limited by inheriting unnecessary BFS semantics. The core issue is a fundamental mismatch: BFS requires strict level-synchronization to compute shortest paths, while CC only needs eventual label consistency without ordering constraints. This semantic mismatch manifests as three critical bottlenecks: (i) severe load imbalance from vertex-centric task allocation, which fails to distribute the massive workload of high-degree hub vertices; (ii) redundant writes from dynamic push/pull mode switching, which necessitates costly frontier reconstruction; and (iii) redundant synchronization and computation from enforcing BFS’s strict ordering guarantees, which are superfluous for CC computation. We introduce FastCC , a lightweight multiprocess system-algorithm co-designed solution that breaks this semantic mismatch. The cornerstone of our approach is the strategic decision to fix the highest-degree vertex as the BFS root, creating a predictable computation topology. This enables three synergistic innovations: (1) hybrid task partitioning that employs edge-centric allocation in the critical first iteration to eliminate load imbalance at its source; (2) predictable mode switching that leverages the deterministic computation graph to bypass expensive frontier reconstruction; and (3) lightweight, custom synchronization primitives that relax BFS’s strict ordering to match CC’s eventual consistency requirements, while allowing earlier label propagation within the same iteration to reduce the overall computational workload. Extensive evaluation on large-scale real-world and synthetic power-law graphs demonstrates that FastCC achieves significant performance improvements, with speedups of 10.6-55.5× faster (average: 37.8×) over state-of-the-art CC implementations including ConnectIt and vGraph. FastCC also reduces peak memory footprint by up to 2.87× and exhibits superior, more predictable scalability. The practical efficacy of our approach is validated by its deployment as the core engine in a top-ranked GreenGraph500 solution.
Menghan Jia, Yongquan Fu, Yiming Zhang 0003, Xinhai Chen 0001, Dongsheng Li 0001
ACM Trans. Archit. Code Optim.4
2025 PEINR: A Physics-enhanced Implicit Neural Representation for High-Fidelity Flow Field Reconstruction
abstract
Implicit neural representation (INR) has now been thrust into the limelight with its flexibility in high-fidelity flow field reconstruction tasks. However, the lack of standard benchmarking datasets and the grid independence assumption for INR-based methods hinder progress and adoption in real-world simulation scenarios. Moreover, naive adoptions of existing INR frameworks suffer from limited accuracy in capturing fine-scale structures and spatiotemporal dynamics. Tacking these issues, we first introduce HFR-Beach, a 5.4 TB public large-scale CFD dataset with 33,600 unsteady 2D and 3D vector fields for reconstructing high-fidelity flow fields. We further present PEINR, a physics-enhanced INR framework, to enrich the flow fields by concurrently enhancing numerical-precision and grid-resolution. Specifically, PEINR is mainly composed of physical encoding and transformer-based spatiotemporal fuser (TransSTF). Physical encoding decouples temporal and spatial components, employing Gaussian coordinate encoding and localized encoding techniques to capture the nonlinear characteristics of spatiotemporal dynamics and the stencil discretization of spatial dimensions, respectively. TransSTF fuses both spatial and temporal information via transformer for capturing long-range temporal dependencies. Qualitative and quantitative experiments and demonstrate that PEINR outperforms state-of-the-art INR-based methods in reconstruction quality.
Liming Shen, Liang Deng, Chongke Bi, Xinhai Chen 0001, Yueqing Wang, Jie Liu 0002
ICML5
2025 VES: Vectorized Sparse General Matrix-Matrix Multiplication on Multi-Core DSPs
abstract
The Sparse General Matrix-Matrix Multiplication (SpGEMM) is widely used in a variety of applications. However, research on optimizing SpGEMM for high-performance digital signal processors (DSPs) has been limited. We present VES, a method to accelerate SpGEMM on multi-core DSPs, using the FT-M7032 platform as a case study. Based on the ESC algorithm, VES enhances computational efficiency through vectorized expansion operations, a double-buffering strategy, and an optimized vectorized sorting method. We provide an in-depth analysis of the bottlenecks in vectorized sorting and introduce an efficient vectorized reduction method that significantly improves instruction-level pipeline throughput. Experimental results show that VES outperforms existing methods HASH, ESC, and SPA by an average of 1.12×, 1.90×, and 22.34× on 1,931 sparse matrices, with maximum speedups of 8.69×, 65.9×, and 821.6×, respectively.
Chuhe Hong, Gencheng Liu, Qingyang Zhang 0009, Xinhai Chen 0001, Jie Liu 0002
ICPP6
2025 IA-Chol: Input-Aware Cholesky Decomposition on CPU and GPU
Jixiao Deng, Lin Chen 0028, Tun Li 0002, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002
ICS6
2025 Spatio-Temporal Prediction of Two-Phase Flow in Heterogeneous Reservoirs Using Deep Neural Networks
abstract
Accurate spatio-temporal prediction of oil-water two-phase flow under heterogeneous reservoir conditions is crucial for optimizing oil recovery strategies. Pioneering studies have begun to explore intelligent methods to reduce the computational cost in traditional numerical simulation. However, existing intelligent methods suffer from error accumulation problems, leading to inaccurate evolution results over time. In this paper, we propose CASTNet, a novel spatio-temporal model for predicting the evolution of oil-water two-phase flow. The model designs an encoder-decoder structure to capture and simulate the spatio-temporal variations of oil-water two-phase flow. We design an attention-enhanced ConvLSTM module to strengthen the identification of heterogeneous regions, tackling the issue of failing to detect spatial information in heterogeneous reservoirs. Experiments demonstrate that CASTNet achieves accurate and efficient spatio-temporal prediction of oil-water two-phase flow. It outperforms existing intelligent methods by maintaining a maximum absolute error of 0.08 during long-term simulation, while it achieves one order-of-magnitude improvement in computational efficiency compared to traditional numerical methods.
Xinhai Chen 0001, Qisong Xiao, Haijian Yang, Jiali Tu, Jie Liu 0002
IJCNN2
2025 Informative Discrimination Network for Efficient Single Image Super-Resolution
abstract
Deploying convolutional neural networks on low-resource mobile devices for single image super-resolution (SISR) faces the issue of how to balance the parameter amount and performance. The default solution is simultaneously condensing both hierarchical representation and attention features into their respective light proxies. The insight underlying this solution lies in the fact that features are redundant since the super-resolution needs plenty of similar pixels. This work takes it to the next step from the viewpoint of informativeness and discrimination. In detail, we propose an informative disrcimination network (IDNet) for SISR. For informativeness, a multi-scale residual block (MRB) is explored to capture informative spatial details via the scale-in-scale structure. It mines rich intra-layer spatial details based on inter-layer ones of the default hierarchical representation. However, it also incurs feature redundancy. Though attention serves to reduce this redundancy, feature discrimination and pixel-wise structural preservation cannot be guaranteed. Here spatial discrimination attention behaves like the biased discriminant classifier to induce spatial discrimination, while the nuclear-norm regularization recovers the image low-rank structure to reduce artifacts or noises. Importantly, no extra network weights are introduced for model efficiency. Experiments show that IDNet delivers sound performance with fewer parameters, as compared to its cousins.
Yuzheng Tu, Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Bo Yang 0023, Xiang Gao 0020, Xiang Zhang 0008
IJCNN2
2025 Dual-Spectral Neural Operator for Solving Partial Differential Equations
abstract
Solving partial differential equations is a key focus of research in scientific computing. Traditional neural operator methods often face challenges in capturing both global features and local features simultaneously, limiting their effectiveness in complex scenarios. In this paper, we propose the Dual-Spectral Neural Operator (DSNO), an innovative method that efficiently addresses these challenges by integrating both Fourier and wavelet transforms. DSNO employs a dynamic adjustment mechanism, incorporating a hybrid multiplicative-additive fusion strategy to optimally balance the contributions of Fourier and wavelet transforms, thereby enhancing its capacity to capture both global and local features. Extensive experiments demonstrate that our method outperforms existing neural operator methods, achieving state-of-the-art predictive performance on a range of challenging tasks.
Xinhai Chen 0001, Jie Liu 0002
IJCNN2
2025 An Efficient Adaptive Dual-Threshold Svm Based on Heterogeneous Collaboration
abstract
Support Vector Machine (SVM) is highly effective at processing high-dimensional, nonlinear data. However, more than 90 % of the training time is spent on kernel matrix calculations, and existing approaches encounter challenges in adapting to heterogeneous architectures. This paper presents an adaptive dual-threshold method leveraging heterogeneous collaboration to accelerate kernel matrix computations. We also analyze the impact of the working set size on training time and accuracy in ThunderSVM to minimize training time. Tasks with distinct characteristics are allocated to appropriate computing cores through heterogeneous collaboration, with dynamic load balancing via adaptive dual thresholds. On a CPU-DSP heterogeneous platform, our method delivers an average speedup of$5.52 \times$compared to the optimal CPU-only implementation.
Chuhe Hong, Gencheng Liu, Xinhai Chen 0001, Jie Liu 0002
IPDPS6
2025 UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss
abstract
Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-varying regions, enhancing both simulation accuracy and computational efficiency. However, traditional approaches suffer from high computational complexity and geometric inflexibility, limiting their applicability, and existing supervised learning-based approaches face challenges in zero-shot generalization across diverse PDEs and mesh topologies. In this paper, we present an $\textbf{U}$nsupervised and $\textbf{G}$eneralizable $\textbf{M}$esh $\textbf{M}$ovement $\textbf{N}$etwork (UGM2N). We first introduce unsupervised mesh adaptation through localized geometric feature learning, eliminating the dependency on pre-adapted meshes. We then develop a physics-constrained loss function, M-Uniform loss, that enforces mesh equidistribution at the nodal level. Experimental results demonstrate that the proposed network exhibits equation-agnostic generalization and geometric independence in efficient mesh adaptation. It demonstrates consistent superiority over existing methods, including robust performance across diverse PDEs and mesh geometries, scalability to multi-scale resolutions and guaranteed error reduction without mesh tangling.
Xinhai Chen 0001, Xiang Gao 0020, Qingyang Zhang 0009, Menghan Jia, Xiang Zhang 0008, Jie Liu 0002
NeurIPS2
2025 3DMeshNet: A three-dimensional differential neural network for structured mesh generation
abstract
Mesh generation is a crucial step in numerical simulations, significantly impacting simulation accuracy and efficiency. However, generating meshes remains time-consuming and requires expensive computational resources. In this paper, we propose a novel method, 3DMeshNet, for three-dimensional structured mesh generation. The method embeds the meshing-related differential equations into the loss function of neural networks, formulating the meshing task as an unsupervised optimization problem. It takes geometric points as input to learn the potential mapping between parametric and computational spaces. After suitable offline training, 3DMeshNet can efficiently output a three-dimensional structured mesh with a user-defined number of quadrilateral/hexahedral cells through the feed-forward neural prediction. To enhance training stability and accelerate convergence, we integrate loss function reweighting through weight adjustments and gradient projection alongside applying finite difference methods to streamline derivative computations in the loss. Experiments on different cases show that 3DMeshNet is robust and fast. It outperforms neural network-based methods and yields superior meshes compared to traditional mesh partitioning methods. 3DMeshNet significantly reduces training times by up to 85% compared to other neural network-based approaches and lowers meshing overhead by 4 to 8 times relative to traditional meshing methods.
Jiaming Peng, Xinhai Chen 0001, Jie Liu 0002
Graph. Model.2
2025 HADF: a hash-adaptive dual fusion implicit network for super-resolution of turbulent flows
abstract
Turbulence, a complex multi-scale phenomenon inherent in fluid flow systems, presents critical challenges and opportunities for understanding physical mechanisms across scientific and engineering domains. Although high-resolution (HR) turbulence data remain indispensable for advancing both theoretical insights and engineering solutions, their acquisition is severely limited by prohibitively high computational costs. While deep learning architectures show transformative potential in reconstructing high-fidelity flow representations from sparse measurements, current methodologies suffer from two inherent constraints: strict reliance on perfectly paired training data and inability to perform multi-scale reconstruction within a unified framework. To address these challenges, we propose HADF, a hash-adaptive dynamic fusion implicit network for turbulence reconstruction. Specifically, we develop a low-resolution (LR) consistency loss that facilitates effective model training under conditions of missing paired data, eliminating the conventional requirement for fully matched LR and HR datasets. We further employ hash-adaptive spatial encoding and dynamic feature fusion to extract turbulence features, mapping them with implicit neural representations for reconstruction at arbitrary resolutions. Experimental results demonstrate that HADF achieves superior performance in global reconstruction accuracy and local physical properties compared to state-of-the-art models. It precisely recovers fine turbulence details for partially unpaired data conditions and diverse resolutions by training only once while maintaining robustness against noise.
Xinhai Chen 0001, Gen Zhang, Qingyang Zhang 0009, Jie Liu 0002
Frontiers Inf. Technol. Electron. Eng.2
2025 An intelligent mesh-smoothing method with graph neural networks
abstract
In computational fluid dynamics (CFD), mesh-smoothing methods are widely used to refine the mesh quality for achieving high-precision numerical simulations. Specifically, optimization-based smoothing is used for high-quality mesh smoothing, but it incurs significant computational overhead. Pioneer works have improved its smoothing efficiency by adopting supervised learning to learn smoothing methods from high-quality meshes. However, they pose difficulties in smoothing the mesh nodes with varying degrees and require data augmentation to address the node input sequence problem. Additionally, the required labeled high-quality meshes further limit the applicability of the proposed method. In this paper, we present graph-based smoothing mesh net (GMSNet), a lightweight neural network model for intelligent mesh smoothing. GMSNet adopts graph neural networks (GNNs) to extract features of the node’s neighbors and outputs the optimal node position. During smoothing, we also introduce a fault-tolerance mechanism to prevent GMSNet from generating negative volume elements. With a lightweight model, GMSNet can effectively smooth mesh nodes with varying degrees and remain unaffected by the order of input data. A novel loss function, MetricLoss, is developed to eliminate the need for high-quality meshes, which provides stable and rapid convergence during training. We compare GMSNet with commonly used mesh-smoothing methods on two-dimensional (2D) triangle meshes. Experimental results show that GMSNet achieves outstanding mesh-smoothing performances with 5% of the model parameters compared to the previous model, but offers a speedup of 13.56 times over the optimization-based smoothing.
Xinhai Chen 0001, Junjun Yan, Jie Liu 0002
Frontiers Inf. Technol. Electron. Eng.2
2024 SuperCSR: A Space-Time-Efficient CSR Representation for Large-scale Graph Applications on Supercomputers
abstract
It is widely accepted that graph representations such as the Compressed Sparse Row (CSR) format, directly affect the space and time complexities of graph processing. However, the standard CSR and its current variations are prone to high memory footprint and complicated calculations, which necessitates the development of more efficient graph processing techniques to save memory and reduce calculations. This paper presents SuperCSR, a more space-time-efficient CSR representation for fast graph processing. SuperCSR’s key idea is to leverage the law of sorted graphs, which would directly access adjacent vertex sets from an active vertex ID without complex indexing calculations and with a lower memory footprint.
Xinbiao Gan, Qiang Zhang 0053, Bo Yang 0023, Xinhai Chen 0001, Jie Liu 0002
ICPP5
2024 MST: Topology-Aware Message Aggregation for Exascale Graph Processing of Traversal-Centric Algorithms
abstract
This article presents MST, a communication-efficient message library for fast graph traversal on exascale clusters. The key idea is to follow the multi-level network topology to perform topology-aware message aggregation, where small messages are gathered and scattered at each level of domain. To facilitate message aggregation, we equip MST with flexible buffer management including active buffer switching and dynamic buffer expansion. We implement MST on the newest-generation Tianhe supercomputer and evaluated its performance using various traversal-centric algorithms on both synthetic trillion-scale graphs and real-world big graphs. The results show that MST-based graph traversal is orders of magnitude faster than that based on Active Messages Library (AML). For the Graph500-BFS benchmark, MST-based Tianhe (with 77.2 K nodes) outperforms the Fugaku supercomputer (with 148.5 K nodes) by 18.53%, while Fugaku is ranked No. 1 in the latest Graph500-BFS ranking (June 2023). MST also greatly improves graph processing performance on other commercial large-scale computing systems at the National Supercomputing Center in Changsha (NSCC) and WuzhenLight.
Xinbiao Gan, Bo Yang 0023, Xinhai Chen 0001, Chunye Gong, Shijie Li 0002, Kai Lu 0001, Qiao Li 0001, Yiming Zhang 0003
ACM Trans. Archit. Code Optim.5
2023 ST-PINN: A Self-Training Physics-Informed Neural Network for Partial Differential Equations
abstract
Partial differential equations (PDEs) are an essential computational kernel in physics and engineering. With the advance of deep learning, physics-informed neural networks (PINNs), as a mesh-free method, have shown great potential for fast PDE solving in various applications. To address the issue of low accuracy and convergence problems of existing PINNs, we propose a self-training physics-informed neural network, ST-PINN. Specifically, ST-PINN introduces a pseudo label based self-learning algorithm during training. It employs governing equation as the pseudo-labeled evaluation index and selects the highest confidence examples from the sample points to attach the pseudo labels. To our best knowledge, we are the first to incorporate a self-training mechanism into physics-informed learning. We conduct experiments on five PDE problems in different fields and scenarios. The results demonstrate that the proposed method allows the network to learn more physical information and benefit convergence. The ST-PINN outperforms existing physics-informed neural network methods and improves the accuracy by a factor of 1.33x-2.54x.
Junjun Yan, Xinhai Chen 0001, Enqiang Zhou, Jie Liu 0002
IJCNN2
2023 Predicting gene regulatory links from single-cell RNA-seq data using graph neural networks
abstract
Single-cell RNA-sequencing (scRNA-seq) has emerged as a powerful technique for studying gene expression patterns at the single-cell level. Inferring gene regulatory networks (GRNs) from scRNA-seq data provides insight into cellular phenotypes from the genomic level. However, the high sparsity, noise and dropout events inherent in scRNA-seq data present challenges for GRN inference. In recent years, the dramatic increase in data on experimentally validated transcription factors binding to DNA has made it possible to infer GRNs by supervised methods. In this study, we address the problem of GRN inference by framing it as a graph link prediction task. In this paper, we propose a novel framework called GNNLink, which leverages known GRNs to deduce the potential regulatory interdependencies between genes. First, we preprocess the raw scRNA-seq data. Then, we introduce a graph convolutional network-based interaction graph encoder to effectively refine gene features by capturing interdependencies between nodes in the network. Finally, the inference of GRN is obtained by performing matrix completion operation on node features. The features obtained from model training can be applied to downstream tasks such as measuring similarity and inferring causality between gene pairs. To evaluate the performance of GNNLink, we compare it with six existing GRN reconstruction methods using seven scRNA-seq datasets. These datasets encompass diverse ground truth networks, including functional interaction networks, Loss of Function/Gain of Function data, non-specific ChIP-seq data and cell-type-specific ChIP-seq data. Our experimental results demonstrate that GNNLink achieves comparable or superior performance across these datasets, showcasing its robustness and accuracy. Furthermore, we observe consistent performance across datasets of varying scales. For reproducibility, we provide the data and source code of GNNLink on our GitHub repository: https://github.com/sdesignates/GNNLink.
Guo Mao, Zhengbin Pang, Ke Zuo, Xiangdong Pei, Xinhai Chen 0001, Jie Liu 0002
Briefings Bioinform.6
2022 CSR&RV: An Efficient Value Compression Format for Sparse Matrix-Vector Multiplication
Junjun Yan, Xinhai Chen 0001, Jie Liu 0002
NPC2
2022 Improving the Performance of Lattice Boltzmann Method with Pipelined Algorithm on A Heterogeneous Multi-zone Processor
Qingyang Zhang 0009, Lei Xu 0035, Rongliang Chen, Lin Chen 0028, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023
PDCAT5
2022 A novel neural network approach for airfoil mesh quality evaluation
Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Yufei Pang, Liang Deng, Lihua Chi, Kenli Li 0001
J. Parallel Distributed Comput.1
2021 An efficient image to column algorithm for convolutional neural networks
abstract
Convolutional Neural Networks (CNNs) are a class of deep neural networks. The image to column (im2col) procedure is an important step for CNN and consumes about 28.8% of the whole inference time. In this paper, we present an efficient im2col algorithm, name im2cole (word “e” means efficient). The condition with different stride and pad in im2cole is well handled and the judgements in the innermost loop are removed. The procedure with pad = 1 is split into three conditions. This will reduce the pause of CPU instruction pipeline. The performances of the presented im2cole algorithm are reported with different inputs. Some discussion and performance issues are also reported. The experimental results show that the overall performance speedup of im2cole ranges from 2.12 to 4.33 compared with the original algorithm. The real application with Darknet shows that im2cole can get 20.75% whole performance improvement.
Chunye Gong, Xinhai Chen 0001, Shuling Lv, Jie Liu 0002, Bo Yang 0023, Weimin Bao, Yufei Pang
IJCNN2
2021 MVE-Net: An Automatic 3-D Structured Mesh Validity Evaluation Framework Using Deep Neural Networks
Xinhai Chen 0001, Jie Liu 0002, Chunye Gong, Shengguo Li, Yufei Pang
Comput. Aided Des.1
2021 A region-based hypergraph network for joint entity-relation extraction
Qian Wan 0007, Luona Wei, Xinhai Chen 0001, Jie Liu 0002
Knowl. Based Syst.3
2021 Configurable Multi-directional Systolic Array Architecture for Convolutional Neural Networks
abstract
The systolic array architecture is one of the most popular choices for convolutional neural network hardware accelerators. The biggest advantage of the systolic array architecture is its simple and efficient design principle. Without complicated control and dataflow, hardware accelerators with the systolic array can calculate traditional convolution very efficiently. However, this advantage also brings new challenges to the systolic array. When computing special types of convolution, such as the small-scale convolution or depthwise convolution, the processing element (PE) utilization rate of the array decreases sharply. The main reason is that the simple architecture design limits the flexibility of the systolic array. In this article, we design a configurable multi-directional systolic array (CMSA) to address these issues. First, we added a data path to the systolic array. It allows users to split the systolic array through configuration to speed up the calculation of small-scale convolution. Second, we redesigned the PE unit so that the array has multiple data transmission modes and dataflow strategies. This allows users to switch the dataflow of the PE array to speed up the calculation of depthwise convolution. In addition, unlike other works, we only make a few changes and modifications to the existing systolic array architecture. It avoids additional hardware overheads and can be easily deployed in application scenarios that require small systolic arrays such as mobile terminals. Based on our evaluation, CMSA can increase the PE utilization rate by up to 1.6 times compared to the typical systolic array when running the last layers of ResNet-18. When running depthwise convolution in MobileNet, CMSA can increase the utilization rate by up to 14.8 times. At the same time, CMSA and the traditional systolic arrays are similar in area and energy consumption.
Sheng Ma, Xinhai Chen 0001, Yang Guo 0003
ACM Trans. Archit. Code Optim.4
2020 OHTMA: an optimized heuristic topology-aware mapping algorithm on the Tianhe-3 exascale supercomputer prototype
abstract
With the rapid increase of the size of applications and the complexity of the supercomputer architecture, topology-aware process mapping becomes increasingly important. High communication cost has become a dominant constraint of the performance of applications running on the supercomputer. To avoid a bad mapping strategy which can lead to terrible communication performance, we propose an optimized heuristic topology-aware mapping algorithm (OHTMA). The algorithm attempts to minimize the hop-byte metric that we use to measure the mapping results. OHTMA incorporates a new greedy heuristic method and pair-exchange-based optimization. It reduces the number of long-distance communications and effectively enhances the locality of the communication. Experimental results on the Tianhe-3 exascale supercomputer prototype indicate that OHTMA can significantly reduce the communication costs.
Yishui Li, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Chunye Gong, Xinbiao Gan, Shengguo Li, Han Xu 0008
Frontiers Inf. Technol. Electron. Eng.2
2020 VBSF: a new storage format for SIMD sparse matrix-vector multiplication on modern processors
Yishui Li, Peizhen Xie, Xinhai Chen 0001, Jie Liu 0002, Bo Yang 0023, Shengguo Li, Chunye Gong, Xinbiao Gan, Han Xu 0008
J. Supercomput.3
2018 TAMM: A New Topology-Aware Mapping Method for Parallel Applications on the Tianhe-2A Supercomputer
Xinhai Chen 0001, Jie Liu 0002, Shengguo Li, Peizhen Xie, Lihua Chi
ICA3PP (1)1
2018 An efficient SIMD compression format for sparse matrix-vector multiplication
abstract
Summary Sparse matrix‐vector multiplication (SpMV) is an essential kernel in sparse linear algebra and has been studied extensively on all modern processor and accelerator architectures. Compressed Sparse Row (CSR) is a frequently used format for sparse matrices storage. However, CSR‐based SpMV has poor performance on processors with vector units. In order to take full advantage of SIMD acceleration technology in SpMV, we proposed a new matrix storage format called CSR‐SIMD. The new storage format compresses the non‐zero elements into many variable‐length data fragments with consecutive memory access addresses. Thus, the data locality of sparse matrix A and dense vector x expands and the floating‐point operations for each fragment can be completely calculated by vectorized implementation on wide SIMD units. Our experimental results indicate that CSR‐SIMD has better storage efficiency and low‐overhead for format conversion. Besides, the new format achieves high scalability on wide SIMD units. In comparison with the CSR‐based and BCSR‐based SpMV, CSR‐SIMD obtains better performance on FT1500A, Intel Xeon, and Intel Xeon Phi.
Xinhai Chen 0001, Peizhen Xie, Lihua Chi, Jie Liu 0002, Chunye Gong
Concurr. Comput. Pract. Exp.1