Hao Ni 0001

dblp:25/9101-1 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0001-5485-4376ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 PathFusion-Net: A Rough Path Theory-Based Deep Learning Model for ECG Arrhythmia Classification
abstract
This study introduces a novel electrocardiogram (ECG) arrhythmia classification model, PathFusion-Net, which integrates Rough Path Theory with deep learning technologies. The model combines Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM), Path Signatures, and Path Development to extract spatial morphological features from ECG images and multi-order temporal representations from ECG signals. By adopting an inter-patient split paradigm, our approach more closely reflects real-world clinical diagnostic settings compared to intra-patient methods. The model demonstrates state-of-the-art overall classification performance on both the MIT-BIH Arrhythmia Database and a private clinical dataset, achieving 94.7% and 95.1% accuracy, respectively, under the AAMI four-class standard with an inter-patient split paradigm. On the MIT-BIH dataset, the proposed method attains competitive precision and recall across multiple arrhythmia types, including 95.2% /87.9% for ventricular ectopic beats (V) and 75.7% /92.3% for supraventricular ectopic beats (S), indicating balanced performance across clinically diverse categories. This research highlights the potential of Rough Path Theory in time-series analysis and offers a novel deep learning framework for automated early detection and monitoring of ECG arrhythmias. The code used in this study is available at: https://github.com/Rand2AI/PathFusion-Net.
Tianlong Feng, Qingchen Li, Yongzhi Liao, Di Lu 0001, Jianqin Zhao, Hao Ni 0001, Hongying Liu 0001, Jingjing Deng 0001
IEEE J. Biomed. Health Informatics9
2025 MCGAN: Enhancing GAN Training with Regression-Based Generator Loss
abstract
Generative adversarial networks (GANs) have emerged as a powerful tool for generating high-fidelity data. However, the main bottleneck of existing approaches is the lack of supervision on the generator training, which often results in undamped oscillation and unsatisfactory performance. To address this issue, we propose an algorithm called Monte Carlo GAN (MCGAN). This approach, utilizing an innovative generative loss function, termed the regression loss, reformulates the generator training as a regression task and enables the generator training by minimizing the mean squared error between the discriminator's output of real data and the expected discriminator of fake data. We demonstrate the desirable analytic properties of the regression loss, including discriminability and optimality, and show that our method requires a weaker condition on the discriminator for effective generator training. These properties justify the strength of this approach to improve the training stability while retaining the optimality of GAN by leveraging strong supervision of the regression loss. Extensive experiments on diverse datasets, including image data (CIFAR-10/100, FFHQ256, ImageNet, and LSUN Bedroom), time series data (VAR and stock data), and video data, are conducted to demonstrate the flexibility and effectiveness of our proposed MCGAN. Numerical results show that the proposed MCGAN is versatile in enhancing a variety of backbone GAN models and achieves consistent and significant improvement in terms of quality, accuracy, training stability, and learned latent space.
Baoren Xiao, Hao Ni 0001
AAAI2
2025 DevInSight: Weaving Path Development Into Online Signature Verification
Yilin Shi, Hao Ni 0001
ICDAR (5)3
2024 Adaptive Global Gesture Paths and Signature Features for Skeleton-based Gesture Recognition
Dongzi Shi, Xin Zhang 0013, Tong Xiong, Hao Ni 0001
ICPR (15)5
2024 NODE-ImgNet: A PDE-informed effective and robust model for image denoising
Xinheng Xie, Yue Wu 0016, Hao Ni 0001, Cuiyu He
Pattern Recognit.3
2024 Skeleton-Based Gesture Recognition With Learnable Paths and Signature Features
abstract
For the skeleton-based gesture recognition, graph convolutional networks (GCNs) have achieved remarkable performance since the human skeleton is a natural graph. However, the biological structure might not be the crucial one for motion analysis. Also, spatial differential information like joint distance and angle between bones may be overlooked during the graph convolution. In this article, we focus on obtaining meaningful joint groups and extracting their discriminative features by the path signature (PS) theory. Firstly, to characterize the constraints and dependencies of various joints, we propose three types of paths, i.e., spatial, temporal, and learnable path. Especially, a learnable path generation mechanism can group joints together that are not directly connected or far away, according to their kinematic characteristic. Secondly, to obtain informative and compact features, a deep integration of PS with few parameters are introduced. All the computational process is packed into two modules, i.e., spatial-temporal path signature module (ST-PSM) and learnable path signature module (L-PSM) for the convenience of utilization. They are plug-and-play modules available for any neural network like CNNs and GCNs to enhance the feature extraction ability. Extensive experiments have conducted on three mainstream datasets (ChaLearn 2013, ChaLearn 2016, and AUTSL). We achieved the state-of-the-art results with simpler framework and much smaller model size. By inserting our two modules into the several GCN-based networks, we can observe clear improvements demonstrating the great effectiveness of our proposed method.
Dongzi Shi, Chenyang Li 0007, Yu Li 0043, Hao Ni 0001, Xin Zhang 0013
IEEE Trans. Multim.5
2023 EnsExam: A Dataset for Handwritten Text Erasure on Examination Papers
Liufeng Huang, Bangdong Chen, Chongyu Liu, Dezhi Peng, Weiying Zhou, Yaqiang Wu, Hao Ni 0001
ICDAR (3)8
2023 SegCTC: Offline Handwritten Chinese Text Recognition via Better Fusion Between Explicit and Implicit Segmentation
Jiarong Huang, Dezhi Peng, Hao Ni 0001
ICDAR (4)4
2022 Path Signature Neural Network of Cortical Features for Prediction of Infant Cognitive Scores
abstract
Studies have shown that there is a tight connection between cognition skills and brain morphology during infancy. Nonetheless, it is still a great challenge to predict individual cognitive scores using their brain morphological features, considering issues like the excessive feature dimension, small sample size and missing data. Due to the limited data, a compact but expressive feature set is desirable as it can reduce the dimension and avoid the potential overfitting issue. Therefore, we pioneer the path signature method to further explore the essential hidden dynamic patterns of longitudinal cortical features. To form a hierarchical and more informative temporal representation, in this work, a novel cortical feature based path signature neural network (CF-PSNet) is proposed with stacked differentiable temporal path signature layers for prediction of individual cognitive scores. By introducing the existence embedding in path generation, we can improve the robustness against the missing data. Benefiting from the global temporal receptive field of CF-PSNet, characteristics consisted in the existing data can be fully leveraged. Further, as there is no need for the whole brain to work for a certain cognitive ability, a top K selection module is used to select the most influential brain regions, decreasing the model size and the risk of overfitting. Extensive experiments are conducted on an in-house longitudinal infant dataset within 9 time points. By comparing with several recent algorithms, we illustrate the state-of-the-art performance of our CF-PSNet (i.e., root mean square error of 0.027 with the time latency of 518 milliseconds for each sample).
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001
IEEE Trans. Medical Imaging3
2021 Logsig-RNN: a novel network for robust and efficient skeleton-based action recognition
Hao Ni 0001, Shujian Liao, Kevin Schlegel, Terry J. Lyons
BMVC1
2020 Signature features with the visibility transformation
abstract
In this paper we put the visibility transformation on a clear theoretical footing and show that this transform is able to embed the effect of the absolute position of the data stream into signature features in a unified and efficient way. The generated feature set is particularly useful in pattern recognition tasks, for its simplifying role in allowing the signature feature set to accommodate nonlinear functions of absolute and relative values.
Yue Wu 0016, Hao Ni 0001, Terry J. Lyons, Robin L. Hudson
ICPR2
2020 Infant Cognitive Scores Prediction with Multi-stream Attention-Based Temporal Path Signature Features
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001
MICCAI (7)3
2020 Simultaneous left atrium anatomy and scar segmentations via deep learning in multiview information with attention
abstract
Three-dimensional late gadolinium enhanced (LGE) cardiac MR (CMR) of left atrial scar in patients with atrial fibrillation (AF) has recently emerged as a promising technique to stratify patients, to guide ablation therapy and to predict treatment success. This requires a segmentation of the high intensity scar tissue and also a segmentation of the left atrium (LA) anatomy, the latter usually being derived from a separate bright-blood acquisition. Performing both segmentations automatically from a single 3D LGE CMR acquisition would eliminate the need for an additional acquisition and avoid subsequent registration issues. In this paper, we propose a joint segmentation method based on multiview two-task (MVTT) recursive attention model working directly on 3D LGE CMR images to segment the LA (and proximal pulmonary veins) and to delineate the scar on the same dataset. Using our MVTT recursive attention model, both the LA anatomy and scar can be segmented accurately (mean Dice score of 93% for the LA anatomy and 87% for the scar segmentations) and efficiently (∼0.27 s to simultaneously segment the LA anatomy and scars directly from the 3D LGE CMR dataset with 60–68 2D slices). Compared to conventional unsupervised learning and other state-of-the-art deep learning based methods, the proposed MVTT model achieved excellent results, leading to an automatic generation of a patient-specific anatomical model combined with scar segmentation for patients in AF.
Guang Yang 0006, Jun Chen 0030, Zhifan Gao, Shuo Li 0001, Hao Ni 0001, Elsa D. Angelini, Tom Wong, Raad Mohiaddin, Eva Nyktari, Rick Wage, Lei Xu 0037, Yanping Zhang 0001, Xiuquan Du, Heye Zhang, David N. Firmin, Jennifer Keegan
Future Gener. Comput. Syst.5
2019 A Path Signature Approach for Speech Emotion Recognition
abstract
Automatic speech emotion recognition (SER) remains a \ndifficult task within human-computer interaction, despite increasing interest in the research community. One key challenge is how to effectively integrate short-term characterisation \nof speech segments with long-term information such as temporal variations. Motivated by the numerical approximation theory of stochastic differential equations (SDEs), we propose the \nnovel use of path signatures. The latter provide a pathwise definition to solve SDEs, for the integration of short speech frames. \nFurthermore we propose a hierarchical tree structure of path signatures, to capture both global and local information. A simple tree-based convolutional neural network (TBCNN) is used \nfor learning the structural information stemming from dyadic \npath-tree signatures. Our experimental results on a widely \nused benchmark dataset demonstrate comparable performance \nto complex neural network based systems.
Bo Wang 0034, Maria Liakata, Hao Ni 0001, Terry J. Lyons, Alejo J. Nevado-Holgado, Kate Saunders
INTERSPEECH3
2018 Multiview Two-Task Recursive Attention Model for Left Atrium and Atrial Scars Segmentation
Jun Chen 0030, Guang Yang 0006, Zhifan Gao, Hao Ni 0001, Elsa D. Angelini, Raad Mohiaddin, Tom Wong, Yanping Zhang 0001, Xiuquan Du, Heye Zhang, Jennifer Keegan, David N. Firmin
MICCAI (2)4
2018 Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
abstract
Online handwritten Chinese text recognition (OHCTR) is a challenging problem as it involves a large-scale character set, ambiguous segmentation, and variable-length input sequences. In this paper, we exploit the outstanding capability of path signature to translate online pen-tip trajectories into informative signature feature maps, successfully capturing the analytic and geometric properties of pen strokes with strong local invariance and robustness. A multi-spatial-context fully convolutional recurrent network (MC-FCRN) is proposed to exploit the multiple spatial contexts from the signature feature maps and generate a prediction sequence while completely avoiding the difficult segmentation problem. Furthermore, an implicit language model is developed to make predictions based on semantic context within a predicting feature sequence, providing a new perspective for incorporating lexicon constraints and prior knowledge about a certain language in the recognition procedure. Experiments on two standard benchmarks, Dataset-CASIA and Dataset-ICDAR, yielded outstanding results, with correct rates of 97.50 and 96.58 percent, respectively, which are significantly better than the best result reported thus far in the literature.
Zecheng Xie, Zenghui Sun, Hao Ni 0001, Terry J. Lyons
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Rotation-free online handwritten character recognition using dyadic path signature features, hanging normalization, and deep neural network
abstract
The path signature feature (PSF) which was initially introduced in rough paths theory as a branch of stochastic analysis, has recently been successfully applied to the field of pattern recognition for extracting sufficient quantity of information contained in a finite trajectory, but with potentially high dimension. In this paper, we propose a variation of path signature representation, namely the dyadic path signature feature (D-PSF), to fully characterize the trajectory using a hierarchical structure to solve the rotation-free online handwritten character recognition (OLHCR) problem. We adopt the deep neural network (DNN) as classifier, and investigate three hanging normalization methods to improve the robustness of the DNN to rotational distortions. Extensive experiments on digits, English letters, and Chinese radicals demonstrated that the proposed D-PSF, jointly with hanging normalization and DNN, achieved very promising results for rotated OLHCR, significantly outperforming previous methods.
Hao Ni 0001, Terry J. Lyons
ICPR3