Tao He 0016

dblp:94/5035-16 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-9405-3979ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 CNM-UNet: Continuous Ordinary Differential Equations for Medical Image Segmentation
abstract
Integrating Ordinary Differential Equations (ODEs) with U-shaped neural networks has emerged as a novel direction in medical image segmentation. Current networks predominantly employ discretization methods incorporating ODEs. However, these methods face inherent trade-offs between model compactness, computational accuracy, and efficiency. Continuous ODE solutions were rarely studied because they face three limitations: high computational costs, long training time, and poor generalization ability. To address these limitations, we propose an innovative Continuous Neural Memory ODE UNet (CNM-UNet), which replaces all hierarchical decoder layers in vanilla UNet with a single Continuous Neural Memory ODEs Block (CNM-Block) decoder, significantly reducing computation costs and improving training efficiency. CNM-UNet leverages ODEs' dynamic properties to establish continuous temporal feature extraction. For alleviating the generalization problem, a DUal SElf-updated (DUSE) strategy based on test-time adaptation principles is introduced to enhance cross-domain generalization. Experimental results demonstrate CNM-UNet's comprehensive advantages in computational capacity, convergence speed, and cross-domain adaptability, offering new insights for practical deployment of continuous ODE methodologies for medical image segmentation.
Yashi Zhu, Quansong He, Kaishen Wang, Zhang Yi 0001, Tao He 0016
AAAI7
2026 Enhancing feature fusion of U-like networks with dynamic skip connections
Quansong He, Kaishen Wang, Jianlong Xiong, Zhang Yi 0001, Tao He 0016
Medical Image Anal.6
2025 FuseUNet: A Multi-Scale Feature Fusion Method for U-like Networks
abstract
Medical image segmentation is a critical task in computer vision, with UNet serving as a milestone architecture. The typical component of UNet family is the skip connection, however, their skip connections face two significant limitations: (1) they lack effective interaction between features at different scales, and (2) they rely on simple concatenation or addition operations, which constrain efficient information integration. While recent improvements to UNet have focused on enhancing encoder and decoder capabilities, these limitations remain overlooked. To overcome these challenges, we propose a novel multi-scale feature fusion method that reimagines the UNet decoding process as solving an initial value problem (IVP), treating skip connections as discrete nodes. By leveraging principles from the linear multistep method, we propose an adaptive ordinary differential equation method to enable effective multi-scale feature fusion. Our approach is independent of the encoder and decoder architectures, making it adaptable to various U-Net-like networks. Experiments on ACDC, KiTS2023, MSD brain tumor, and ISIC2017/2018 skin lesion segmentation datasets demonstrate improved feature utilization, reduced network parameters, and maintained high performance. The code is available at https://github.com/nayutayuki/FuseUNet.
Quansong He, Xiangde Min, Kaishen Wang, Tao He 0016
ICML4
2025 MobileODE: An Extra Lightweight Network
abstract
Depthwise-separable convolution has emerged as a significant milestone in the lightweight development of Convolutional Neural Networks (CNNs) over the past decade. This technique consists of two key components: depthwise convolution, which captures spatial information, and pointwise convolution, which enhances channel interactions. In this paper, we propose a novel method to lightweight CNNs through the discretization of Ordinary Differential Equations (ODEs). Specifically, we optimize depthwise-separable convolution by replacing the pointwise convolution with a discrete ODE module, termed the \emph{\textbf{C}hannelwise \textbf{O}DE \textbf{S}olver (COS)}. The COS module is constructed by a simple yet efficient direct differentiation Euler algorithm, using learnable increment parameters. This replacement reduces parameters by over $98.36$\% compared to conventional pointwise convolution. By integrating COS into MobileNet, we develop a new extra lightweight network called MobileODE. With carefully designed basic and inverse residual blocks, the resulting MobileODEV1 and MobileODEV2 reduce channel interaction parameters by $71.0$\% and $69.2$\%, respectively, compared to MobileNetV1, while achieving higher accuracy across various tasks, including image classification, object detection, and semantic segmentation. The code is available at {\url{https://github.com/cashily/MobileODE}}.
Bo Gou, Xiangde Min, Lei Zhang 0005, Zhang Yi 0001, Tao He 0016
NeurIPS7
2025 Neural Memory State Space Models for Medical Image Segmentation
abstract
With the rapid advancement of deep learning, computer-aided diagnosis and treatment have become crucial in medicine. UNet is a widely used architecture for medical image segmentation, and various methods for improving UNet have been extensively explored. One popular approach is incorporating transformers, though their quadratic computational complexity poses challenges. Recently, State-Space Models (SSMs), exemplified by Mamba, have gained significant attention as a promising alternative due to their linear computational complexity. Another approach, neural memory Ordinary Differential Equations (nmODEs), exhibits similar principles and achieves good results. In this paper, we explore the respective strengths and weaknesses of nmODEs and SSMs and propose a novel architecture, the nmSSM decoder, which combines the advantages of both approaches. This architecture possesses powerful nonlinear representation capabilities while retaining the ability to preserve input and process global information. We construct nmSSM-UNet using the nmSSM decoder and conduct comprehensive experiments on the PH2, ISIC2018, and BU-COCO datasets to validate its effectiveness in medical image segmentation. The results demonstrate the promising application value of nmSSM-UNet. Additionally, we conducted ablation experiments to verify the effectiveness of our proposed improvements on SSMs and nmODEs.
Jingjun Gu, Quansong He, Tianli Zhao, Jialong Guo, Tao He 0016, Jiajun Bu
Int. J. Neural Syst.8
2025 A pseudo-3D coarse-to-fine architecture for 3D medical landmark detection
Boyan Liu, Guikun Xu, Jixiang Guo, Tao He 0016
Neurocomputing6
2025 Neural Memory Self-Supervised State Space Models With Learnable Gates
abstract
Discrete Ordinary Differential Equations (ODEs) have been employed to develop lightweight deep neural networks in recent years. In this letter, we introduce a novel lightweight UNet variant called the neural memory Self-Supervised State Space Model (nmS4M-UNet) for medical image segmentation, where discrete ODEs serve as the decoder. The proposed nmS4M block has learnable gates and performs multi-head computation to enhance memory updates. Additionally, the nmS4M-UNet incorporates a self-supervised learning branch to improve feature extraction capabilities. The intermediate features are reused as partial input to the decoder, helping to mitigate network overfitting. The nmS4M-UNet reduces the number of parameters by 29.70% compared to the standard UNet. Experimental results on the PH2, ISIC2018, and BU-COCO datasets demonstrate that the proposed nmS4M-UNet achieves performance comparable to state-of-the-art models.
Zhang Yi 0001, Tao He 0016, Jiajun Bu
IEEE Signal Process. Lett.4
2024 A Lightweight U-like Network Utilizing Neural Memory Ordinary Differential Equations for Slimming the Decoder
Quansong He, Zhang Yi 0001, Tao He 0016
IJCAI5
2024 Strengthening Layer Interaction via Dynamic Layer Attention
Kaishen Wang, Xun Xia, Jian Liu 0041, Zhang Yi 0001, Tao He 0016
IJCAI5
2024 A novel masking model for Buddhist literature understanding by using Generative Adversarial Networks
Chaowen Yan, Lili Chang, Tao He 0016
Expert Syst. Appl.5
2024 A Bidirectional Feedforward Neural Network Architecture Using the Discretized Neural Memory Ordinary Differential Equation
abstract
Deep Feedforward Neural Networks (FNNs) with skip connections have revolutionized various image recognition tasks. In this paper, we propose a novel architecture called bidirectional FNN (BiFNN), which utilizes skip connections to aggregate features between its forward and backward paths. The BiFNN accepts any FNN as a plugin that can incorporate any general FNN model into its forward path, introducing only a few additional parameters in the cross-path connections. The backward path is implemented as a nonparameter layer, utilizing a discretized form of the neural memory Ordinary Differential Equation (nmODE), which is named [Formula: see text]-net. We provide a proof of convergence for the [Formula: see text]-net and evaluate its initial value problem. Our proposed architecture is evaluated on diverse image recognition datasets, including Fashion-MNIST, SVHN, CIFAR-10, CIFAR-100, and Tiny-ImageNet. The results demonstrate that BiFNNs offer significant improvements compared to embedded models such as ConvMixer, ResNet, ResNeXt, and Vision Transformer. Furthermore, BiFNNs can be fine-tuned to achieve comparable performance with embedded models on Tiny-ImageNet and ImageNet-1K datasets by loading the same pretrained parameters.
Zhang Yi 0001, Tao He 0016
Int. J. Neural Syst.3
2024 Anchor Ball Regression Model for large-scale 3D skull landmark detection
Tao He 0016, Guikun Xu, Jixiang Guo
Neurocomputing1
2023 An automatic methodology for full dentition maturity staging from OPG images using deep learning
Wenxuan Dong, Meng You, Tao He 0016, Jiaqi Dai, Yueting Tang, Yuchao Shi, Jixiang Guo
Appl. Intell.3
2023 Cascade-refine model for cephalometric landmark detection in high-resolution orthodontic images
Tao He 0016, Jixiang Guo, Fanxin Zeng, Zhang Yi 0001
Knowl. Based Syst.1
2022 Subtraction Gates: Another Way to Learn Long-Term Dependencies in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) can remember temporal contextual information over various time steps. The well-known gradient vanishing/explosion problem restricts the ability of RNNs to learn long-term dependencies. The gate mechanism is a well-developed method for learning long-term dependencies in long short-term memory (LSTM) models and their variants. These models usually take the multiplication terms as gates to control the input and output of RNNs during forwarding computation and to ensure a constant error flow during training. In this article, we propose the use of subtraction terms as another type of gates to learn long-term dependencies. Specifically, the multiplication gates are replaced by subtraction gates, and the activations of RNNs input and output are directly controlled by subtracting the subtrahend terms. The error flows remain constant, as the linear identity connection is retained during training. The proposed subtraction gates have more flexible options of internal activation functions than the multiplication gates of LSTM. The experimental results using the proposed Subtraction RNN (SRNN) indicate comparable performances to LSTM and gated recurrent unit in the Embedded Reber Grammar, Penn Tree Bank, and Pixel-by-Pixel MNIST experiments. To achieve these results, the SRNN requires approximate three-quarters of the parameters used by LSTM. We also show that a hybrid model combining multiplication forget gates and subtraction gates could achieve good performance.
Tao He 0016, Hua Mao 0001, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Cephalometric landmark detection by considering translational invariance in the two-stage framework
Tao He 0016, Zhang Yi 0001, Jixiang Guo
Neurocomputing1
2020 Multilabel classification by exploiting data-driven pair-wise label dependence
abstract
Exploiting label dependence is a widely used approach to boost classification performance for multilabel classification problems. However, most of the traditional label dependence methods have high time complexity, especially when combined with deep neural networks (DNNs). Thus they usually can not be efficiently applied in large-scale data sets. Recent advances in large-scale multilabel classification widely developed pair-wise ranking and structure-driven methods, but label dependence was little exploited. In most of the structure-driven methods, binary relevance (BR) with multiple binary cross-entropy (BCE) loss functions, a simple but effective method, is still the prior solution incorporation with DNNs in large-scale data sets. In this paper, we propose a novel loss function called label dependent cross-entropy (LDCE), which directly introduces label dependence to BCE loss function by data-driven conditional probability. Combined with deep convolutional neural networks (DCNNs), LDCE introduces no extra parameters and induces very little extra computational complexity. Moreover, we develop its tiny variant with sparse label dependence and its learnable version for automatic learning pair-wise label dependence. Within the BR scheme, LDCE outperforms BCE on seven widely used benchmark datasets. We also perform two large-scale multilabel image classification tasks (VOC 2007 and ChestX-ray14) with DCNNs, and LDCE outperforms BCE and achieves comparable results to the state-of-the-art.
Tao He 0016, Lei Zhang 0005, Jixiang Guo, Zhang Yi 0001
Int. J. Intell. Syst.1
2020 Multi-task learning for the segmentation of organs at risk with label dependence
Tao He 0016, Junjie Hu 0004, Jixiang Guo, Zhang Yi 0001
Medical Image Anal.1
2020 MediMLP: Using Grad-CAM to Extract Crucial Variables for Lung Cancer Postoperative Complication Prediction
abstract
Lung cancer postoperative complication prediction (PCP) is significant for decreasing the perioperative mortality rate after lung cancer surgery. In this paper we concentrate on two PCP tasks: (1) the binary classification for predicting whether a patient will have postoperative complications; and (2) the three-class multi-label classification for predicting which postoperative complication a patient will experience. Furthermore, an important clinical requirement of PCP is the extraction of crucial variables from electronic medical records. We propose a novel multi-layer perceptron (MLP) model called medical MLP (MediMLP) together with the gradient-weighted class activation mapping (Grad-CAM) algorithm for lung cancer PCP. The proposed MediMLP, which involves one locally connected layer and fully connected layers with a shortcut connection, simultaneously extracts crucial variables and performs PCP tasks. The experimental results indicated that MediMLP outperformed normal MLP on two PCP tasks and had comparable performance with existing feature selection methods. Using MediMLP and further experimental analysis, we found that the variable of "time of indwelling drainage tube" was very relevant to lung cancer postoperative complications.
Tao He 0016, Jixiang Guo, Xiuyuan Xu, Zihuai Wang, Kaiyu Fu, Lunxu Liu, Zhang Yi 0001
IEEE J. Biomed. Health Informatics1
2017 Cell tracking using deep neural networks with multi-task learning
Tao He 0016, Hua Mao 0001, Jixiang Guo, Zhang Yi 0001
Image Vis. Comput.1
2017 Moving object recognition using multi-view three-dimensional convolutional neural networks
Tao He 0016, Hua Mao 0001, Zhang Yi 0001
Neural Comput. Appl.1
2016 Learning to Appreciate the Aesthetic Effects of Clothing
abstract
How do people describe clothing? The words like “formal”or "casual" are usually used. However, recent works often focus on recognizing or extracting visual features (e.g., sleeve length, color distribution and clothing pattern) from clothing images accurately. How can we bridge the gap between the visual features and the aesthetic words? In this paper, we formulate this task to a novel three-level framework: visual features(VF) - image-scale space (ISS) - aesthetic words space(AWS). Leveraging the art-field image-scale space served as an intermediate layer, we first propose a Stacked Denoising Autoencoder Guided by CorrelativeLabels (SDAE-GCL) to map the visual features to the image-scale space; and then according to the semantic distances computed byWordNet::Similarity, we map the most often used aesthetic words in online clothing shops to the image-scale space too. Employing upper body menswear images downloaded from several global online clothing shops as experimental data, the results indicate that the proposed three-level framework can help to capture the subtle relationship between visual features and aesthetic words better compared to several baselines. To demonstrate that our three-level framework and its implementation methods are universally applicable, we finally present some interesting analyses on the fashion trend of menswear in the last 10 years.
Jia Jia 0001, Guangyao Shen, Tao He 0016, Zhiyuan Liu 0001, Huan-Bo Luan
AAAI4
2016 Inferring users' emotions for human-mobile voice dialogue applications
abstract
In this paper, we tackle the problem of inferring users' emotions in real-world Voice Dialogue Applications (VDAs, Siri1, Cortana2, etc.). We first conduct an investigation, indicating that besides the text information of users' queries, the acoustic information and query attributes are very important in inferring emotions in VDAs. To integrate the information above, we propose a Hybrid Emotion Inference Model (HEIM), which involves a Latent Dirichlet Allocation (LDA) to extract text features and a Long Short-Term Memory (LSTM) to model the acoustic features. To further improve accuracy, a Recurrent Autoencoder Guided by Query Attributes (RAGQA) which incorporates other emotion-related query attributes is proposed in HEIM to pre-train LSTM. The accuracy of HEIM on a data set collected from Sogou Voice Assistant3(Chinese Siri) containing 93,000 utterances achieves 75.2%, which outperforms state-of-the-art methods for 33.5–38.5%. Specifically, we discover that on average, the acoustic information enhances the performance for 46.6%, while query attributes further enhance the performance for 6.5%.
Boya Wu, Jia Jia 0001, Tao He 0016, Xiaoyuan Yi, Yishuang Ning
ICME3