Along He

dblp:243/9296 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-1356-8757ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TAPE: A multi-agent framework for task-adaptive planning and execution in resource-constrained environments
Along He, Haobin Wang
Expert Syst. Appl.3
2026 Federated semi-supervised calibrated efficient fine-tuning of foundation models for medical image classification
Along He, Yanlin Wu, LinLin Shen, Ke Zou, Huazhu Fu
Knowl. Based Syst.1
2025 HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration
abstract
Mixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top-k routing mechanisms.Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant performance degradation.To address this trade-off between efficiency and performance, we propose Hook-MoE, a plug-and-play single-layer compensation framework that effectively restores performance using only a small post-training calibration set.Our method strategically inserts a lightweight trainable Hook module immediately preceding selected transformer blocks.Comprehensive evaluations on four popular MoE models, with an average performance degradation of only 2.5% across various benchmarks, our method reduces the number of activated experts by more than 50% and achieves a 1.42× inference speed-up during the prefill stage.Through systematic analysis, we further reveal that the upper layers require fewer active experts, offering actionable insights for refining dynamic expert selection strategies and enhancing the overall efficiency of MoE models.We make our code available at https://github.com/KerwinKai/HookMoE.
Longkai Cheng, Along He, Mulin Li, Xueshuo Xie, Tao Li 0022
EMNLP2
2025 A Novel Framework for Data Augmentation on Appearance Changes for Long-Term Person Re-Identification
abstract
The general Re-ID works including datasets and networks have an assumption that the pedestrians do not change their appearances throughout the research. The lack of diversity about the appearance of the Re-ID datasets significantly limits the performance of models in cross-appearance Re-ID. Therefore, this paper conduct data augmentation in the appearance dimension to the datasets by the generation framework CAG(cross-appearance images generation) to support the studies about cross-appearance Re-ID. To demonstrate the effectiveness of the generation framework, we selected several classical and SOTA Re-ID models and conducted experiments on three commonly used cross-appearance Re-ID datasets, NKUP+IPRCC/DeepChange. The results show that the cross-appearance Re-ID images generated by CAG can help the models to obtain the robust pedestrian feature, and the diversity of appearance is a universal method for cross-appearance person re-identification.
Tao Li 0022, Along He, Tehui Huang, Qiankun Dong
JCC4
2025 Towards Automated Pediatric Dental Development Staging: A Dataset and Model
Peng Wang 0178, Along He, Anli Wang, Zhenhuan Zhou, Xiaohang Guan, Tao Li 0022
MICCAI (13)2
2025 GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
abstract
Medical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in multi-modal learning have significantly improved performance, current methods still suffer from limited answer reliability and poor interpretability, impairing the ability of clinicians and patients to understand and trust model outputs. To address these limitations, this work first proposes a Region-Aware Multimodal Chain-of-Thought (RMCoT) dataset, in which the process of producing an answer is preceded by a sequence of intermediate reasoning steps that explicitly ground relevant visual regions of the medical image, thereby providing fine-grained explainability. Furthermore, we introduce a novel verifiable reward mechanism for reinforcement learning to guide post-training, improving the alignment between the model's reasoning process and its final answer. Remarkably, our method achieves comparable performance using only one-eighth of the training data, demonstrating the efficiency and effectiveness of the proposal. The dataset is available at https://www.med-vqa.com/GEMeX/.
Bo Liu 0113, Along He, Huazhu Fu, Xiao-Ming Wu 0003
ACM Multimedia3
2025 Trans-SAM: Transfer Segment Anything Model to medical image segmentation with Parameter-Efficient Fine-Tuning
Yanlin Wu, Xiongfeng Yang, Hong Kang, Along He, Tao Li 0022
Knowl. Based Syst.5
2025 AdaptFRCNet: Semi-supervised adaptation of pre-trained model with frequency and region consistency for medical image segmentation
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu
Medical Image Anal.1
2025 DVPT: Dynamic Visual Prompt Tuning of large pre-trained models for medical image analysis
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu
Neural Networks1
2024 Spatial-Frequency Dual Domain Attention Network For Medical Image Segmentation
abstract
In medical images, various types of lesions often manifest significant differences in their shape and texture. Accurate medical image segmentation demands deep learning models with robust capabilities in multi-scale and boundary feature learning. However, previous models still have limitations in addressing the above issues. The majority of medical image segmentation networks exclusively learn features in the spatial domain, disregarding the abundant global information in the frequency domain. This results in a bias towards low-frequency components, neglecting crucial high-frequency information. To address these problems, we introduce SF-UNet, a spatial-frequency dual-domain attention network. It comprises two main components: the Multi-scale Progressive Channel Attention (MPCA) block, which progressively extract multi-scale features across adjacent encoder layers, and the lightweight Frequency-Spatial Attention (FSA) block, with only 0.05M parameters, enabling concurrent learning of texture and boundary features from both spatial and frequency domains. We validate the effectiveness of the proposed SF-UNet on three public datasets. Experimental results show that compared to previous state-of-the-art medical image segmentation networks, SF-UNet achieves the best performance, and achieves up to 9.4% and 10.78% improvement in DSC and IOU. Codes will be released at https://github.com/nkicsl/SF-UNet.
Zhenhuan Zhou, Along He, Yanlin Wu, Rui Yao 0010, Xueshuo Xie, Tao Li 0022
BIBM2
2024 FRCNet: Frequency and Region Consistency for Semi-supervised Medical Image Segmentation
Along He, Tao Li 0022, Yanlin Wu, Ke Zou, Huazhu Fu
MICCAI (8)1
2024 Open-Set Semi-supervised Medical Image Classification with Learnable Prototypes and Outlier Filter
Along He, Tao Li 0022, Yitian Zhao, Junyong Zhao, Huazhu Fu
MICCAI (11)1
2024 Resfusion: Denoising Diffusion Probabilistic Models for Image Restoration Based on Prior Residual Noise
abstract
Recently, research on denoising diffusion models has expanded its application to the field of image restoration. Traditional diffusion-based image restoration methods utilize degraded images as conditional input to effectively guide the reverse generation process, without modifying the original denoising diffusion process. However, since the degraded images already include low-frequency information, starting from Gaussian white noise will result in increased sampling steps. We propose Resfusion, a general framework that incorporates the residual term into the diffusion forward process, starting the reverse process directly from the noisy degraded images. The form of our inference process is consistent with the DDPM. We introduced a weighted residual noise, named resnoise, as the prediction target and explicitly provide the quantitative relationship between the residual term and the noise term in resnoise. By leveraging a smooth equivalence transformation, Resfusion determine the optimal acceleration step and maintains the integrity of existing noise schedules, unifying the training and inference processes. The experimental results demonstrate that Resfusion exhibits competitive performance on ISTD dataset, LOL dataset and Raindrop dataset with only five sampling steps. Furthermore, Resfusion can be easily applied to image generation and emerges with strong versatility. Our code and model are available at https://github.com/nkicsl/Resfusion.
Zhenning Shi, Haoshuai Zheng, Changsheng Dong, Bin Pan, Xueshuo Xie, Along He, Tao Li 0002, Huazhu Fu
NeurIPS7
2024 NKUT: Dataset and Benchmark for Pediatric Mandibular Wisdom Teeth Segmentation
abstract
Germectomy is a common surgery in pediatric dentistry to prevent the potential dangers caused by impacted mandibular wisdom teeth. Segmentation of mandibular wisdom teeth is a crucial step in surgery planning. However, manually segmenting teeth and bones from 3D volumes is time-consuming and may cause delays in treatment. Deep learning based medical image segmentation methods have demonstrated the potential to reduce the burden of manual annotations, but they still require a lot of well-annotated data for training. In this paper, we initially curated a Cone Beam Computed Tomography (CBCT) dataset, NKUT, for the segmentation of pediatric mandibular wisdom teeth. This marks the first publicly available dataset in this domain. Second, we propose a semantic separation scale-specific feature fusion network named WTNet, which introduces two branches to address the teeth and bones segmentation tasks. In WTNet, We design a Input Enhancement (IE) block and a Teeth-Bones Feature Separation (TBFS) block to solve the feature confusions and semantic-blur problems in our task. Experimental results suggest that WTNet performs better on NKUT compared to previous state-of-the-art segmentation methods (such as TransUnet), with a maximum DSC lead of nearly 16%.
Zhenhuan Zhou, Along He, Xitao Que, Kai Wang 0001, Rui Yao 0010, Tao Li 0022
IEEE J. Biomed. Health Informatics3
2024 Bilateral Supervision Network for Semi-Supervised Medical Image Segmentation
abstract
Massive high-quality annotated data is required by fully-supervised learning, which is difficult to obtain for image segmentation since the pixel-level annotation is expensive, especially for medical image segmentation tasks that need domain knowledge. As an alternative solution, semi-supervised learning (SSL) can effectively alleviate the dependence on the annotated samples by leveraging abundant unlabeled samples. Among the SSL methods, mean-teacher (MT) is the most popular one. However, in MT, teacher model's weights are completely determined by student model's weights, which will lead to the training bottleneck at the late training stages. Besides, only pixel-wise consistency is applied for unlabeled data, which ignores the category information and is susceptible to noise. In this paper, we propose a bilateral supervision network with bilateral exponential moving average (bilateral-EMA), named BSNet to overcome these issues. On the one hand, both the student and teacher models are trained on labeled data, and then their weights are updated with the bilateral-EMA, and thus the two models can learn from each other. On the other hand, pseudo labels are used to perform bilateral supervision for unlabeled data. Moreover, for enhancing the supervision, we adopt adversarial learning to enforce the network generate more reliable pseudo labels for unlabeled data. We conduct extensive experiments on three datasets to evaluate the proposed BSNet, and results show that BSNet can improve the semi-supervised segmentation performance by a large margin and surpass other state-of-the-art SSL methods.
Along He, Tao Li 0022, Juncheng Yan, Kai Wang 0001, Huazhu Fu
IEEE Trans. Medical Imaging1
2023 H2Former: An Efficient Hierarchical Hybrid Transformer for Medical Image Segmentation
abstract
Accurate medical image segmentation is of great significance for computer aided diagnosis. Although methods based on convolutional neural networks (CNNs) have achieved good results, it is weak to model the long-range dependencies, which is very important for segmentation task to build global context dependencies. The Transformers can establish long-range dependencies among pixels by self-attention, providing a supplement to the local convolution. In addition, multi-scale feature fusion and feature selection are crucial for medical image segmentation tasks, which is ignored by Transformers. However, it is challenging to directly apply self-attention to CNNs due to the quadratic computational complexity for high-resolution feature maps. Therefore, to integrate the merits of CNNs, multi-scale channel attention and Transformers, we propose an efficient hierarchical hybrid vision Transformer (H2Former) for medical image segmentation. With these merits, the model can be data-efficient for limited medical data regime. The experimental results show that our approach exceeds previous Transformer, CNNs and hybrid methods on three 2D and two 3D medical image segmentation tasks. Moreover, it keeps computational efficiency in model parameters, FLOPs and inference time. For example, H2Former outperforms TransUNet by 2.29% in IoU score on KVASIR-SEG dataset with 30.77% parameters and 59.23% FLOPs.
Along He, Kai Wang 0001, Tao Li 0022, Chengkun Du, Huazhu Fu
IEEE Trans. Medical Imaging1
2022 Progressive Multiscale Consistent Network for Multiclass Fundus Lesion Segmentation
abstract
Effectively integrating multi-scale information is of considerable significance for the challenging multi-class segmentation of fundus lesions because different lesions vary significantly in scales and shapes. Several methods have been proposed to successfully handle the multi-scale object segmentation. However, two issues are not considered in previous studies. The first is the lack of interaction between adjacent feature levels, and this will lead to the deviation of high-level features from low-level features and the loss of detailed cues. The second is the conflict between the low-level and high-level features, this occurs because they learn different scales of features, thereby confusing the model and decreasing the accuracy of the final prediction. In this paper, we propose a progressive multi-scale consistent network (PMCNet) that integrates the proposed progressive feature fusion (PFF) block and dynamic attention block (DAB) to address the aforementioned issues. Specifically, PFF block progressively integrates multi-scale features from adjacent encoding layers, facilitating feature learning of each layer by aggregating fine-grained details and high-level semantics. As features at different scales should be consistent, DAB is designed to dynamically learn the attentive cues from the fused features at different scales, thus aiming to smooth the essential conflicts existing in multi-scale features. The two proposed PFF and DAB blocks can be integrated with the off-the-shelf backbone networks to address the two issues of multi-scale and feature inconsistency in the multi-class segmentation of fundus lesions, which will produce better feature representation in the feature space. Experimental results on three public datasets indicate that the proposed method is more effective than recent state-of-the-art methods.
Along He, Kai Wang 0001, Tao Li 0022, Wang Bo, Hong Kang, Huazhu Fu
IEEE Trans. Medical Imaging1
2021 CABNet: Category Attention Block for Imbalanced Diabetic Retinopathy Grading
abstract
Diabetic Retinopathy (DR) grading is challenging due to the presence of intra-class variations, small lesions and imbalanced data distributions. The key for solving fine-grained DR grading is to find more discriminative features corresponding to subtle visual differences, such as microaneurysms, hemorrhages and soft exudates. However, small lesions are quite difficult to identify using traditional convolutional neural networks (CNNs), and an imbalanced DR data distribution will cause the model to pay too much attention to DR grades with more samples, greatly affecting the final grading performance. In this article, we focus on developing an attention module to address these issues. Specifically, for imbalanced DR data distributions, we propose a novel Category Attention Block (CAB), which explores more discriminative region-wise features for each DR grade and treats each category equally. In order to capture more detailed small lesion information, we also propose the Global Attention Block (GAB), which can exploit detailed and class-agnostic global attention feature maps for fundus images. By aggregating the attention blocks with a backbone network, the CABNet is constructed for DR grading. The attention blocks can be applied to a wide range of backbone networks and trained efficiently in an end-to-end manner. Comprehensive experiments are conducted on three publicly available datasets, showing that CABNet produces significant performance improvements for existing state-of-the-art deep architectures with few additional parameters and achieves the state-of-the-art results for DR grading. Code and models will be available at https://github.com/he2016012996/CABnet.
Along He, Tao Li 0022, Kai Wang 0001, Huazhu Fu
IEEE Trans. Medical Imaging1