EDBT 2026 Demo / reviewers in the wild / expert
Yan Yan 0022
dblp:13/3953-22
· DBLP profile ↗
25ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-6344-136XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BACFormer: A robust boundary-aware transformer for medical image segmentation
Zhiyong Huang 0004, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Knowl. Based Syst. | 8 |
| 2026 | Bridging teacher-student representation domains for manifold-aware knowledge distillation
Shuai Miao, Zhiyong Huang 0004, Daidi Zhong, Mingyang Hou, Penghao Jia, Yan Yan 0022, Yushi Liu 0001 |
Knowl. Based Syst. | 7 |
| 2026 | Eliminating domain-related confounding factors in cross-domain one-shot medical image segmentation via causal inference
Mingyang Hou, Zhiyong Huang 0004, Daidi Zhong, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Medical Image Anal. | 7 |
| 2026 | Integrating visual and language cues via state space models for medical image segmentation
Mingyang Hou, Zhiyong Huang 0004, Daidi Zhong, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Neural Networks | 7 |
| 2026 | Enhancing Radiography-Report foundation model via Multi-View masked contrastive learning
Daidi Zhong, Zhiyong Huang 0004, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Pattern Recognit. | 9 |
| 2026 | Overcoming Limitations in One-Shot Semantic Segmentation via Adaptive Visual-Text Guided Prototype Relationship OptimizationabstractThe scarcity of high-quality annotated data in medical imaging significantly constrains the performance of deep learning-based segmentation models. While few-shot medical image segmentation (FSMIS) has emerged as a promising solution, existing methods exhibit critical limitations when handling scenarios with inter-class similarity between foreground-background regions and intra-class heterogeneity within foreground objects. Current prototype-based approaches focus primarily on the holistic extraction of the prototype from support images, failing to distinguish subtle anatomical variations and complex feature representations effectively. The AVT-ProNet features three innovative components: 1) An Adaptive Visual-Text Prototype Generation (AVPG) module leveraging CLIP’s cross-modal guide capabilities through adaptive prompting strategies; 2) a graph-based multiregion prototyping relationship optimization (GMPRO) module establishing structural relationships between decomposed subregion prototypes via graph neural networks; 3) a foreground-background prototyping contrast learning (FBPCL) strategy implementing dual-space optimization through inter-class separation and intra-class compactness. The synergistic integration of multi-modal guidance, structural relationship modeling, and contrastive prototype refinement enables our framework to overcome existing limitations in FSMIS. Comprehensive evaluations across multiple clinical scenarios (CHAOS, SABS, and CMR datasets under diverse training configurations) demonstrate superior performance over state-of-the-art approaches, including PANet, CAT-Net, DMAP, and recent PAMI baselines, Source code is available at https://github.com/394481125/AVT-ProNet. Mingyang Hou, Zhiyong Huang 0004, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | GDPA: A Parameter-Efficient Gated Dual-Path Spatiotemporal Adapter for Surgical Workflow RecognitionabstractAutomated surgical workflow recognition is a key enabler for context-aware computer-assisted surgery. Due to the scarcity of labeled surgical videos, training models from scratch or fully fine-tuning large vision architectures is often resourceintensive and can lead to overfitting. To alleviate this bottleneck, we insert a lightweight spatiotemporal adapter module into an image-pretrained backbone to inject temporal reasoning capability into the backbone and fine-tune the entire network. This not only preserves the backbone's original strong spatial representation capability, but also enables the model to learn such issues as the characteristic long-range temporal dependencies in surgical scenarios and the fine-grained interactions between instruments and organs, thereby achieving robust transfer to surgical video understanding tasks. Specifically, we propose a parameter-efficient Gated Dual-Path Spatiotemporal Adapter (GDPA) for surgical workflow recognition. GDPA is a lightweight module inserted into a frozen image-pretrained backbone to enable parallel spatial and temporal modeling. Within each GDPA block, the spatial and temporal branches operate in parallel, and their outputs are fused by an adaptive gating mechanism. This gating mechanism adaptively allocates the contributions of the two branches according to phase-dependent requirements for fine-grained spatial semantics and long-range temporal cues. Compared with previous methods, GDPA requires only a small number of trainable parameters and low computational cost. By retaining the backbone's pretrained visual knowledge and training only lightweight adapters in an end-to-end manner, it achieves higher accuracy even in data-limited scenarios. We evaluate our method on the Cholec80 and AutoLaparo surgical video benchmarks. The results show that even with a parameter budget far below that of full fine-tuning, GDPA still achieves competitive or even superior performance. Shimei Wang, Yongde Guo, Huasong Shao, Yushi Liu 0001, Jing Xiong 0001, Lei Wang 0029, Yan Yan 0022 |
BIBM | 8 |
| 2025 | Spectral Adaptive Hypergraphs for Skeleton-Based Action RecognitionabstractGraph based models have significantly advanced skeleton-based action recognition, yet most methods rely on pairwise connections and struggle to capture higher order, rhythm consistent dependencies. We introduce the Spectral Adaptive Hypergraph Convolution Network (SA-HyperGCN), which constructs hyperedges in the frequency domain to model coordinated multi joint relations. By transforming joint trajectories into spectral representations and adaptively grouping joints via low-frequency similarity, SA-HyperGCN discovers actiondriven structures beyond spatial adjacency. The resulting spectral hypergraph is fused with the static skeleton topology through multi-head hypergraph convolution for expressive message passing. We further incorporate virtual joints into the hypergraph and apply a topology-aware regularization that discourages overly similar geometric configurations. This constraint helps maintain diversity among the learned virtual joint structures and stabilizes training. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120 and NW-UCLA show that SA-HyperGCN consistently outperforms strong GCN and hypergraph baselines, validating the effectiveness of spectral modeling and topologyguided regularization. Yongde Guo, Huasong Shao, Yushi Liu 0001, Shimei Wang, Jing Xiong 0001, Lei Wang 0029, Yan Yan 0022 |
BIBM | 8 |
| 2025 | A Multi-Domain Patch-Differentiated Transformer for vehicle re-identification
Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001, Daming Sun, Hans Gregersen |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | WTSF-ReID: Depth-driven Window-oriented Token Selection and Fusion for multi-modality vehicle re-identification with knowledge consistency constraint
Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | SCFMUNet: A fusion architecture based on multi-scale state space model and channel attention for medical image segmentation
Zhiyong Huang 0004, Mingyang Hou, Shiyao Zhou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
Neural Networks | 7 |
| 2025 | Feature-Tuning Hierarchical Transformer via token communication and sample aggregation constraint for object re-identification
Zhiyong Huang 0004, Mingyang Hou, Jiaming Pei, Yan Yan 0022, Yushi Liu 0001, Daming Sun |
Neural Networks | 5 |
| 2025 | Representation Selective Coupling via Token Sparsification for Multi-Spectral Object Re-IdentificationabstractTo tackle the challenge of single-spectral object re-identification in complex and dynamic lighting scenarios, multi-spectral object re-identification, which integrates visible light and infrared information, is gradually taking the lead. Nevertheless, the significant heterogeneity across spectra causes formidable obstacles for this task. Most existing approaches alleviate inter-spectral disparities by amalgamating representations from different spectra, ignoring the selection of spectrum-specific crucial information. To address this issue, we propose a novel Representation Selective Coupling Network (RSCNet) for multi-spectral object re-identification. Specifically, we design an Attention-Fourier Token Sparsification (AFTS) module to adaptively sparse and join tokens from multi-spectral images in the attention domain and Fourier domain. This not only preserves spectrum-specific crucial information but also reduces inter-spectral gaps by selective coupling of multi-spectral representation. Meanwhile, to further align multi-spectral information and guide the model to learn more discriminative representation, we propose an Information Unification Constraint (IUC) learning strategy. Both feature-level information constraint and distribution-level information constraint are simultaneously deployed in IUC. Finally, we conduct extensive experiments on three multi-spectral object re-identification benchmarks, and the experimental results verify the effectiveness of our proposed method. Zhiyong Huang 0004, Mingyang Hou, Jiaming Pei, Yan Yan 0022, Yushi Liu 0001, Daming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MCECF: A Multiscale Complementary Enhanced Context Fusion Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) holds significant research value in remote sensing (RS) image processing. In recent years, many researchers have achieved remarkable results in RSCD tasks using methods based on convolutional neural networks (CNNs) or Transformers. Considering the limited receptive field of CNN models and the high computational cost of Transformers, many researchers have combined the two approaches, yielding promising results. However, most current RSCD-based models focus solely on change and temporal information, overlooking their complementary relationship. Additionally, some multiscale feature fusion methods emphasize enhancing individual scales while neglecting the correlations between different scales. To address the above issues, we propose a multiscale complementary enhanced context fusion (MCECF) network. The network first introduces a global-local context aggregation module (GLCAM) to capture global-local context information while extracting multilevel feature maps. Subsequently, a complementary enhancement difference module (CEDM) is employed to complementarily aggregate the captured change and temporal information of bi-temporal RS image features. To fully leverage the correlations between multiscale features, a progressive decoder comprising a supervised spatial attention (SSA) mechanism and a multiscale complementary enhanced fusion module (MCEFM) was developed. Moreover, to tackle the disparity between changed and unchanged regions, a dual-branch dynamic attention fusion module (DAFM) was designed to enhance the model’s adaptability to diverse scenarios. We conducted comparative experiments on five RSCD datasets against nine state-of-the-art (SOTA) methods, and the results confirmed the effectiveness of the proposed MCECF in RSCD tasks. Our code will be made available athttps://github.com/kakuqikaduo/MCECF Zhiyong Huang 0004, Hongjiang Qiu, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | MSCD-VM-UNet: A Vision Mamba Combining Multi-Scale Global and Local Feature Extraction With Cross-Domain Feature Fusion for Medical Image SegmentationabstractAccurate segmentation of tissues and lesions is essential for diagnosis and treatment. State Space Models (SSMs) have gained attention for their linear complexity and ability to model long-range dependencies. However, the existing Mamba architecture relies on direct skip connections, which limits its ability to integrate multi-scale and multi-level features and handle boundary details effectively. To address these limitations, we propose the MSCD-VM-UNet architecture, which incorporates three novel modules: the Spatial Group Multi-Scale Attention Module (SGMAM), the Cross-Domain Feature Fusion Module (CDFFM), and the Attention-Based Feature Injection Module (ABFIM). The SGMAM captures multi-scale global and local information and adaptively adjusts feature importance to highlight key regions while suppressing noise. The CDFFM enhances boundary and detail handling by aligning semantic features from both the frequency and spatial domains. The ABFIM utilizes attention mechanisms to adaptively fuse and weigh features from different scales and semantics, promoting feature collaboration and improving the model's robustness in complex tasks. Experiments on multiple datasets show that these modules significantly enhance the accuracy of MSCD-VM-UNet, setting a new benchmark for medical image segmentation. Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation ModelabstractHuman activity recognition is closely related to human health and is a hot research topic. With the continuous development of AI technology, large models have shown great potential in various training tasks. However, there is limited analysis of human sensor motion data. In this study, we fine-tuned large models on five motion datasets, designed an adaptation layer suitable for time series data to extract action representations fully, and proposed a novel HAR framework. We conducted action recognition experiments on five foundation models and compared them with ten classical algorithm models. The experimental results show that the accuracy of action recognition in small-parameter pre-trained large models can reach 98.8%, indicating significant research potential. Fuhai Xiong, Junxian Wang, Yushi Liu 0001, Kamen Ivanov, Lei Wang 0029, Yan Yan 0022 |
BIBM | 7 |
| 2024 | CSwT-SR: Conv-Swin Transformer for Blind Remote Sensing Image Super-Resolution With Amplitude-Phase Learning and Structural Detail Alternating LearningabstractImage super-resolution (SR) stands as a pivotal process in the domains of image processing and computer vision, finding diverse applications in film, television, photography, surveillance, medical imaging, and remote sensing. In the context of remote sensing images (RSIs), the inherent challenge arises from low spatial resolution caused by factors such as sensor noise, orbit height, and weather conditions, necessitating SR reconstruction. An evident limitation of prevailing methods lies in their dependence on idealized fixed degradation models, which fail to capture the intricate degradation processes unique to remote sensing scenes. In response to these constraints, this article introduces an innovative blind image super-resolution reconstruction method tailored for remote sensing images. The proposed approach integrates convolution with a transformer and incorporates an amplitude-phase learning module (ALM) to comprehensively capture local and long-range dependencies while enhancing frequency information. The iterative optimization strategy refines texture information by carefully balancing structural and detail elements. Key contributions include a holistic approach to remote sensing image SR, ALM integration for precise feature representation, and the introduction of a patch-based frequency loss mechanism for evaluating frequency-domain features. Rigorous experiments demonstrate that compared with other state-of-the-art (SOTA) methods, the proposed algorithm delivers SR results with exceptional visual perception quality across three distinct remote sensing datasets. Mingyang Hou, Zhiyong Huang 0004, Yan Yan 0022, Yunlan Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Noninvasive Blood Glucose Monitoring Using Spatiotemporal ECG and PPG Feature Fusion and Weight-Based Choquet Integral Multimodel Approachabstractchange of blood glucose (BG) level stimulates the autonomic nervous system leading to variation in both human's electrocardiogram (ECG) and photoplethysmogram (PPG). In this article, we aimed to construct a novel multimodal framework based on ECG and PPG signal fusion to establish a universal BG monitoring model. This is proposed as a spatiotemporal decision fusion strategy that uses weight-based Choquet integral for BG monitoring. Specifically, the multimodal framework performs three-level fusion. First, ECG and PPG signals are collected and coupled into different pools. Second, the temporal statistical features and spatial morphological features in the ECG and PPG signals are extracted through numerical analysis and residual networks, respectively. Furthermore, the suitable temporal statistical features are determined with three feature selection techniques, and the spatial morphological features are compressed by deep neural networks (DNNs). Lastly, weight-based Choquet integral multimodel fusion is integrated for coupling different BG monitoring algorithms based on the temporal statistical features and spatial morphological features. To verify the feasibility of the model, a total of 103 days of ECG and PPG signals encompassing 21 participants were collected in this article. The BG levels of participants ranged between 2.2 and 21.8 mmol/L. The results obtained show that the proposed model has excellent BG monitoring performance with a root-mean-square error (RMSE) of 1.49 mmol/L, mean absolute relative difference (MARD) of 13.42%, and Zone A + B of 99.49% in tenfold cross-validation. Therefore, we conclude that the proposed fusion approach for BG monitoring has potentials in practical applications of diabetes management. Jingzhen Li, Olatunji Mumini Omisore, Yuhang Liu 0007, Huajie Tang, Pengfei Ao, Yan Yan 0022, Lei Wang 0029, Ze-dong Nie |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Learning Topological Representation of Sensor Network with Persistent Homology in HCI SystemsabstractHand gesture and movement analysis is a crucial learning task in Human-computer interaction (HCI) applications. Sensor-based HCI systems simultaneously capture the information with multiple locations to track the coordination of different regions of muscles. Based on the fact that there exists a temporal correlation between the regions, the connectivity analysis of sensor signals builds a network. The graph-based approach for analyzing the sensor network has provided novel insight into the learning in HCI, which has not been broadly investigated in hand gesture recognition tasks. This work proposes a topological representation learning scheme as a graph-based approach for sensor network analysis. Through investigation of the topological properties with persistent homology, the spatial-temporal characteristics are well described to build recognition models. Experiments on the NinaPro DB-2, DB-4, DB-5, and DB-7 datasets with sensor networks built with sEMG signal and IMU signal demonstrate exceptional performance of the proposed topological approach. The topological features are effective in graph representation learning with sensor networks used in hand gesture recognition. The proposed work provides a novel learning scheme in HCI systems and human-in-the-loop studies. Yan Yan 0022, Chengdong Li, Jing Xiong 0001, Lei Wang 0029 |
BIBM | 1 |
| 2023 | Persistence Landscape-based Topological Data Analysis for Personalized Arrhythmia ClassificationabstractHuman ECG sensing signals can be regarded as the observed variables of the human heart’s nonlinear dynamic system, which can effectively reflect the state changes of the heart system. They can be used for heart health monitoring and related disease identification. Due to the robust chaos, nonlinearity, and complexity of ECG signals, it is challenging to express them by standard features. Therefore, this paper proposes a nonlinear topological data analysis method to model ECG signals and extract nonlinear features for ECG anomaly detection. Firstly, we use the time delay embedding approach to map the ECG time series to the topological space for phase space reconstruction to form the ECG point cloud. Then, based on the point cloud information in space, the persistent homology method was used to construct the topological imprint of ECG data. Finally, the persistence landscape in the topological impression was extracted as the topological feature of the ECG signal for ECG anomaly detection. With only 20% of the total training dataset, it achieves a 100% accuracy for normal heartbeats, 98.75% for ventricular beats, 95.88% for supra-ventricular moments, and 91.97% for fusion beats. Thus the method can be trained for a single individual, allowing for personalized analysis systems. With the present study, TDA could be a valuable tool for biomedical signal analysis, with potential application in customized data processing. Yushi Liu 0001, Lei Wang 0029, Yan Yan 0022 |
BSN | 3 |
| 2023 | Topological Nonlinear Analysis of Dynamical Systems in Wearable Sensor-Based Human Physical Activity InferenceabstractThis work presents a topological nonlinear analysis approach for dynamical system measurements, frequently appearing in sensor-based inference tasks in human physical activity analysis. Traditional approaches to dynamical modeling included linear and nonlinear methods with specific representational abilities and some drawbacks. A novel approach we investigate is using topological descriptors of the shape of the dynamical attractor to represent the nature of dynamics. The proposed framework has three essential advantages compared to previous approaches: 1) with nonlinear phase space reconstruction, the dynamics descriptor is derived from the observation time series without any statistical assumption; 2) with the topological data analysis technique, the phase space topological properties are described in an intrinsic multiresolution analytical way, which brings novel information compared to traditional phase-space modeling techniques; 3) with different types of measurement sensing signals, the proposed approach shows stability in activities state inference. We illustrate our idea with the physical activity recognition tasks with wearable sensors, where the topological characteristics of reconstructed phase state space show strong representational ability for activity type inference. Yan Yan 0022, Yi-Chun Huang, Yushi Liu 0001, Jing Xiong 0001, Lei Wang 0029 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2022 | A Pupil Segmentation Framework with Masked Image Modeling Enhanced Swin-TransformerabstractDetecting pupil from the image is critical in human-machine interaction and biomedical computing applications, which is supposed to be an actual image segmentation problem. Recently developed deep learning models provide a variety of novel approaches to the pupil segmentation task. However, dataset preparation and annotation acquirement to build pupil image datasets are labor-intensive and time-consuming. The shortage of labeled samples restricted the improvement of deep learning models. In this work, we use a mask image modeling mechanism to learn the latent representation from limited data samples, which significantly helps train deep models. Further, we propose a novel pupil segmentation model based on the recently proposed Swin-Transformer to validate the improvement validity of the mask mechanism. The proposed computational framework achieves better performance on the pupil segmentation tasks based on the LPW dataset through comparison experiments with other related deep learning models. The proposed framework is a promising solution for pupil segmentation and detection in small-sample learning applications. Yongde Guo, LuYu Tang, Jing Xiong 0001, Yan Yan 0022 |
BIBM | 6 |
| 2022 | Deep Transfer Learning with Graph Neural Network for Sensor-Based Human Activity RecognitionabstractThe sensor-based human activity recognition (HAR) in mobile application scenarios is often confronted with variation in sensing modalities and deficiencies in annotated samples. To address these two challenging problems, we devised a graph-inspired deep learning approach that uses data from human-body mounted wearable sensors. As a step toward a complete HAR solution, the proposed method was further used to build a deep transfer learning model. Specifically, we present a multi-layer residual structure involving graph convolutional neural network (ResGCNN) toward the sensor-based HAR tasks, namely the HAR-ResGCNN approach. Experimental results on the PAMAP2 and mHealth data sets demonstrate that our ResGCNN is effective at capturing the characteristics of actions with comparable results compared to other sensor-based HAR models (with an average accuracy of 98.18% and 99.07%, respectively). More importantly, the parameter-based transfer learning experiments using the ResGCNN model show excellent transferability and small sample learning ability, which is a promising solution in sensor-based HAR applications. Tianzheng Liao, Yushi Liu 0001, Kamen Ivanov, Jing Xiong 0001, Yan Yan 0022 |
BIBM | 6 |
| 2020 | Towards adequate prediction of prediabetes using spatiotemporal ECG and EEG feature analysis and weight-based multi-model approach
Tobore Igbe, Abhishek Kandwal, Jingzhen Li, Yan Yan 0022, Olatunji Mumini Omisore, Efetobore Enitan, Sinan Li, Yuhang Liu 0007, Lei Wang 0029, Ze-dong Nie |
Knowl. Based Syst. | 4 |
| 2015 | A restricted Boltzmann machine based two-lead electrocardiography classificationabstractAn restricted Boltzmann machine learning algorithm were proposed in the two-lead heart beat classification problem. ECG classification is a complex pattern recognition problem. The unsupervised learning algorithm of restricted Boltzmann machine is ideal in mining the massive unlabelled ECG wave beats collected in the heart healthcare monitoring applications. A restricted Boltzmann machine (RBM) is a generative stochastic artificial neural network that can learn a probability distribution over its set of inputs. In this paper a deep belief network was constructed and the RBM based algorithm was used in the classification problem. Under the recommended twelve classes by the ANSI/AAMI EC57: 1998/(R)2008 standard as the waveform labels, the algorithm was evaluated on the two-lead ECG dataset of MIT-BIH and gets the performance with accuracy of 98.829%. The proposed algorithm performed well in the two-lead ECG classification problem, which could be generalized to multi-lead unsupervised ECG classification or detection problems. Yan Yan 0022, Xinbing Qin, Yige Wu, Jianping Fan 0002, Lei Wang 0029 |
BSN | 1 |