Yu Wang 0108

dblp:02/5889-108 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
17since 2021 · last 2026
0009-0002-9663-2231ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Cross-modal multitask learning for automated quantitative characterization of infrastructure airhole defects
Yu Wang 0108, Yingchao Dai, Xiaodong Gan, Zhengtao Yang, Zhou Wu 0001
Expert Syst. Appl.1
2026 Attention-Guided Spatiotemporal Information Fusion of GB-SAR Data for Landslide Displacement Prediction
abstract
Accurate landslide displacement prediction is a key component of IoT-enabled landslide monitoring and warning support frameworks, supporting risk-informed warning analysis and risk-informed decision-making. However, precise forecasting remains challenging due to the non-stationary characteristics of displacement data and the complex spatiotemporal correlations among monitoring points. To this end, this article proposes a novel attention-guided spatiotemporal fusion framework, named VMAG (Variational Mode Decomposition and Multi-head Attention-based GRU), for accurate multi-step landslide displacement prediction. Specifically, a differential fluctuation sequence is first constructed using Ground-Based Synthetic Aperture Radar (GB-SAR) displacement observations to enhance the perceptibility of displacement mutations. Then, based on the inherent characteristics of the time series, the Variational Mode Decomposition (VMD) algorithm is applied to decompose the sequence into multi-scale components to mitigate non-stationarity. Subsequently, a multi-head attention mechanism is employed to dynamically extract spatial dependencies between the target node and reference monitoring nodes, which are then fed into a Gated Recurrent Unit (GRU) network to capture temporal evolutions. Experimental results on datasets from the Moshi Gully (MSG) landslide demonstrate that the proposed VMAG significantly outperforms mainstream benchmarks. For instance, the Root Mean Square Error (RMSE) is reduced to 1.6960 mm, and the Mean Absolute Percentage Error (MAPE) achieves 0.63%, showing superior accuracy compared to ANN, LSTM, and standard GRU models.
Xiaodong Gan, Yingchao Dai, Yu Wang 0108, Reza Malekian, Zhou Wu 0001
IEEE Internet Things J.3
2025 Take the Bull by the Horns: Learning to Segment Hard Samples
abstract
Medical image segmentation is vital for clinical applications, with hard samples playing a key role in segmentation accuracy. We propose an effective image segmentation framework that includes mechanisms for identifying and segmenting hard samples. It derives a novel image segmentation paradigm: 1) Learning to identify hard samples: automatically selecting inherent hard samples from different datasets, and 2) Learning to segment hard samples: achieving the segmentation of hard samples through effective feature augmentation on dedicated networks. We name our method ‘Learning to Segment hard samples’ (L2S). The hard sample identification module comprises a backbone model and a classifier, which dynamically uncovers inherent dataset patterns. The hard sample segmentation module utilizes the diffusion process for feature augmentation and incorporates a more sophisticated segmentation network to achieve precise segmentation. We justify our motivation through solid theoretical analysis and extensive experiments. Evaluations across various modalities show that our L2S outperforms other SOTA methods, particularly by substantially improving the segmentation accuracy of hard samples. On ISIC dataset, our L2S improves the Dice score on hard samples and overall segmentation by 8.97% and 1.01%, respectively, compared to SOTA methods. The code is available at https://github.com/TqlYuanGie/L2S.
Jingyu Kong, Yu Wang 0108, Yuping Duan
CVPR3
2025 Free Scale 2D-3D Regional Retrieval Based on Cross Modal Information Fusion
abstract
2D-3D cross modal retrieval (CMR) aims to retrieve query image matching points from a 3D reference map. Existing classical CMR datasets and methods commonly support database-based retrieval only, i.e., the point cloud retrieval results are fixed-scale geometric surfaces. The failure to consider geometric regions and information scales fundamentally limits the practical deployment of CMR in engineering systems that require dynamic spatial reasoning, such as autonomous navigation or three-dimensional industrial measurement. In this article, we introduce a new benchmark called cross modal regional retrieval, which extends the classic CMR to allow the free retrieval of associated regions within the point cloud from images. Toward this, a multiview training paradigm is proposed in the training phase, which enables the model to identify occluded points in the region based on a single view. Autoencoders are utilized to learn the mapping of fusion features from a single view to multiple views. We also convert the image retrieval task within the scene cloud into a point classification task in the image to implement global free retrieval. The information fusion and guidance provided by the global point cloud enhances the capability of image cross-modal retrieval. To match the input patterns of the model, we propose a method for constructing datasets from three benchmark sources. Extensive experiments demonstrate that our method achieves state-of-the-art performance compared to existing methods for 2D-3D cross modal regional retrieval.
Zhou Wu 0001, Yu Wang 0108, Hongtuo Qi, Liang Feng 0001, Jiepeng Liu
IEEE Trans. Ind. Informatics2
2025 RemixFormer++: A Multi-Modal Transformer Model for Precision Skin Tumor Differential Diagnosis With Memory-Efficient Attention
abstract
Diagnosing malignant skin tumors accurately at an early stage can be challenging due to ambiguous and even confusing visual characteristics displayed by various categories of skin tumors. To improve diagnosis precision, all available clinical data from multiple sources, particularly clinical images, dermoscopy images, and medical history, could be considered. Aligning with clinical practice, we propose a novel Transformer model, named RemixFormer++ that consists of a clinical image branch, a dermoscopy image branch, and a metadata branch. Given the unique characteristics inherent in clinical and dermoscopy images, specialized attention strategies are adopted for each type. Clinical images are processed through a top-down architecture, capturing both localized lesion details and global contextual information. Conversely, dermoscopy images undergo a bottom-up processing with two-level hierarchical encoders, designed to pinpoint fine-grained structural and textural features. A dedicated metadata branch seamlessly integrates non-visual information by encoding relevant patient data. Fusing features from three branches substantially boosts disease classification accuracy. RemixFormer++ demonstrates exceptional performance on four single-modality datasets (PAD-UFES-20, ISIC 2017/2018/2019). Compared with the previous best method using a public multi-modal Derm7pt dataset, we achieved an absolute 5.3% increase in averaged F1 and 1.2% in accuracy for the classification of five skin tumors. Furthermore, using a large-scale in-house dataset of 10,351 patients with the twelve most common skin tumors, our method obtained an overall classification accuracy of 92.6%. These promising results, on par or better with the performance of 191 dermatologists through a comprehensive reader study, evidently imply the potential clinical usability of our method.
Kai Huang 0008, Lianzhen Zhong, Yuan Gao 0017, Wei Liu 0127, Yanjie Zhou, Wenchao Guo, Yuanqiang Zou, Yuping Duan, Le Lu 0001, Yu Wang 0108
IEEE Trans. Medical Imaging13
2024 Evolutionary Multitasking With Centralized Learning for Large-Scale Combinatorial Multiobjective Optimization
abstract
Evolutionary multitasking (EMT) has attracted much attention in the community of evolutionary computation recently. It intends to improve the performance of evolutionary optimization on multiple problems via knowledge learning and transfer across them while the optimization processes progress online. Existing EMT paradigms can be classified as explicit EMT (EEMT) and implicit EMT (IEMT) according to the mechanisms adopted in the knowledge transfer. With additional knowledge learning and transfer modules, the EEMT often brings flexible algorithmic designs and effective knowledge transfer against the IEMT. However, most of the existing EEMT studies are designed for continuous optimization problems. Due to the difficulty of learning problem-specific mappings across combinatorial optimization problems, EEMT for combinatorial optimization is still in the nascent stage. Furthermore, it is worth noting that, with the growing number of tasks in today’s real-world applications and the enlarged number of decision variables in each of the tasks, learning mappings across tasks becomes more challenging. Keeping the above in mind, this paper presents a novel EEMT algorithm with centralized learning for solving the large-scale and multi-objective combinatorial optimization in many-task manner, in which knowledge transfer across tasks is conducted based on a centralized learning model, instead of task-specific mappings which are required in existing EEMT studies. To investigate the performance of the proposed centralized learning assisted EEMT, comprehensive empirical studies have been conducted on the large-scale and multi-objective knapsack problems. Lastly, the efficacy of our proposed method is further validated on a real-world combinatorial optimization application.
Wei Zhou 0001, Yu Wang 0108, Min Li 0056, Liang Feng 0001, Kay Chen Tan
IEEE Trans. Evol. Comput.3
2023 Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation
abstract
Recent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identically applying them in multiple feature stages. Our preliminary experiments show that with the strong long-range dependency modeling capacity of Transformer, simply concatenating the two modality features and feeding them to vanilla Transformers for feature fusion can distinctly benefit the performance but at a cost of heavy computation. Through further empirical analysis, we find that attention dependencies learned in Transformer in different stages exhibit completely different properties: global query-independent dependency in the low-level stages and semantic-specific dependency in the high-level stages. Motivated by the observations, we propose two Transformer variants: i) Context-Sharing Transformer (CST) that learns the global-shared contextual information within image frames with a lightweight computation. ii) Semantic Gathering-Scattering Transformer (SGST) that models the semantic correlation separately for the foreground and background and reduces the computation cost with a soft token merging mechanism. We apply CST and SGST for low-level and high-level feature fusions, respectively, formulating a level-isomerous Transformer framework for ZVOS task. Compared with the baseline that uses vanilla Transformers for multi-stage fusion, ours significantly increase the speed by 13× and achieves new state-of-the-art ZVOS performance. Code is available at https://github.com/DLUT-yyc/Isomer.
Yifan Wang 0004, Lijun Wang 0001, Xiaoqi Zhao 0003, Huchuan Lu, Yu Wang 0108, Weibo Su, Lei Zhang 0006
ICCV6
2023 A Novel Multi-task Model Imitating Dermatologists for Accurate Differential Diagnosis of Skin Diseases in Clinical Images
Yan-Jie Zhou, Wei Liu 0127, Yuan Gao 0017, Le Lu 0001, Yuping Duan, Na Jin, Xiaoyong Man, Yu Wang 0108
MICCAI (6)11
2023 Fast Vehicle Routing via Knowledge Transfer in a Reproducing Kernel Hilbert Space
abstract
Vehicle routing problems (VRPs) are essential in logistics. In the literature, many exact and heuristic optimization algorithms have been proposed to solve the VRPs. These traditional approaches, however, generally start the optimization from scratch and ignore the experiences of solving related VRPs, which may lead to unnecessary computational costs in searching repeated problems and reduce the efficiency of vehicle routing. Recently, transfer optimization (TO) has been presented to speed up vehicle routing by reusing the knowledge learned from similarly solved VRPs. However, existing TO methods build connections across VRPs in a low-dimensional Euclidean space, which has limited modeling ability in the cases of having nonlinear correlations. Keeping this in mind, this article presents a study of TO equipped with the kernel method for fast vehicle routing. In contrast to existing TO methods, in this work, the learning of connections across VRPs for knowledge transfer is conducted in a reproducing kernel Hilbert space (RKHS), which thus has greater modeling capacity in nonlinear customer relationships between VPRs. To evaluate the performance of the proposed method, comprehensive empirical studies have been conducted using well-known VRP benchmarks, against existing state-of-the-art TO methods for vehicle routing. Finally, a well-known real-world VRP application given by a routing company (Jingdong), namely, the package delivery problem (PDP), is investigated to further assess the efficacy of our proposed method.
Liang Feng 0001, Min Li 0056, Yu Wang 0108, Zexuan Zhu 0001, Kay Chen Tan
IEEE Trans. Syst. Man Cybern. Syst.4
2022 RemixFormer: A Transformer Model for Precision Skin Tumor Differential Diagnosis via Multi-modal Imaging and Non-imaging Data
Yuan Gao 0017, Wei Liu 0127, Kai Huang 0008, Le Lu 0001, Xiaosong Wang 0001, Xian-Sheng Hua 0001, Yu Wang 0108
MICCAI (3)9
2022 Multi-space evolutionary search with dynamic resource allocation strategy for large-scale optimization
Qingxia Shang, Junwei Dong, Yaqing Hou, Yu Wang 0108, Min Li 0056, Liang Feng 0001
Neural Comput. Appl.5
2021 Mobile-based Clock Drawing Test for Detecting Early Signs of Dementia
abstract
Dementia is one of the major causes of disability and dependency among older people. Early detection is the key for preserving the quality of life of the patients and reducing caring costs. The Clock Drawing Test (CDT) is commonly used by clinicians to screen for early signs of dementia. We build an automated CDT that runs on mobile platforms, enabling convenient and frequent self-monitoring and testing at minimal costs. Our system combines both a spatial-temporal approach and a purely image-based deep learning approach to analyze and evaluate the hand-drawn clocks based on established clinical criteria. Our system produces scores that are highly correlated with expert human raters.
Hongchao Jiang, Yanci Zhang, Jun Ji, Yu Wang 0108, Ying Chi, Chunyan Miao
AAAI5
2021 SpineOne: A One-Stage Detection Framework for Degenerative Discs and Vertebrae
abstract
Spinal degeneration plagues many elders, office workers, and even the younger generations. Effective pharmic or surgical interventions can help relieve degenerative spine conditions. However, the traditional diagnosis procedure is often too laborious. Clinical experts need to localize discs and vertebrae as a preliminary step of pathological diagnosis. Machine learning systems have been developed to aid this procedure generally following a two-stage methodology: first perform anatomical localization, then pathological classification. Towards more efficient and accurate diagnosis, we propose a one-stage detection framework termed SpineOne to simultaneously localize and classify degenerative discs and vertebrae from magnetic resonance imaging (MRI) slices. SpineOne is built upon the following three key techniques: 1) a new design of the keypoint heatmap to facilitate simultaneous keypoint localization and classification; 2) the use of attention modules to better differentiate the representations between discs and vertebrae; and 3) a novel gradient-guided objective association mechanism to associate multiple learning objectives at the later training stage. Empirical results on the Spinal Disease Intelligent Diagnosis Tianchi Competition (SDID-TC) dataset of 550 exams demonstrate that our approach surpasses existing methods by a large margin.
Jiabo He, Wei Liu 0127, Yu Wang 0108, Xingjun Ma, Xian-Sheng Hua 0001
BIBM3
2021 Clustering-Augmented Multi-instance Learning for Neural Relation Extraction
Qi Zhang 0001, Siliang Tang, Jinquan Sun, Yu Wang 0108, Lei Zhang 0006
ECIR (2)4
2021 Towards Parkinson's Disease Prognosis Using Self-Supervised Learning and Anomaly Detection
abstract
Parkinson’s disease (PD) is a chronic disease with a high risk of incidence after the age of 60 and is a problem for many countries facing an aging population. Current works have mainly focused on supervised learning using data collected from various sensors to differentiate between PD and healthy subjects. However, such supervised methods are not ideal for prognosis where there are no labels (i.e., we do not know in advance which subjects will develop PD in the future). We propose to tackle the problem as a semi-supervised anomaly detection task, where we model the physiological patterns of healthy subjects instead. A self-supervised learning technique first learns a good representation of the sensor signals. The representations are then adapted to capture inter-class patterns for anomaly detection. Evaluation on a large-scale PD dataset shows that our approach can learn discriminative features.
Hongchao Jiang, Wei Yang Bryan Lim, Jer Shyuan Ng, Yu Wang 0108, Ying Chi, Chunyan Miao
ICASSP4
2021 Development and validation of a practical instrument for evaluating players' familiarity with exergames
Hao Zhang 0049, Di Wang 0004, Yu Wang 0108, Ying Chi, Chunyan Miao
Int. J. Hum. Comput. Stud.3
2021 Learned snakes for 3D image segmentation
Lihong Guo, Yueyun Liu, Yu Wang 0108, Yuping Duan, Xue-Cheng Tai
Signal Process.3
2020 Explainable and Argumentation-based Decision Making with Qualitative Preferences for Diagnostics and Prognostics of Alzheimer's Disease
abstract
Argumentation has gained traction as a formalism to make more transparent decisions and provide formal explanations recently. In this paper, we present an argumentation-based approach to decision making that can support modelling and automated reasoning about complex qualitative preferences and offer dialogical explanations for the decisions made. We first propose Qualitative Preference Decision Frameworks (QPDFs). In a QPDF, we use contextual priority to represent the relative importance of combinations of goals in different contexts and define associated strategies for deriving decision preferences based on prioritized goal combinations. To automate the decision computation, we map QPDFs to Assumption-based Argumentation (ABA) frameworks so that we can utilize existing ABA argumentative engines for our implementation. We implemented our approach for two tasks, diagnostics and prognostics of Alzheimer's Disease (AD), and evaluated it with real-world datasets. For each task, one of our models achieves the highest accuracy and good precision and recall for all classes compared to common machine learning models. Moreover, we study how to formalize argumentation dialogues that give contrastive, focused and selected explanations for the most preferred decisions selected in given contexts.
Zhiqi Shen 0001, Benny Toh Hsiang Tan, Jing Jih Chin, Cyril Leung, Yu Wang 0108, Ying Chi, Chunyan Miao
KR6
2020 Landmarks Detection with Anatomical Constraints for Total Hip Arthroplasty Preoperative Measurements
Wei Liu 0127, Yu Wang 0108, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001
MICCAI (4)2
2019 Learned Full-Sampling Reconstruction
Weilin Cheng, Yu Wang 0108, Ying Chi, Xuansong Xie, Yuping Duan
MICCAI (5)2