Wentao Lei

dblp:274/2405 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling
abstract
Molecular representation learning, a cornerstone for downstream tasks like molecular captioning and molecular property prediction, heavily relies on Graph Neural Networks (GNN). However, GNN suffers from the over-smoothing problem, where node-level features collapse in deep GNN layers. While existing feature projection methods with cross-attention have been introduced to mitigate this issue, they still perform poorly in deep features. This motivated our exploration of using Mamba as an alternative projector for its ability to handle complex sequences. However, we observe that while Mamba excels at preserving global topological information from deep layers, it neglects fine-grained details in shallow layers. The capabilities of Mamba and cross-attention exhibit a global-local trade-off. To resolve this critical global-local trade-off, we propose Hierarchical and Structure-Aware Network (HSA-Net), a novel framework with two modules that enables a hierarchical feature projection and fusion. Firstly, a Hierarchical Adaptive Projector (HAP) module is introduced to process features from different graph layers. It learns to dynamically switch between a cross-attention projector for shallow layers and a structure-aware Graph-Mamba projector for deep layers, producing high-quality, multi-level features. Secondly, to adaptively merge these multi-level features, we design a Source-Aware Fusion (SAF) module, which flexibly selects fusion experts based on the characteristics of the aggregation features, ensuring a precise and effective final representation fusion. Extensive experiments demonstrate that our HSA-Net framework quantitatively and qualitatively outperforms current state-of-the-art (SOTA) methods.
Zihang Shao, Wentao Lei, Wencai Ye
AAAI2
2025 Teaching Others Teaches Yourself: Semi-supervised Ensembled Pseudo-labeling Method for Image Classification
abstract
Semi-supervised methods have recently received significant attention in deep learning because they are able to reduce the dependence on labeled data while ensuring good performance. The pseudo-label based method has been widely used as a classic semi-supervised method but still suffers from the confirmation bias problem, which causes significant damage to the model’s training process. To solve this problem, we propose an ensembled semi-supervised framework, which can effectively improve the quality of pseudo-labels by ensembling the prediction results of multiple models to generate pseudo-labels. Simultaneously, we innovatively design a Ensembled Divergence Promotion (EDP) method to increase the diversity among different models for better ensembling results. Extensive experiments are conducted on CIFAR-10 and CIFAR-100, showing superior performance of the proposed method to the state-of-the-art (SOTA) semi-supervised method. Meanwhile, extensive ablation experiments and visualization results are conducted to prove the effectiveness of our method.
Wentao Lei, Li Liu 0036
ICASSP1
2025 Multi-Modal Rhythmic Generative Model for Chinese Cued Speech Gestures Generation
abstract
Cued speech (CS) is a novel visual coding system, which combines lip reading with several specific hand codings to help hearing-impaired people to communicate effectively. This work focuses on the audio/text-driven CS gestures (i.e., continuous lip and hand gestures movements) generation. Previous work used template-based statistical methods for the French CS generation. However, these methods are fragile since they need careful hand-crafted pre-processing to fit models, resulting in poor robustness. Furthermore, the natural rhythm in generated CS gesture sequences, which is essential for a coding system of spoken languages, was overlooked in prior studies. To solve the above-mentioned problems, we innovatively propose a two-branched rhythmic CS gesture generation framework, which contains a multi-modal adversarial semantic generator (MASG) to generate accurate multi-modal CS gestures (i.e., lip, hand shape and hand position movements), and an audio-driven rhythm generator (ARG) to extract the rhythm information. Moreover, we design a new Gesture Audio Difference (GAD) metric to evaluate the rhythm coherence considering the issue of asynchrony between CS hand gestures and lip movements. Extensive experimental results are presented on two datasets of two tasks (a CS dataset named MCCS-2024 and a co-speech TED dataset) with comprehensive ablation analysis and user study, demonstrating the effectiveness of our method. The code and dataset with multi-modal annotations were made public at https://mccs-2024.github.io/.
Li Liu 0036, Wentao Lei, Wenwu Wang 0001
ICASSP2
2025 FFD: Fine-Finger Diffusion Model for Music to Fine-grained Finger Dance Generation
Boya Dong, Wentao Lei
INTERSPEECH2
2024 Bridge to Non-Barrier Communication: Gloss-Prompted Fine-Grained Cued Speech Gesture Generation with Diffusion Model
Wentao Lei, Li Liu 0036
IJCAI1
2023 Understanding Strategies and Challenges of Conducting Daily Data Analysis (DDA) Among Blind and Low-vision People
abstract
Being able to analyze and derive insights from data, which we call Daily Data Analysis (DDA), is an increasingly important skill in everyday life. While the accessibility community has explored ways to make data more accessible to blind and low-vision (BLV) people, little is known about how BLV people perform DDA. Knowing BLV people’s strategies and challenges in DDA would allow the community to make DDA more accessible to them. Toward this goal, we conducted a mixed-methods study of interviews and think-aloud sessions with BLV people (N=16). Our study revealed five key approaches for DDA (i.e., overview obtaining, column comparison, key statistics identification, note-taking, and data validation) and the associated challenges. We discussed the implications of our findings and highlighted potential directions to make DDA more accessible for BLV people.
Chutian Jiang, Wentao Lei, Emily Kuang, Teng Han, Mingming Fan 0001
ASSETS2
2023 Spatio-Temporal Structure Consistency for Semi-Supervised Medical Image Classification
abstract
Intelligent medical diagnosis has shown remarkable progress on the large-scale datasets with full annotations. However, very few labeled images are available due to significantly expensive annotations by experts. To efficiently leverage abundant unlabeled data, we propose a novel Spatio-Temporal Structure Consistent (STSC) learning framework to combine both spatial and temporal structure consistency. Specifically, a gram matrix is derived to capture the structural similarity among different training samples in the representation space. At the spatial level, our framework explicitly enforces the consistency of structural similarity among different samples under perturbations. At the temporal level, we desire to maintain the consistency of the structural similarity in different training iterations by digging out the stable sub-structures in a relation graph. Experiments on two medical image datasets (i.e., ISIC 2018 and ChestX-ray14) show that our method outperforms state-of-the-art Semi-Supervised Learning (SSL) methods. Furthermore, extensive qualitative analysis on the Gram matrices and heatmaps by Grad-CAM are presented to validate the effectiveness of our method.
Wentao Lei, Lei Liu 0049, Li Liu 0036
ICASSP1
2022 "I Shake The Package To Check If It's Mine": A Study of Package Fetching Practices and Challenges of Blind and Low Vision People in China
abstract
With about 230 million packages delivered per day in 2020, fetching packages has become a routine for many city dwellers in China. When fetching packages, people usually need to go to collection sites of their apartment complexes or a KuaiDiGui, an increasingly popular type of self-service package pickup machine. However, little is known whether such processes are accessible to blind and low vision (BLV) city dwellers. We interviewed BLV people (N=20) living in a large metropolitan area in China to understand their practices and challenges of fetching packages. Our findings show that participants encountered difficulties in finding the collection site and localizing and recognizing their packages. When fetching packages from KuaiDiGuis, they had difficulty in identifying the correct KuaiDiGui, interacting with its touch screen, navigating the complex on-screen workflow, and opening the target compartment. We discuss design considerations to make the package fetching process more accessible to the BLV community.
Wentao Lei, Mingming Fan 0001, Juliann Thang
CHI1
2020 Semi-Supervised Active Learning for COVID-19 Lung Ultrasound Multi-symptom Classification
abstract
Ultrasound (US) is a non-invasive yet effective medical diagnostic imaging technique for the COVID-19 global pandemic. However, due to complex feature behaviors and expensive annotations of US images, it is difficult to apply Artificial Intelligence (AI) assisting approaches for the lung's multi-symptom (multi-label) classification. To overcome these difficulties, we propose a novel semi-supervised Two-Stream Active Learning (TSAL) method to model complicated features and reduce labeling costs in an iterative manner. The core component of TSAL is the multi-label learning mechanism, in which label correlation information is used to design a multi-label margin (MLM) strategy and a confidence validation for automatically selecting informative samples and confident labels. In this framework, a multi-symptom multi-label (MSML) classification network is proposed to learn discriminative features of lung symptoms, and a human-machine interaction (HMI) is exploited to confirm the final annotations that are used to fine-tune MSML. Moreover, a novel lung US dataset named COVID19-LUSMS is built, currently containing 71 clinical patients with 6,836 images sampled from 678 videos. Experimental evaluations show that TSAL can achieve superior performance to the baseline and the state-of-the-art using only 20% data. Qualitatively, visualization of the attention map confirms a good consistency between the model prediction and the clinical knowledge.
Lei Liu 0049, Wentao Lei, Li Liu 0036, Yongfang Luo
ICTAI2