Wei Xu 0038

dblp:32/1213-38 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-4705-7189ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EduArt-Bench: A Benchmark and Lightweight Scoring Calibration for K-12 Art Education
Zekun Huang, Wei Xu 0038
AIED (3)5
2026 RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-Aware Learning
abstract
Image-to-3D generation aims to predict a geometrically and perceptually plausible 3D model from a single 2D image. Conventional approaches typically follow a cascaded pipeline: initially generating multi-view projections from the single input image through view synthesis, followed by optimizing 3D geometry and appearance strictly using these projections. However, such deterministic optimization neglects epistemic uncertainty from imperfectly generated data, particularly due to limited observations and inconsistent content. To address this issue, we propose an uncertainty-aware optimization framework that explicitly models and mitigates epistemic uncertainty, leading to more robust and reliable 3D generation. For epistemic uncertainty arising from incomplete viewpoint coverage, we employ a progressive sampling strategy that sinusoidally varies camera elevations and progressively integrates diverse viewpoints into training, enhancing viewpoint coverage and stabilizing optimization. For epistemic uncertainty caused by the deterministic optimization on the noisy and inconsistent generated multi-view frames, we estimate an uncertainty map from the discrepancies between two independently optimized Gaussian models. This map is incorporated into uncertainty-aware regularization, adaptively adjusting loss weights to suppress unreliable supervision. Furthermore, we provide a theoretical analysis of uncertainty-aware optimization by deriving a probabilistic upper bound on the expected generation error, providing insights into its effectiveness. Extensive experiments demonstrate that our method significantly reduces artifacts and inconsistencies, leading to higher-quality 3D generation. More visual results are available at our website https://rigi3d.github.io/.
Zhedong Zheng, Wei Xu 0038, Ping Liu 0004
IEEE Trans. Image Process.3
2025 DD-RobustBench: An Adversarial Robustness Benchmark for Dataset Distillation
abstract
Dataset distillation techniques have revolutionized the way of utilizing large datasets by compressing them into smaller, yet highly effective subsets that preserve the original datasets' accuracy. However, while these methods have proven effective in reducing data size and training times, the robustness of these distilled datasets against adversarial attacks remains underexplored. This vulnerability poses significant risks, particularly in security-sensitive applications. To address this critical gap, we introduce DD-RobustBench, a novel and comprehensive benchmark specifically designed to evaluate the adversarial robustness of distilled datasets. Our benchmark is the most extensive of its kind and integrates a variety of dataset distillation techniques, including recent advancements such as TESLA, DREAM, SRe2L, and D4M, which have shown promise in enhancing model performance. DD-RobustBench also rigorously tests these datasets against a diverse array of adversarial attack methods to ensure broad applicability. Our evaluations cover a wide spectrum of datasets, including but not limited to, the widely used ImageNet-1K. This allows us to assess the robustness of distilled datasets in scenarios mirroring real-world applications. Furthermore, our detailed quantitative analysis investigates how different components involved in the distillation process, such as data augmentation, downsampling, and clustering, affect dataset robustness. Our findings provide critical insights into which techniques enhance or weaken the resilience of distilled datasets against adversarial threats, offering valuable guidelines for developing more robust distillation methods in the future. Through DD-RobustBench, we aim not only to benchmark but also to push the boundaries of dataset distillation research by highlighting areas for improvement and suggesting pathways for future innovations in creating datasets that are not only compact and efficient but also secure and resilient to adversarial challenges. The implementation details and essential instructions are available on DD-RobustBench.
Yifan Wu 0037, Jiawei Du 0002, Ping Liu 0004, Yuewei Lin, Wei Xu 0038, Wenqing Cheng
IEEE Trans. Image Process.5
2024 Piano Transcription with Harmonic Attention
abstract
Automatic Music Transcription (AMT) aims to convert music audio into digital sheet music. Piano transcription is a popular but challenging subtask of AMT. For every piano pitch, the harmonic structure is fixed in the frequency domain, while the Transformer based on self-attention has great potential to extract features in the long sequence. In this paper, we propose piano harmonic attention, a mask self-attention, for better capturing harmonic features. The mask matrix is designed with the harmonic prior to pre-modeling the harmonic structure during calculating attention scores. To verify its effectiveness, we append the harmonic attention-based Transformer after every convolutional neural network block of the High-resolution piano transcription system. The evaluation results on the MAESTRO dataset show that the proposed model achieves comprehensive improvements over the baseline, with a note F1 score of 97.33%, which is comparable to the state-of-the-art system.
Ruimin Wu, Xianke Wang, Wei Xu 0038, Wenqing Cheng
ICASSP4
2024 Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance
abstract
Existing text-to-image editing methods tend to excel either in rigid or non-rigid editing but encounter challenges when combining both, resulting in misaligned outputs with the provided text prompts. In addition, integrating reference images for control remains challenging. To address these issues, we present a versatile image editing framework capable of executing both rigid and non-rigid edits, guided by either textual prompts or reference images. We leverage a dual-path injection scheme to handle diverse editing scenarios and introduce an integrated self-attention mechanism for fusion of appearance and structural information. To mitigate potential visual artifacts, we further employ latent fusion techniques to adjust intermediate latents. Compared to previous work, our approach represents a significant advance in achieving precise and versatile image editing. Comprehensive experiments validate the efficacy of our method, showcasing competitive or superior results in text-based editing and appearance transfer tasks, encompassing both rigid and non-rigid settings.
Ping Liu 0004, Wei Xu 0038
ICME3
2024 CNN-Transformer Ensemble: Advancing Visual Piano Transcription with Global and Local Features
abstract
Piano transcription is a significant problem in music information retrieval, which aims to infer the note sequence from recorded music signals. Recently, visual piano transcription heavily relies on Convolutional Neural Networks (CNNs) for local feature detection, facing challenges in capturing global features of the piano keyboard. Therefore, we propose a novel visual piano transcription model that integrates CNN and transformer branches. The CNN branch has inductive bias and extracts local features of the piano key, while the transformer branch aggregates global features over the keyboard. The proposed model integrates local features and global features at different resolutions, enhancing the performance of the visual transcription model. Finally, our model achieves a remarkable F1-score of 92.35% on the OMAPS2 dataset and attains state-of-the-art results on other datasets. This substantiates the model’s innovative approach and its potential to advance visual piano transcription within music information retrieval.
Xianke Wang, Ruimin Wu, Wei Xu 0038
IJCNN5
2024 Active Speaker Detection in Fisheye Meeting Scenes with Scene Spatial Spectrums
abstract
Active Speaker Detection (ASD) plays a crucial role in scene understanding tasks by determining whether an on-screen person in a given scene is speaking.In this work, to address the ASD in the context of multi-party roundtable meetings, we propose a novel approach that incorporates the fusion of spatial information of the scenes.To leverage the multiple data sources of the scenes, our method involves generating audio spatial spectrum heatmaps from the multi-channel audio and integrating them with the panoramic images.Additionally, we propose the novel FisheyeMeeting dataset, which combines fisheye panoramic video recordings with muti-channel audio captured from a six-channel circular microphone array.By enabling the multi-modal model to capture audio-visual cues in multi-party meeting scenes, our approach achieves an impressive 89.11% mAP on the FisheyeMeeting dataset.Notably, this outperforms the current SOTA methods by a significant 2.3% mAP improvement.
Xinghao Huang, Long Rao, Wei Xu 0038, Wenqing Cheng
INTERSPEECH4
2024 A Two-Stage Audio-Visual Fusion Piano Transcription Model Based on the Attention Mechanism
abstract
Piano transcription is a significant problem in the field of music information retrieval, aiming to obtain symbolic representations of music from captured audio or visual signals. Previous research has mainly focused on single-modal transcription methods using either audio or visual information, yet there is a small number of studies based on audio-visual fusion. To leverage the complementary advantages of both modalities and achieve higher transcription accuracy, we propose a two-stage audio-visual fusion piano transcription model based on the attention mechanism, utilizing both audio and visual information from the piano performance. In the first stage, we propose an audio model and a visual model. The audio model utilizes frequency domain sparse attention to capture harmonic relationships in the frequency domain, while the visual model includes both CNN and Transformer branches to merge local and global features at different resolutions. In the second stage, we employ cross-attention to learn the correlations between different modalities and the temporal relationships of the sequences. Experimental results on the OMAPS2 dataset show that our model achieves an F1-score of 98.60%, demonstrating significant improvement compared with the single-modal transcription models.
Xianke Wang, Ruimin Wu, Wei Xu 0038, Wenqing Cheng
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Text-Guided Eyeglasses Manipulation With Spatial Constraints
abstract
Virtual try-on of eyeglasses involves placing eyeglasses of different shapes and styles onto a face image without physically trying them on. While existing methods have shown impressive results, the variety of eyeglasses styles is limited and the interactions are not always intuitive or efficient. To address these limitations, we propose GlassesCLIP, a text-guided eyeglasses manipulation method with spatial constraints, which allows for control of the eyeglasses shape and style based on a binary mask and text, respectively. Specifically, we introduce a mask encoder to extract mask conditions and a modulation module that enables simultaneous injection of text and mask conditions. This design allows for fine-grained control of the eyeglasses' appearance based on both textual descriptions and spatial constraints. Our approach includes a disentangled mapper and a decoupling strategy that preserves irrelevant areas, resulting in better local editing. We employ a two-stage training scheme to handle the different convergence speeds of the various modality conditions, successfully controlling both the shape and style of eyeglasses. Extensive comparison experiments and ablation analyses demonstrate the effectiveness of our approach in achieving diverse eyeglasses styles while preserving irrelevant areas.
Ping Liu 0004, Jingen Liu, Wei Xu 0038
IEEE Trans. Multim.4
2023 A Dual-Path Approach for Gaze Following in Fisheye Meeting Scenes
Long Rao, Xinghao Huang, Shipeng Cai, Wei Xu 0038, Wenqing Cheng
PRCV (5)5
2023 MusicYOLO: A Vision-Based Framework for Automatic Singing Transcription
abstract
Automatic singing transcription (AST), which refers to the process of inferring the onset, offset, and pitch from the singing audio, is of great significance in music information retrieval. Most AST models use the convolutional neural network to extract spectral features and predict the onset and offset moments separately. The frame-level probabilities are inferred first, and then the note-level transcription results are obtained through post-processing. In this paper, a new AST framework called MusicYOLO is proposed, which obtains the note-level transcription results directly. The onset/offset detection is based on the object detection model YOLOX, and the pitch labeling is completed by a spectrogram peak search. Compared with previous methods, the MusicYOLO detects note objects rather than isolated onset/offset moments, thus greatly enhancing the transcription performance. On the sight-singing vocal dataset (SSVD) established in this paper, the MusicYOLO achieves an 84.60% transcription F1-score, which is the state-of-the-art method.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 A Multi-Stage Automatic Evaluation System for Sight-Singing
abstract
Sight-singing exercises are a fundamental part of music education. In this paper, we present an objective and complete automatic evaluation system for sight-singing, which has two critical stages: note transcription and note alignment. In the first stage, we use an onset detector based on the convolutional recurrent neural network (CRNN) for note segmentation and the pitch extractor described in (Kimet al.2018) for note labeling. In the second stage, an alignment algorithm based on relative pitch modeling is proposed. Due to the lack of datasets for sight-singing note alignment and the overall system evaluation, we construct the sight-singing vocal dataset (SSVD). Each module of the system and the entire system are tested on this dataset. The onset detector achieves an F-measure of 90.61%, and the stages of note transcription and note alignment achieve an F-measure of 88.42% and 94.79%, respectively. In addition, we propose an objective criterion for the sight-singing evaluation system. Based on this criterion, our automatic sight-singing system achieves an F-measure of 77.95% on the SSVD dataset.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
IEEE Trans. Multim.4
2022 Musicyolo: A Sight-Singing Onset/Offset Detection Framework Based on Object Detection Instead of Spectrum Frames
abstract
In this paper, we propose MusicYOLO based on object detection to detect the onset and offset in singing for the first time. The onset of the vocal is not as stable and clear as that of musical instruments, which makes the frame-based onset/offset detection methods often not work well. Compared with the previous onset/offset detection methods, MusicYOLO detects the whole note object in the spectrogram image instead of transient frame features around onset/offset, improving the onset/offset detection performance significantly. The experiment results show that the MusicYOLO framework has obtained a 94.16% F1 score of onset detection and a 91.35% F1 score of offset detection on the ISMIR2014 dataset, which proves that MusicYOLO is the state-of-the-art onset/offset detection framework for singing situation.
Xianke Wang, Wei Xu 0038, Wenqing Cheng
ICASSP2
2022 SingMaster: A Sight-singing Evaluation System of "Shoot and Sing" Based on Smartphone
abstract
Based on the smart phone, this paper integrates OMR (Optical Music Recognition) with sight-singing evaluation, and develops a "shoot and sing" practice APP called SingMaster. This system is mainly composed of three modules: OMR, evaluation and user interface. The OMR module converts the score photographed in the real scene into a note reference sequence. The sight-sing evaluation module first completes the note transcription of the sound spectrum through onset detection and pitch extraction, then aligns the transcribed note sequence with the reference sequence, and performs the evaluation. Finally, the evaluation results are visually fed back to the practitioners through the user interface module. It can provide guidance for practitioners at any time, any place and on any score instead of a real teacher.
Wei Xu 0038, Lijie Luo, Xianke Wang
ACM Multimedia1
2021 Transition-Aware: A More Robust Approach for Piano Transcription
abstract
Piano transcription is a classic problem in music information retrieval. More and more transcription methods based on deep learning have been proposed in recent years. In 2019, Google Brain published a larger piano transcription dataset, MAESTRO. On this dataset, Onsets and Frames transcription approach proposed by Hawthorne achieved a stunning onset F1 score of 94.73%. Unlike the annotation method of Onsets and Frames, Transition-aware model presented in this paper annotates the attack process of piano signals called atack transition in multiple frames, instead of only marking the onset frame. In this way, the piano signals around onset time are taken into account, enabling the detection of piano onset more stable and robust. Transition-aware achieves a higher transcription F1 score than Onsets and Frames on MAESTRO dataset and MAPS dataset, reducing many extra note detection errors. This indicates that Transition-aware approach has better generalization ability on different datasets.
Xianke Wang, Wei Xu 0038, Juanting Liu, Wenqing Cheng
DAFx2
2021 An Audio-Visual Fusion Piano Transcription Approach Based on Strategy
abstract
Piano transcription is a fundamental problem in the field of music information retrieval. At present, a large number of transcriptional studies are mainly based on audio or video, yet there is a small number of discussion based on audio-visual fusion. In this paper, a piano transcription model based on strategy fusion is proposed, in which the transcription results of the video model are used to assist audio transcription. Due to the lack of datasets currently used for audio-visual fusion, the OMAPS data set is proposed in this paper. Meanwhile, our strategy fusion model achieves a 92.07% F1 score on OMAPS dataset. The transcription model based on feature fusion is also compared with the one based on strategy fusion. The experiment results show that the transcription model based on strategy fusion achieves better results than the one based on feature fusion.
Xianke Wang, Wei Xu 0038, Juanting Liu, Wenqing Cheng
DAFx2
2012 A simple way to improve lookup performance in KAD
abstract
KAD is the largest DHT system with several million simultaneous users. The dynamics of peer participation which is called churn affects the performance of lookup operations in P2P systems, since some individual peers in the routing tables might be missing or stale. In this paper, we performed a simple way to improve lookup performance in KAD, by taking highly available contacts as lookup entries instead of stale ones. We track highly available peers in KAD by a special designed crawler. When a stale contact is encountered in the lookup process, the closest XOR-distance highly available peer of the target will be found to replace the stale contact. The measurement study, compared with the normal lookup process, shows that it is much effective.
Fangfang Liao, Wei Xu 0038, Wenqing Cheng
APCC3
2006 A Transaction-Aware Coordination Protocol for Web Services Composition
Wei Xu 0038, Wenqing Cheng, Wei Liu 0004
WISE1