EDBT 2026 Demo / reviewers in the wild / expert
Anfeng Xu
dblp:180/5858
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the GlobeabstractWe present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the RAIL license at https://github.com/tiantiaf0627/voxlect. Tiantian Feng, Anfeng Xu, Xuan Shi, Thanathai Lertpetchpun, Yoonjeong Lee, Dani Byrd, Shri Narayanan |
KDD (1) | 3 |
| 2025 | Data Efficient Child-Adult Speaker Diarization with Simulated ConversationsabstractAutomating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies "who spoke when", is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available. Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Catherine Lord, Shri Narayanan |
ICASSP | 1 |
| 2025 | Effective Integration of KAN for Keyword SpottingabstractKeyword spotting (KWS) is an important speech processing component for smart devices with voice assistance capability. In this paper, we investigate if Kolmogorov-Arnold Networks (KAN) can be used to enhance the performance of KWS. We explore various approaches to integrate KAN for a model architecture based on 1D Convolutional Neural Networks (CNN). We find that KAN is effective at modeling high-level features in lower-dimensional spaces, resulting in improved KWS performance when integrated appropriately. The findings shed light on understanding KAN for speech processing tasks and on other modalities for future researchers. Anfeng Xu, Biqiao Zhang, Shuyu Kong, Yiteng Huang, Sangeeta Srivastava, Ming Sun 0013 |
ICASSP | 1 |
| 2025 | Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
Tiantian Feng, Anfeng Xu, Xuan Shi, Somer Bishop, Shri Narayanan |
INTERSPEECH | 2 |
| 2025 | Examining Test-Time Adaptation for Personalized Child Speech Recognition
Zhonghao Shi, Xuan Shi, Anfeng Xu, Tiantian Feng, Harshvardhan Srivastava, Shri Narayanan, Maja J. Mataric |
INTERSPEECH | 3 |
| 2025 | Large Language Models based ASR Error Correction for Child Conversations
Anfeng Xu, Tiantian Feng, So Hyun Kim, Somer Bishop, Catherine Lord, Shri Narayanan |
INTERSPEECH | 1 |
| 2024 | Audio-Visual Child-Adult Speaker Classification in Dyadic InteractionsabstractInteractions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer scale and robustness across diverse and wide-ranging conditions. Identifying the speech segments belonging to the child is a critical step in such modeling. Conventional child-adult speaker classification typically relies on audio modeling approaches, overlooking visual signals that convey speech articulation information, such as lip motion. Building on the foundation of an audio-only child-adult speaker classification pipeline, we propose incorporating visual cues through active speaker detection and visual processing models. Our framework involves video preprocessing, utterance-level child-adult speaker detection, and late fusion of modality-specific predictions. We demonstrate from extensive experiments that a visually aided classification pipeline enhances the accuracy and robustness of the classification. We show relative improvements of 2.38% and 3.97% in F1 macro score when one face and two faces are visible, respectively. Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Shri Narayanan |
ICASSP | 1 |
| 2024 | Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
Anfeng Xu, Tiantian Feng, Lue Shen, Helen Tager-Flusberg, Shri Narayanan |
INTERSPEECH | 1 |
| 2023 | Understanding Spoken Language Development of Children with ASD Using Pre-trained Speech Embeddings
Anfeng Xu, Rajat Hebbar, Rimita Lahiri, Tiantian Feng, Lindsay Butler, Lue Shen, Helen Tager-Flusberg, Shri Narayanan |
INTERSPEECH | 1 |
| 2023 | MM-AU: Towards Multimodal Understanding of Advertisement VideosabstractAdvertisement videos (ads) play an integral part in the domain of Internet e-commerce, as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise narrative structures. The narrative structures of advertisements involve several elements like reasoning about the broad content (topic and the underlying message) and examining fine-grained details involving the transition of perceived tone due to the sequence of events and interaction among characters. In this work, to facilitate the understanding of advertisements along the three dimensions of topic categorization, perceived tone transition, and social message detection, we introduce a multimodal multilingual benchmark called MM-AU comprised of 8.4 K videos (147hrs) curated from multiple web-based sources. We explore multiple zero-shot reasoning baselines through the application of large language models on the ads transcripts. Further, we demonstrate that leveraging signals from multiple modalities, including audio, video, and text, in multimodal transformer-based supervised models leads to improved performance compared to unimodal approaches. Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shri Narayanan |
ACM Multimedia | 5 |
| 2020 | A Visual Tracking Method Based on an Adaptive Overlapping Correlation Filter for Robotic Real-Time Cognitive ImagingabstractComputer vision is a very important research direction in the cognitive computing field. Robots encounter various target-tracking problems with computer vision systems. Robust scale estimation is an important issue in tracking algorithms. Most of the available methods have difficulty addressing even reasonable changes of scale in complex videos. In this paper, we propose a visual tracking method based on robust scale estimation, which uses a discriminant correlation filter based on a time-dependent scale-space filter and an adaptive cross-correlation filter. The tracker uses separate essential filters for sample migration and scale estimation. Furthermore, the built-in scale estimation method can be introduced into other tracking algorithms. We validate the proposed method on the UAV123 dataset. The results of comparison experiments with the traditional correlation filter tracking method demonstrate that the proposed method improves the success rate and tracking accuracy while controlling the computational complexity; its success rate measured by the area under the curve is 0.638, while at a location error precision of 20%, it is 0.649. Yihua Lan, Pianpian Ma, Anfeng Xu |
Wirel. Commun. Mob. Comput. | 3 |
| 2018 | Cross domain adaptation by learning partially shared classifiers and weighting source data points in the shared subspaces
Anfeng Xu, Sunny Chughtai |
Neural Comput. Appl. | 2 |