EDBT 2026 Demo / reviewers in the wild / expert
Qingming Tang
dblp:79/8473
· DBLP profile ↗
24ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-2670-4917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effective Techniques for Scaling Audio Encoder PretrainingabstractThis work presents advancements in audio pretraining objectives designed to generate semantically rich embeddings, capable of addressing a wide range of audio-related tasks. Despite significant progress in the field, current methods often emphasize full fine-tuning in downstream applications, which can obscure the true potential of pretrained audio encoders. In this study, we present an audio encoder that achieves stateof-the-art (SOTA) performances in both fine-tuning and linear probing, utilizing a carefully curated set of pragmatic techniques. Building on previous research, we incorporate masked prediction and introduce SpecAug within a curriculum masking strategy at the patch level, which progressively increases training difficulty, along with a mask-aware position bias. To comprehensively assess the encoder’s capabilities, we examine the impact of scaling both the dataset size and model capacity, conducting linear probing evaluations while keeping the encoder frozen as well as full fine-tuning. Our model demonstrates superior performance compared to recent SOTA methods across various downstream tasks. Additionally, we explore the potential of tokenizing the resulting audio embeddings for use as discrete inputs, enhancing our understanding of the model’s capabilities. Byeonggeun Kim, Andrew Bydlon, Qingming Tang, Huy Phan, Chieh-Chi Kao |
ICASSP | 3 |
| 2025 | IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion ModelingabstractText-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio generation. Despite achieving high audio fidelity, they incur significant inference latency due to the slow diffusion sampling process. MAGNET, a mask-based model operating on discrete tokens, addresses slow inference through iterative mask-based parallel decoding. However, its audio quality still lags behind that of diffusion-based models. In this work, we introduce IMPACT, a text-to-audio generation framework that achieves high performance in audio quality and fidelity while ensuring fast inference. IMPACT utilizes iterative mask-based parallel decoding in a continuous latent space powered by diffusion modeling. This approach eliminates the fidelity constraints of discrete tokens while maintaining competitive inference speed. Results on AudioCaps demonstrate that IMPACT achieves state-of-the-art performance on key metrics including Fréchet Distance (FD) and Fréchet Audio Distance (FAD) while significantly reducing latency compared to prior models. The project website is available at https://audio-impact.github.io/. Kuan-Po Huang, Shu-Wen Yang, Huy Phan, Bo-Ru Lu, Byeonggeun Kim, Sashank Macha, Qingming Tang, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao |
ICML | 7 |
| 2025 | Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionabstractAutoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous nature. We research audio generation with a causal language model (LM) without discrete tokens. We leverage token-wise diffusion to model the continuous distribution of the next continuous-valued token. Our approach delivers significant improvements over previous discrete solution, AudioGen, achieving 20% and 40% relative gains on AudioCaps in Frechet Audio Distance (FAD) and Kullback-Leibler (KL) divergence, respectively. Additionally, we propose a novel masked next-token prediction task that incorporates masked prediction into the causal LM framework. On AudioCaps, the innovation yields 41% and 33% relative FAD improvements over AudioGen Base (285M) and AudioGen Large (1B) models, respectively, and is on par with the state-of-the-art (SOTA) diffusion models. Furthermore, we achieve these results with significantly fewer parameters—193M for our Base and 462M for our Large models. Shu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang, Huy Phan, Bo-Ru Lu, Harshavardhan Sundar, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao |
ICML | 4 |
| 2024 | Learning for Transductive Threshold Calibration in Open-World RecognitionabstractIn deep metric learning for visual recognition, the calibration of distance thresholds is crucial for achieving desired model performance in the true positive rates (TPR) or true negative rates (TNR). However, calibrating this threshold presents challenges in open-world scenarios, where the test classes can be entirely disjoint from those encountered during training. We define the problem of finding distance thresholds for a trained embedding model to achieve target performance metrics over unseen open-world test classes as open-world threshold calibration. Existing posthoc threshold calibration methods, reliant on inductive inference and requiring a calibration dataset with a similar distance distribution as the test data, often prove ineffective in open-world scenarios. To address this, we introduce OpenGCN, a Graph Neural Network-based transductive threshold calibration method with enhanced adaptability and robustness. OpenGCN learns to predict pairwise connectivity for the unlabeled test instances embedded in a graph to determine its TPR and TNR at various distance thresholds, allowing for transductive inference of the distance thresholds which also incorporates test-time information. Extensive experiments across open-world visual recognition benchmarks validate OpenGCN's superiority over existing posthoc calibration methods for open-world threshold calibration. Dongsheng An, Tianjun Xiao, Tong He 0002, Qingming Tang, Ying Nian Wu, Joseph Tighe, Yifan Xing |
CVPR | 5 |
| 2024 | Cross-Triggering Issue in Audio Event Detection and MitigationabstractCross-triggering is a critical problem for applications of audio event detection (AED), particularly in low-resource settings. However, not much attention (if not none) has been paid to this problem in the AED research community. In this work, we tackle this problem via a regularization approach. We propose a regularizer, namely mutual exclusivity regularizer, that is able to enforce pairwise exclusivity between two event classes when they do not co-occur. When the regularizer is added to the loss function for network training, in effect, the increase in the score of one event class will result in the decrease of the other and vice versa. To quantify the effectiveness of the proposed regularizer, we developed an AED system based on convolutional neural network (CNN) for the detection of hand clap and door knock, two transient audio events that share similar spectro-temporal profiles, and conducted experiments on a large-scale real-world dataset (around 274.2 hours). The experimental results show that the proposed approach is able to largely mitigate the cross-triggering issue in various experimental settings. Furthermore, the reduction in cross-triggering, as a result, leads to improvement in the detection performance. Huy Phan, Byeonggeun Kim, Andrew Bydlon, Qingming Tang, Chieh-Chi Kao |
ICASSP | 5 |
| 2024 | On-Device Constrained Self-Supervised Learning for Keyword Spotting via Quantization Aware Pre-Training and Fine-TuningabstractLarge self-supervised models have excelled in various speech processing tasks, but their deployment on resource-limited devices is often impractical due to their substantial memory footprint. Previous studies have demonstrated the effectiveness of self-supervised pre-training for keyword spotting, even with constrained model capacity. In our pursuit of maintaining high performance while minimizing the model’s resource demands, we investigate the implementation of Quantization Aware Training for both self-supervised pre-training and fine-tuning, specifically tailored to fit within the constraints of on-device model budget. Our experiments emphasize the critical role of selecting and synchronizing QAT methods throughout both stages of model training and tuning. We evaluate our methodology on a 16.6k-hour in-house keyword spotting dataset, and show that there is no decline in performance, even when the bit size of model weights and activations is cut by a factor of four. Gene-Ping Yang, Sashank Macha, Qingming Tang, Yuzong Liu |
ICASSP | 4 |
| 2024 | Threshold-Consistent Margin Loss for Open-World Deep Metric LearningabstractExisting losses used in deep metric learning (DML) for image retrieval often lead to highly non-uniform intra-class and inter-class representation structures across test classes and data distributions. When combined with the common practice of using a fixed threshold to declare a match, this gives rise to significant performance variations in terms of false accept rate (FAR) and false reject rate (FRR) across test classes and data distributions. We define this issue in DML as threshold inconsistency. In real-world applications, such inconsistency often complicates the threshold selection process when deploying large-scale image retrieval systems. To measure this inconsistency, we propose a novel variance-based metric called Operating-Point-Inconsistency-Score (OPIS) that quantifies the variance in the operating characteristics across classes. Using the OPIS metric, we find that achieving high accuracy levels in a DML model does not automatically guarantee threshold consistency. In fact, our investigation reveals a Pareto frontier in the high-accuracy regime, where existing methods to improve accuracy often lead to degradation in threshold consistency. To address this trade-off, we introduce the Threshold-Consistent Margin (TCM) loss, a simple yet effective regularization technique that promotes uniformity in representation structures across classes by selectively penalizing hard sample pairs. Large-scale experiments demonstrate TCM's effectiveness in enhancing threshold consistency while preserving accuracy, simplifying the threshold selection process in practical DML settings. Linghan Xu, Qingming Tang, Ying Nian Wu, Joseph Tighe, Yifan Xing |
ICLR | 4 |
| 2023 | FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event ClassificationabstractPerformance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers and client devices in a privacy preserving manner. In this work we propose a novel Federated Relaxed Pareto Optimization (FedRPO) method for semi-supervised FL with heterogeneous client data. In contrast to federated averaging class of FL algorithms (fedAvg) that perform unconstrained weighted aggregation across all data sources, FedRPO enables special treatment of data with high quality annotations vs. data with pseudo-labels of unknown, varying qualities. In particular, FedRPO computes the updates to the global model solving a constrained linear program, with explicit Pareto constraints to prevent performance degradation on annotated data, and controlled relaxation of the Pareto constraints on pseudo-labeled data to prevent learning of patterns in conflict with the annotated data. We show FedRPO significantly outperforms FedAvg on Amazon internal de-identified dataset on AEC tasks. On supervised learning, FedRPO improved precision by 32.5% over FedAvg when maintaining recall at 90%. Combined with FixMatch [1] for semi-supervised learning, FedRPO outperformed FedAvg on precision by 50.5% at 90% recall. Meng Feng, Chieh-Chi Kao, Qingming Tang, Amit Solomon, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 3 |
| 2023 | Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device ConstraintsabstractAcoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that are tuned to achieve optimal accuracy/computation trade-off for each given computation resource constraint. In this paper, we introduce a Once-For-All (OFA) Neural Architecture Search (NAS) framework for AEC. Specifically, we first train a weight-sharing supernet that supports different model architectures, followed by automatically searching for a model given specific computational resource constraints. Our experimental results showed that by just training once, the resulting model from NAS significantly outperforms both models trained individually from scratch and knowledge distillation (25.4% and 7.3% relative improvement). We also found that the benefit of weight-sharing supernet training of ultra-small models comes not only from searching but from optimization. Guan-Ting Lin, Qingming Tang, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 2 |
| 2023 | On-Device Constrained Self-Supervised Speech Representation Learning for Keyword Spotting via Knowledge Distillation
Gene-Ping Yang, Qingming Tang, Dongsu Du, Yuzong Liu |
INTERSPEECH | 3 |
| 2022 | Federated Self-Supervised Learning for Acoustic Event ClassificationabstractStandard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate the feasibility of applying FL to improve AEC performance while no customer data can be directly uploaded to the server. We assume no pseudo labels can be inferred from on-device user inputs, aligning with the typical use cases of AEC. We adapt self-supervised learning to the FL framework for on-device continual learning of representations, and it results in improved performance of the downstream AEC classifiers with- out labeled/pseudo-labeled data available. Compared to the baseline w/o FL, the proposed method improves precision up to 20.3% relatively while maintaining the recall. Our work differs from prior work in FL in that our approach does not require user-generated learning targets, and the data we use is collected from our Beta program and is de-identified, to maximally simulate the production settings. Meng Feng, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 3 |
| 2022 | Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event ClassificationabstractAcoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wikipedia-based text embeddings as auxiliary information to improve AEC. We describe how to extract label embeddings from multiple Wikipedia texts, and formulate the multi-view aligned AEC problem based on VGGish model. We show that our "wikiTAG" embeddings encode rich semantic information and are more informative than label embeddings for AEC tasks. Compared to a supervised baseline on AudioSet, the multi-view model with "wikiTAG" embeddings achieves 7.3% and 1.3% relative improvement in mean average precision (mAP) using 10% and full AudioSet for training, respectively. To the author’s knowledge, this is the first work in the AEC domain on building large-scale label representations by leveraging Wikipedia data in a systematic fashion. Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018 |
ICASSP | 2 |
| 2022 | Improved Representation Learning For Acoustic Event Classification Using Tree-Structured OntologyabstractAcoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is not available in audios and plain tags. We show that by organizing audio representations with a human-curated tree ontology, we can improve the quality of the learned audio representations for downstream AEC tasks. We use consistency training to use large amounts of unlabeled data for structured representation manifold learning. Experimental results indicate that our framework learns high quality representations which enable us to achieve comparable performance in discriminative tasks as fully supervised baselines. Moreover, our framework can better handle audios with unseen tags by confidently assigning a super-category (internal node like "animal" in Fig. 1) tag to the audio. Arman Zharmagambetov, Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Viktor Rozgic, Jasha Droppo, Chao Wang 0018 |
ICASSP | 2 |
| 2021 | Multi-Task Self-Supervised Pre-Training for Music ClassificationabstractDeep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks. Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Brian McFee, Juan Pablo Bello, Chao Wang 0018 |
ICASSP | 3 |
| 2020 | Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked ReconstructionabstractWe propose an approach for pre-training speech representations via a masked reconstruction loss. Our pre-trained encoder networks are bidirectional and can therefore be used directly in typical bidirectional speech recognition models. The pre-trained networks can then be fine-tuned on a smaller amount of supervised data for speech recognition. Experiments with this approach on the LibriSpeech and Wall Street Journal corpora show promising results. We find that the main factors that lead to speech recognition improvements are: masking segments of sufficient width in both time and frequency, pre-training on a much larger amount of unlabeled data than the labeled data, and domain adaptation when the unlabeled and labeled data come from different domains. The gain from pre-training is additive to that of supervised data augmentation. Qingming Tang, Karen Livescu |
ICASSP | 2 |
| 2019 | Controllable Paraphrase Generation with a Syntactic ExemplarabstractPrior work on controllable text generation usually assumes that the controlled attribute can take on one of a small set of values known a priori.In this work, we propose a novel task, where the syntax of a generated sentence is controlled rather by a sentential exemplar.To evaluate quantitatively with standard metrics, we create a novel dataset with human annotations.We also develop a variational model with a neural module specifically designed for capturing syntactic knowledge and several multitask training objectives to promote disentangled representation learning.Empirically, the proposed model is observed to achieve improvements over baselines and learn to capture desirable characteristics. Mingda Chen, Qingming Tang, Sam Wiseman, Kevin Gimpel |
ACL (1) | 2 |
| 2019 | Hierarchical Residual-pyramidal Model for Large Context Based Media Presence DetectionabstractWe study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non-human sounds), and the recorded sound can be a mixture of media and non-media sound.Different from speech recognition, where the recognizer needs to detect local phonetic variation, the key features used to distinguish media and non-media sounds are non-local features. Motivated by this, we propose a hierarchical model to learn representation of each pre-chunked segment within a long recorded stream jointly, and encourage every local representation to be not sensitive to variations within each segment. We also further explore the effects of techniques including stream based normalization and iteratively imputing missing labels of training dataset. Experimental results indicate that our proposed contextual based methods are effective for media presence detection. Qingming Tang, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 1 |
| 2018 | Dependency-Aware Attention Control for Unconstrained Face Recognition with Image Sets
Xiaofeng Liu 0001, B. V. K. Vijaya Kumar, Chao Yang 0011, Qingming Tang, Jane You |
ECCV (11) | 4 |
| 2018 | Variational Sequential Labelers for Semi-Supervised LearningabstractWe introduce a family of multitask variational methods for semi-supervised sequence labeling.Our model family consists of a latentvariable generative model and a discriminative labeler.The generative models use latent variables to define the conditional probability of a word given its context, drawing inspiration from word prediction objectives commonly used in learning word embeddings.The labeler helps inject discriminative information into the latent space.We explore several latent variable configurations, including ones with hierarchical structure, which enables the model to account for both label-specific and word-specific information.Our models consistently outperform standard sequential baselines on 8 sequence labeling datasets, and improve further with unlabeled data. Mingda Chen, Qingming Tang, Karen Livescu, Kevin Gimpel |
EMNLP | 2 |
| 2018 | Acoustic Feature Learning Using Cross-Domain Articulatory MeasurementsabstractPrevious work has shown that it is possible to improve speech recognition by learning acoustic features from paired acoustic-articulatory data, for example by using canonical correlation analysis (CCA) or its deep extensions. One limitation of this prior work is that the learned feature models are difficult to port to new datasets or domains, and articulatory data is not available for most speech corpora. In this work we study the problem of acoustic feature learning in the setting where we have access to an external, domain-mismatched dataset of paired speech and articulatory measurements, either with or without labels. We develop methods for acoustic feature learning in these settings, based on deep variational CCA and extensions that use both source and target domain data and labels. Using this approach, we improve phonetic recognition accuracies on both TIMIT and Wall Street Journal and analyze a number of design choices. Qingming Tang, Karen Livescu |
ICASSP | 1 |
| 2017 | Acoustic Feature Learning via Deep Variational Canonical Correlation AnalysisabstractWe study the problem of acoustic feature learning in the setting where we have access to another (non-acoustic) modality for feature learning but not at test time.We use deep variational canonical correlation analysis (VCCA), a recently proposed deep generative method for multi-view representation learning.We also extend VCCA with improved latent variable priors and with adversarial learning.Compared to other techniques for multi-view feature learning, VCCA's advantages include an intuitive latent variable interpretation and a variational lower bound objective that can be trained end-to-end efficiently.We compare VCCA and its extensions with previous feature learning methods on the University of Wisconsin X-ray Microbeam Database, and show that VCCA-based feature learning improves over previous methods for speaker-independent phonetic recognition. Qingming Tang, Karen Livescu |
INTERSPEECH | 1 |
| 2017 | BASIC: BCR assembly from single cellsabstractMotivation: The B-cell receptor enables individual B cells to identify diverse antigens, including bacterial and viral proteins. While advances in RNA-sequencing (RNA-seq) have enabled high throughput profiling of transcript expression in single cells, the unique task of assembling the full-length heavy and light chain sequences from single cell RNA-seq (scRNA-seq) in B cells has been largely unstudied. Results: We developed a new software tool, BASIC, which allows investigators to use scRNA-seq for assembling BCR sequences at single-cell resolution. To demonstrate the utility of our software, we subjected nearly 200 single human B cells to scRNA-seq, assembled the full-length heavy and the light chains, and experimentally confirmed these results by using single-cell primer-based nested PCRs and Sanger sequencing. Availability and Implementation: http://ttic.uchicago.edu/∼aakhan/BASIC Contact: [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Stefan Canzar, Karlynn E. Neu, Qingming Tang, Patrick C. Wilson, Aly Azeem Khan |
Bioinform. | 3 |
| 2015 | Learning Scale-Free Networks by Dynamic Node Specific Degree PriorabstractLearning network structure underlying data is an important problem in machine learning. This paper presents a novel degree prior to study the inference of scale-free networks, which are widely used to model social and biological networks. In particular, this paper formulates scale-free network inference using Gaussian Graphical model (GGM) regularized by a node degree prior. Our degree prior not only promotes a desirable global degree distribution, but also exploits the estimated degree of an individual node and the relative strength of all the edges of a single node. To fulfill this, this paper proposes a ranking-based method to dynamically estimate the degree of a node, which makes the resultant optimization problem challenging to solve. To deal with this, this paper presents a novel ADMM (alternating direction method of multipliers) procedure. Our experimental results on both synthetic and real data show that our prior not only yields a scale-free network, but also produces many more correctly predicted edges than existing scale-free inducing prior, hub-inducing prior and the l_1 norm. Qingming Tang, Jinbo Xu |
ICML | 1 |
| 2015 | Exact Hybrid Covariance Thresholding for Joint Graphical Lasso
Qingming Tang, Jian Peng 0001, Jinbo Xu |
ECML/PKDD (2) | 1 |