EDBT 2026 Demo / reviewers in the wild / expert
Chieh-Chi Kao
dblp:38/8116
· DBLP profile ↗
32ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-1161-3959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effective Techniques for Scaling Audio Encoder PretrainingabstractThis work presents advancements in audio pretraining objectives designed to generate semantically rich embeddings, capable of addressing a wide range of audio-related tasks. Despite significant progress in the field, current methods often emphasize full fine-tuning in downstream applications, which can obscure the true potential of pretrained audio encoders. In this study, we present an audio encoder that achieves stateof-the-art (SOTA) performances in both fine-tuning and linear probing, utilizing a carefully curated set of pragmatic techniques. Building on previous research, we incorporate masked prediction and introduce SpecAug within a curriculum masking strategy at the patch level, which progressively increases training difficulty, along with a mask-aware position bias. To comprehensively assess the encoder’s capabilities, we examine the impact of scaling both the dataset size and model capacity, conducting linear probing evaluations while keeping the encoder frozen as well as full fine-tuning. Our model demonstrates superior performance compared to recent SOTA methods across various downstream tasks. Additionally, we explore the potential of tokenizing the resulting audio embeddings for use as discrete inputs, enhancing our understanding of the model’s capabilities. Byeonggeun Kim, Andrew Bydlon, Qingming Tang, Huy Phan, Chieh-Chi Kao |
ICASSP | 5 |
| 2025 | IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion ModelingabstractText-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio generation. Despite achieving high audio fidelity, they incur significant inference latency due to the slow diffusion sampling process. MAGNET, a mask-based model operating on discrete tokens, addresses slow inference through iterative mask-based parallel decoding. However, its audio quality still lags behind that of diffusion-based models. In this work, we introduce IMPACT, a text-to-audio generation framework that achieves high performance in audio quality and fidelity while ensuring fast inference. IMPACT utilizes iterative mask-based parallel decoding in a continuous latent space powered by diffusion modeling. This approach eliminates the fidelity constraints of discrete tokens while maintaining competitive inference speed. Results on AudioCaps demonstrate that IMPACT achieves state-of-the-art performance on key metrics including Fréchet Distance (FD) and Fréchet Audio Distance (FAD) while significantly reducing latency compared to prior models. The project website is available at https://audio-impact.github.io/. Kuan-Po Huang, Shu-Wen Yang, Huy Phan, Bo-Ru Lu, Byeonggeun Kim, Sashank Macha, Qingming Tang, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao |
ICML | 10 |
| 2025 | Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionabstractAutoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous nature. We research audio generation with a causal language model (LM) without discrete tokens. We leverage token-wise diffusion to model the continuous distribution of the next continuous-valued token. Our approach delivers significant improvements over previous discrete solution, AudioGen, achieving 20% and 40% relative gains on AudioCaps in Frechet Audio Distance (FAD) and Kullback-Leibler (KL) divergence, respectively. Additionally, we propose a novel masked next-token prediction task that incorporates masked prediction into the causal LM framework. On AudioCaps, the innovation yields 41% and 33% relative FAD improvements over AudioGen Base (285M) and AudioGen Large (1B) models, respectively, and is on par with the state-of-the-art (SOTA) diffusion models. Furthermore, we achieve these results with significantly fewer parameters—193M for our Base and 462M for our Large models. Shu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang, Huy Phan, Bo-Ru Lu, Harshavardhan Sundar, Shalini Ghosh, Hung-yi Lee, Chieh-Chi Kao |
ICML | 10 |
| 2024 | Cross-Triggering Issue in Audio Event Detection and MitigationabstractCross-triggering is a critical problem for applications of audio event detection (AED), particularly in low-resource settings. However, not much attention (if not none) has been paid to this problem in the AED research community. In this work, we tackle this problem via a regularization approach. We propose a regularizer, namely mutual exclusivity regularizer, that is able to enforce pairwise exclusivity between two event classes when they do not co-occur. When the regularizer is added to the loss function for network training, in effect, the increase in the score of one event class will result in the decrease of the other and vice versa. To quantify the effectiveness of the proposed regularizer, we developed an AED system based on convolutional neural network (CNN) for the detection of hand clap and door knock, two transient audio events that share similar spectro-temporal profiles, and conducted experiments on a large-scale real-world dataset (around 274.2 hours). The experimental results show that the proposed approach is able to largely mitigate the cross-triggering issue in various experimental settings. Furthermore, the reduction in cross-triggering, as a result, leads to improvement in the detection performance. Huy Phan, Byeonggeun Kim, Andrew Bydlon, Qingming Tang, Chieh-Chi Kao |
ICASSP | 6 |
| 2023 | FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event ClassificationabstractPerformance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers and client devices in a privacy preserving manner. In this work we propose a novel Federated Relaxed Pareto Optimization (FedRPO) method for semi-supervised FL with heterogeneous client data. In contrast to federated averaging class of FL algorithms (fedAvg) that perform unconstrained weighted aggregation across all data sources, FedRPO enables special treatment of data with high quality annotations vs. data with pseudo-labels of unknown, varying qualities. In particular, FedRPO computes the updates to the global model solving a constrained linear program, with explicit Pareto constraints to prevent performance degradation on annotated data, and controlled relaxation of the Pareto constraints on pseudo-labeled data to prevent learning of patterns in conflict with the annotated data. We show FedRPO significantly outperforms FedAvg on Amazon internal de-identified dataset on AEC tasks. On supervised learning, FedRPO improved precision by 32.5% over FedAvg when maintaining recall at 90%. Combined with FixMatch [1] for semi-supervised learning, FedRPO outperformed FedAvg on precision by 50.5% at 90% recall. Meng Feng, Chieh-Chi Kao, Qingming Tang, Amit Solomon, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 2 |
| 2023 | Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device ConstraintsabstractAcoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that are tuned to achieve optimal accuracy/computation trade-off for each given computation resource constraint. In this paper, we introduce a Once-For-All (OFA) Neural Architecture Search (NAS) framework for AEC. Specifically, we first train a weight-sharing supernet that supports different model architectures, followed by automatically searching for a model given specific computational resource constraints. Our experimental results showed that by just training once, the resulting model from NAS significantly outperforms both models trained individually from scratch and knowledge distillation (25.4% and 7.3% relative improvement). We also found that the benefit of weight-sharing supernet training of ultra-small models comes not only from searching but from optimization. Guan-Ting Lin, Qingming Tang, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 3 |
| 2022 | Federated Self-Supervised Learning for Acoustic Event ClassificationabstractStandard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate the feasibility of applying FL to improve AEC performance while no customer data can be directly uploaded to the server. We assume no pseudo labels can be inferred from on-device user inputs, aligning with the typical use cases of AEC. We adapt self-supervised learning to the FL framework for on-device continual learning of representations, and it results in improved performance of the downstream AEC classifiers with- out labeled/pseudo-labeled data available. Compared to the baseline w/o FL, the proposed method improves precision up to 20.3% relatively while maintaining the recall. Our work differs from prior work in FL in that our approach does not require user-generated learning targets, and the data we use is collected from our Beta program and is de-identified, to maximally simulate the production settings. Meng Feng, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 2 |
| 2022 | Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event ClassificationabstractAcoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wikipedia-based text embeddings as auxiliary information to improve AEC. We describe how to extract label embeddings from multiple Wikipedia texts, and formulate the multi-view aligned AEC problem based on VGGish model. We show that our "wikiTAG" embeddings encode rich semantic information and are more informative than label embeddings for AEC tasks. Compared to a supervised baseline on AudioSet, the multi-view model with "wikiTAG" embeddings achieves 7.3% and 1.3% relative improvement in mean average precision (mAP) using 10% and full AudioSet for training, respectively. To the author’s knowledge, this is the first work in the AEC domain on building large-scale label representations by leveraging Wikipedia data in a systematic fashion. Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018 |
ICASSP | 3 |
| 2022 | Improved Representation Learning For Acoustic Event Classification Using Tree-Structured OntologyabstractAcoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is not available in audios and plain tags. We show that by organizing audio representations with a human-curated tree ontology, we can improve the quality of the learned audio representations for downstream AEC tasks. We use consistency training to use large amounts of unlabeled data for structured representation manifold learning. Experimental results indicate that our framework learns high quality representations which enable us to achieve comparable performance in discriminative tasks as fully supervised baselines. Moreover, our framework can better handle audios with unseen tags by confidently assigning a super-category (internal node like "animal" in Fig. 1) tag to the audio. Arman Zharmagambetov, Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Viktor Rozgic, Jasha Droppo, Chao Wang 0018 |
ICASSP | 3 |
| 2021 | Multi-Task Self-Supervised Pre-Training for Music ClassificationabstractDeep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks. Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Brian McFee, Juan Pablo Bello, Chao Wang 0018 |
ICASSP | 2 |
| 2020 | A Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event ClassificationabstractAcoustic event classification (AEC) and acoustic event detection (AED) refer to the task of detecting whether specific target events occur in audios. As long short-term memory (LSTM) leads to state-of-the-art results in various speech related tasks, it is employed as a popular solution for AEC as well. This paper focuses on investigating the dynamics of LSTM model on AEC tasks. It includes a detailed analysis on LSTM memory retaining, and a benchmarking of nine different pooling methods on LSTM models using 1.7M generated mixture clips of multiple events with different signal-to-noise ratios. This paper focuses on understanding: 1) utterance-level classification accuracy; 2) sensitivity to event position within an utterance. The analysis is done on the dataset for the detection of rare sound events from DCASE 2017 Challenge. We find max pooling on the prediction level to perform the best among the nine pooling approaches in terms of classification accuracy and insensitivity to event position within an utterance. To authors’ best knowledge, this is the first kind of such work focused on LSTM dynamics for AEC tasks. Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018 |
ICASSP | 1 |
| 2020 | Few-Shot Acoustic Event Detection Via Meta LearningabstractWe study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition has been under-studied. We formulate few-shot AED problem and explore different ways of utilizing traditional supervised methods for this setting as well as a variety of meta-learning approaches, which are conventionally used to solve few-shot classification problem. Compared to supervised baselines, meta-learning models achieve superior performance, thus showing its effectiveness on generalization to new audio events. Our analysis including impact of initialization and domain discrepancy further validate the advantage of meta-learning approaches in few-shot AED. Bowen Shi 0002, Ming Sun 0007, Krishna C. Puvvada, Chieh-Chi Kao, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 4 |
| 2020 | Intra-Utterance Similarity Preserving Knowledge Distillation for Audio TaggingabstractKnowledge Distillation (KD) is a popular area of research for reducing the size of large models while still maintaining good performance.The outputs of larger teacher models are used to guide the training of smaller student models.Given the repetitive nature of acoustic events, we propose to leverage this information to regulate the KD training for Audio Tagging.This novel KD method, Intra-Utterance Similarity Preserving KD (IUSP), shows promising results for the audio tagging task.It is motivated by the previously published KD method: Similarity Preserving KD (SP).However, instead of preserving the pairwise similarities between inputs within a mini-batch, our method preserves the pairwise similarities between the frames of a single input utterance.Our proposed KD method, IUSP, shows consistent improvements over SP across student models of different sizes on the DCASE 2019 Task 5 dataset for audio tagging.There is a 27.1% to 122.4% percent increase in improvement of micro AUPRC over the baseline relative to SPs improvement of over the baseline. Chun-Chieh Chang, Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018 |
INTERSPEECH | 2 |
| 2020 | On Front-End Gain Invariant Modeling for Wake Word SpottingabstractWake word (WW) spotting is challenging in far-field due to the complexities and variations in acoustic conditions and the environmental interference in signal transmission. A suite of carefully designed and optimized audio front-end (AFE) algorithms help mitigate these challenges and provide better quality audio signals to the downstream modules such as WW spotter. Since the WW model is trained with the AFE-processed audio data, its performance is sensitive to AFE variations, such as gain changes. In addition, when deploying to new devices, the WW performance is not guaranteed because the AFE is unknown to the WW model. To address these issues, we propose a novel approach to use a new feature called $\Delta$LFBE to decouple the AFE gain variations from the WW model. We modified the neural network architectures to accommodate the delta computation, with the feature extraction module unchanged. We evaluate our WW models using data collected from real household settings and showed the models with the $\Delta$LFBE is robust to AFE gain changes. Specifically, when AFE gain changes up to $\pm$12dB, the baseline CNN model lost up to relative 19.0% in false alarm rate or 34.3% in false reject rate, while the model with $\Delta$LFBE demonstrates no performance loss. Noah D. Stein, Chieh-Chi Kao, Yunliang Cai, Ming Sun 0007, Shiv Vitaladevuni |
INTERSPEECH | 3 |
| 2020 | A Joint Framework for Audio Tagging and Weakly Supervised Acoustic Event Detection Using DenseNet with Global Average PoolingabstractThis paper proposes a network architecture mainly designed for audio tagging, which can also be used for weakly supervised acoustic event detection (AED).The proposed network consists of a modified DenseNet as the feature extractor, and a global average pooling (GAP) layer to predict frame-level labels at inference time.This architecture is inspired by the work proposed by Zhou et al., a well-known framework using GAP to localize visual objects given image-level labels.While most of the previous works on weakly supervised AED used recurrent layers with attention-based mechanism to localize acoustic events, the proposed network directly localizes events using the feature map extracted by DenseNet without any recurrent layers.In the audio tagging task of DCASE 2017, our method significantly outperforms the state-of-the-art method in F1 score by 5.3% on the dev set, and 6.0% on the eval set in terms of absolute values.For weakly supervised AED task in DCASE 2018, our model outperforms the state-of-the-art method in event-based F1 by 8.1% on the dev set, and 0.5% on the eval set in terms of absolute values, by using data augmentation and tri-training to leverage unlabeled data. Chieh-Chi Kao, Bowen Shi 0002, Ming Sun 0007, Chao Wang 0018 |
INTERSPEECH | 1 |
| 2020 | Patch-Based Image Hallucination for Super Resolution With Detail Reconstruction From Similar Sample ImagesabstractImage hallucination and super-resolution have been studied for decades, and many approaches have been proposed to upsample low-resolution images using information from the images themselves, multiple example images, or large image databases. However, most of this work has focused exclusively on small magnification levels because the algorithms simply sharpen the blurry edges in the upsampled images - no actual new detail is typically reconstructed in the final result. In this paper, we present a patch-based algorithm for image hallucination which, for the first time, properly synthesizes novel high frequency detail. To do this, we pose the synthesis problem as a patch-based optimization which inserts coherent, high-frequency detail from contextually-similar images of the same physical scene/subject provided from either a personal image collection or a large online database. The resulting image is visually plausible and contains coherent high frequency information. We demonstrate the robustness of our algorithm by testing it on a large number of images and show that its performance is considerably superior to all state-of-the-art approaches, a result that is verified to be statistically significant through a randomized user study. Chieh-Chi Kao, Yu-Xiang Wang 0003, Jonathan Waltman, Pradeep Sen |
IEEE Trans. Multim. | 1 |
| 2019 | Semi-supervised Acoustic Event Detection Based on Tri-trainingabstractThis paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and relevant acoustic event audios can be limited, especially for rare events. In this paper we leverage an Internet-scale un-labeled dataset with potential domain shift to improve the detection of acoustic events. Based on the classic tri-training approach, our proposed method shows accuracy improvement over both the supervised training baseline, and semi-supervised self-training set-up, in all pre-defined acoustic event detection tasks. As our approach relies on ensemble models, we further show the improvements can be distilled to a single model via knowledge distillation, with the resulting single student model maintaining high accuracy of teacher ensemble models. Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 3 |
| 2019 | Hierarchical Residual-pyramidal Model for Large Context Based Media Presence DetectionabstractWe study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non-human sounds), and the recorded sound can be a mixture of media and non-media sound.Different from speech recognition, where the recognizer needs to detect local phonetic variation, the key features used to distinguish media and non-media sounds are non-local features. Motivated by this, we propose a hierarchical model to learn representation of each pre-chunked segment within a long recorded stream jointly, and encourage every local representation to be not sensitive to variations within each segment. We also further explore the effects of techniques including stream based normalization and iteratively imputing missing labels of training dataset. Experimental results indicate that our proposed contextual based methods are effective for media presence detection. Qingming Tang, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018 |
ICASSP | 3 |
| 2019 | Sub-Band Convolutional Neural Networks for Small-Footprint Spoken Term ClassificationabstractThis paper proposes a Sub-band Convolutional Neural Network for spoken term classification.Convolutional neural networks (CNNs) have proven to be very effective in acoustic applications such as spoken term classification, keyword spotting, speaker identification, acoustic event detection, etc.Unlike applications in computer vision, the spatial invariance property of 2D convolutional kernels does not fit acoustic applications well since the meaning of a specific 2D kernel varies a lot along the feature axis in an input feature map.We propose a sub-band CNN architecture to apply different convolutional kernels on each feature sub-band, which makes the overall computation more efficient.Experimental results show that the computational efficiency brought by sub-band CNN is more beneficial for smallfootprint models.Compared to a baseline full band CNN for spoken term classification on a publicly available Speech Commands dataset, the proposed sub-band CNN architecture reduces the computation by 39.7% on commands classification, and 49.3% on digits classification with accuracy maintained. Chieh-Chi Kao, Ming Sun 0007, Shiv Vitaladevuni, Chao Wang 0018 |
INTERSPEECH | 1 |
| 2019 | Compression of Acoustic Event Detection Models with Quantized DistillationabstractAcoustic Event Detection (AED), aiming at detecting categories of events based on audio signals, has found application in many intelligent systems. Recently deep neural network significantly advances this field and reduces detection errors to a large scale. However how to efficiently execute deep models in AED has received much less attention. Meanwhile state-of-the-art AED models are based on large deep models, which are computational demanding and challenging to deploy on devices with constrained computational resources. In this paper, we present a simple yet effective compression approach which jointly leverages knowledge distillation and quantization to compress larger network (teacher model) into compact network (student model). Experimental results show proposed technique not only lowers error rate of original compact network by 15% through distillation but also further reduces its model size to a large extent (2% of teacher, 12% of full-precision student) through quantization. Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
INTERSPEECH | 3 |
| 2018 | Localization-Aware Active Learning for Object Detection
Chieh-Chi Kao, Teng-Yok Lee, Pradeep Sen, Ming-Yu Liu 0001 |
ACCV (6) | 1 |
| 2018 | R-CRNN: Region-based Convolutional Recurrent Neural Network for Audio Event DetectionabstractThis paper proposes a Region-based Convolutional Recurrent Neural Network (R-CRNN) for audio event detection (AED).The proposed network is inspired by Faster-RCNN [1], a wellknown region-based convolutional network framework for visual object detection.Different from the original Faster-RCNN, a recurrent layer is added on top of the convolutional network to capture the long-term temporal context from the extracted highlevel features.While most of the previous works on AED generate predictions at frame level first, and then use post-processing to predict the onset/offset timestamps of events from a probability sequence; the proposed method generates predictions at event level directly and can be trained end-to-end with a multitask loss, which optimizes the classification and localization of audio events simultaneously.The proposed method is tested on DCASE 2017 Challenge dataset [2].To the best of our knowledge, R-CRNN is the best performing single-model method among all methods without using ensembles both on development and evaluation sets.Compared to the other region-based network for AED (R-FCN [3]) with an event-based error rate (ER) of 0.18 on the development set, our method reduced the ER to half. Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018 |
INTERSPEECH | 1 |
| 2018 | A Simple Model for Detection of Rare Sound EventsabstractWe propose a simple recurrent model for detecting rare sound events, when the time boundaries of events are available for training. Our model optimizes the combination of an utterance-level loss, which classifies whether an event occurs in an utterance, and a frame-level loss, which classifies whether each frame corresponds to the event when it does occur. The two losses make use of a shared vectorial representation the event, and are connected by an attention mechanism. We demonstrate our model on Task 2 of the DCASE 2017 challenge, and achieve competitive performance. Chieh-Chi Kao, Chao Wang 0018 |
INTERSPEECH | 2 |
| 2014 | VLSI Architecture Design of Guided Filter for 30 Frames/s Full-HD VideoabstractFiltering is widely used in image and video processing for various applications. Recently, the guided filter has been proposed and became one of the popular filtering methods. In this paper, to achieve the computation demand of guided filtering in full-HD video, a double integral image architecture for guided filter ASIC design is proposed. In addition, a reformation of the guided filter formula is proposed, which can prevent the error resulted from truncation in the fractional part and modify the regularization parameter ε on user's demand. The hardware architecture of the guided image filter is then proposed and can be embedded in mobile devices to achieve real-time HD applications. To the best of our knowledge, this paper is also the first ASIC design for guided image filter. With a TSMC 90-nm cell library, the design can operate at 100 MHz and support for Full-HD (1920 × 1080) 30 frame/s with 92.9K gate counts and 3.2 KB on-chip memory. Moreover, for the hardware efficiency, our architecture is also the best compared to other previous works with bilateral filter. Chieh-Chi Kao, Jui-Hsin Lai, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Sampling Technique Analysis of Nyström Approximation in Pixel-Wise Affinity MatrixabstractSpectral graph methods are widely employed in image segmentation, and they exhibit excellent performance. However, for high-resolution images, it is impractical to directly calculate the eigenvectors of the affinity matrix owing to the high computational requirements. The Nystrom method provides an efficient way to approximate the large-scale affinity matrix by low-rank approximation. In the machine learning field, previous studies have mainly focused on less data points with high dimensional features. To the best of our knowledge, this is the first study to discuss the performance of sampling methods for Nystrom approximation, in which we focus on the pixel-wise affinity matrix for a single image. In this paper, we propose a mean-shift segmentation-based Nystrom sampling technique for image analysis. The experimental results show that for images with simple compositions and backgrounds, k-means sampling performs better, whereas for images with more complicated compositions and backgrounds, the proposed method can perform better. Chieh-Chi Kao, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 1 |
| 2012 | Tennis Real PlayabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose the process for database normalization and the behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose the framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Min-Chun Hu 0001, Shao-Yi Chien |
IEEE Trans. Multim. | 4 |
| 2012 | Preference-Aware View Recommendation System for Scenic Photos Based on Bag-of-Aesthetics-Preserving FeaturesabstractIn this paper, the framework for a real-time view recommendation system is proposed. The proposed system comprises two parts: offline aesthetic modeling stage and efficient online aesthetic view finding process. A preference-aware aesthetic model is proposed to suggest views according to varied user-favorite photographic styles, where a bottom-up approach is developed to construct an aesthetic feature library with bag-of-aesthetics-preserving features instead of top-down methods that implement the heuristic guidelines (rule-specific features) listed in photography literatures, which is employed in previous works. A collection of scenic photos is used as the test set; however, the proposed method can be employed to other types of photo collection according to different application scenarios. The proposed model can cover both implicit and explicit aesthetic features and can adapt to users' preferences with a learning process. In the second part, the learned model is employed in a view finder to help the user to locate the most aesthetic view while taking a photograph. The experimental results show that the proposed features in the library (92.06% in accuracy) outperform the state-of-the-art rule-specific features (83.63% in accuracy) significantly in the photo aesthetic quality classification task, and the rule-specific features are also proved to be encompassed by the proposed features. Meanwhile, it is observed from experiments that the features extracted for contrast information are more effective than those for absolute information, which is consistent with the properties of human visual systems. Furthermore, the user studies for the view recommendation task confirm that the suggested views are consistent with users' preferences (81.25% agreements). Hsiao-Hang Su, Tse-Wei Chen 0001, Chieh-Chi Kao, Winston H. Hsu, Shao-Yi Chien |
IEEE Trans. Multim. | 3 |
| 2011 | Automatic object segmentation with salient color modelabstractImage segmentation is a well-developing topic in the image processing, and a number of previous works have been proposed and achieved high performance. However, most previous works needed user-assistance to provide the prior information of the target object in the segmentation. In this paper we propose an unsupervised scheme, combining the salient object detection and segmentation method, to segment the target object without any prior information from users. The experimental results show that the proposed salient color model derived with salient features can provide a prior information with high confidence to generate precise segmentation automatically. The proposed color model of salient objects can not only be applied with Min-Cut algorithm, but also extended to more segmentation algorithms, like matting or non-parametric model. Chieh-Chi Kao, Jui-Hsin Lai, Shao-Yi Chien |
ICME | 1 |
| 2011 | Tennis real play: an interactive tennis game with models from real videosabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose a database normalization process and a behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose a framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Shao-Yi Chien |
ACM Multimedia | 4 |
| 2011 | Scenic photo quality assessment with bag of aesthetics-preserving featuresabstractIn this paper, an aesthetic modeling method for scenic photographs is proposed. A bottom-up approach is developed to construct an aesthetic library with bag-of-aesthetics preserving features instead of top-down methods that implement the heuristic guidelines (rule-specific features) listed in the photography literature, which is employed in previous works. The proposed method can cover both implicit and explicit aesthetic features with a learning process. The experimental results show that the proposed features in the library (92.06% in accuracy) outperform the state-of-the-art rule-specific features (83.63% in accuracy) significantly in the aesthetic quality assessment for scenic photos, and the rule-specific features are also proved to be encompassed by the proposed features. Meanwhile, it is observed from experiments that the features extracted for contrast information are more effective than those for absolute information, which is consistent with the properties of human visual systems. Hsiao-Hang Su, Tse-Wei Chen 0001, Chieh-Chi Kao, Winston H. Hsu, Shao-Yi Chien |
ACM Multimedia | 3 |
| 2011 | Tennis Video 2.0: A new presentation of sports videos with content separation and rendering
Jui-Hsin Lai, Chieh-Li Chen, Chieh-Chi Kao, Shao-Yi Chien |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Super-resolution sprite with foreground removalabstractSprite is an image constructed from video clips and is also a medium for multimedia applications. An automatic sprite generation with foreground removal and super-resolution is proposed in this paper. To remove the foreground objects, each pixel-value on the sprite is iteratively updated by the value with maximum appearance probability on temporal and spatial distribution. By storing the half-pixel, superresolution sprite has less blurring-defect from source video. In the result, the generated sprite preserves the complete scenes of background and has higher image quality, and it can used to increase the visual quality in current sprite applications and also employed to facilitate video segmentation. Jui-Hsin Lai, Chieh-Chi Kao, Shao-Yi Chien |
ICME | 2 |