Chunxi Liu

dblp:05/3283 · DBLP profile ↗
← Back
36ranked-venue papers
17as first author
10since 2021 · last 2023
0000-0001-5441-9374ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 16 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 2 since 2021
YearPublicationVenuePosition
2023 TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for Pytorch
abstract
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance.
Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao
ASRU14
2023 Learning ASR Pathways: A Sparse Multilingual ASR Model
abstract
Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-specific parameters. In this work, we present ASR pathways, a sparse multilingual ASR model that activates language-specific sub-networks ("pathways"), such that the parameters for each language are learned explicitly. With the overlapping sub-networks, the shared parameters can also enable knowledge transfer for lower-resource languages via joint multilingual training. We propose a novel algorithm to learn ASR pathways, and evaluate the proposed method on 4 languages with a streaming RNN-T model. Our proposed ASR pathways outperform both dense models and a language-agnostically pruned model, and provide better performance on low-resource languages compared to the monolingual sparse models.
Mu Yang, Andros Tjandra, Chunxi Liu, Ozlem Kalinli
ICASSP3
2023 Multi-Head State Space Model for Speech Recognition
Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer, Mark J. F. Gales
INTERSPEECH7
2022 Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions
abstract
The problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conversations” – a publicly released 846 hour corpus designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of metadata, including age, gender, and skin tone. The entire corpus has been manually transcribed, allowing for detailed ASR evaluations across these metadata. Multiple ASR models are evaluated, including models trained on LibriSpeech, 14,000 hour transcribed, and over 2 million hour untranscribed social media videos. Significant differences in word error rate across gender and skin tone are observed at times for all models. We are releasing human transcripts from the Casual Conversations dataset to encourage the community to develop a variety of techniques to reduce these statistical biases.
Chunxi Liu, Michael Picheny, Leda Sari, Pooja Chitkara, Alex Xiao, Xiaohui Zhang 0007, Mark Chou, Andres Alvarado, Caner Hazirbas, Yatharth Saraf
ICASSP1
2022 Streaming Transformer Transducer based Speech Recognition Using Non-Causal Convolution
abstract
This paper improves the streaming transformer transducer for speech recognition using non-causal convolution. Many works apply the causal convolution to improve streaming transformer ignoring the lookahead context. We propose to use non-causal convolution to process the center block and lookahead context separately. This method leverages the lookahead context in convolution and maintains similar training and decoding efficiency. Given the similar latency, using the non-causal convolution with lookahead context gives better accuracy than causal convolution, especially for open-domain dictation. Besides, this paper applies talking-head attention and a novel history context compression scheme to further improve the performance. The talking-head attention improves the multi-head self-attention by transferring information among different heads. The history context compression method introduces more extended history context compactly. On our in-house data, the proposed methods improve a small Emformer baseline with lookahead context by relative WERR 5.1%, 14.5%, 8.4% on open-domain dictation, assistant general scenarios, and assistant calling scenarios respectively.
Yangyang Shi, Chunyang Wu, Dilin Wang, Alex Xiao, Jay Mahadeokar, Xiaohui Zhang 0007, Chunxi Liu, Ke Li 0023, Yuan Shangguan, Varun Nagaraja, Ozlem Kalinli, Mike Seltzer
ICASSP7
2022 Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks
abstract
Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have comprehensively analyzed audio representation learning for non-speech audio tasks. In this paper, we propose a self-supervised audio representation learning method and apply it to a variety of downstream non-speech audio tasks. We combine the well-known wav2vec 2.0 framework, which has shown success in self-supervised learning for speech tasks, with parameter-efficient conformer architectures. Our self-supervised pre-training can reduce the need for labeled data by two-thirds. On the AudioSet benchmark, we achieve a mean average precision (mAP) score of 0.415, which is a new state-of-the-art on this dataset through audio-only self-supervised learning. Our fine-tuned conformers also surpass or match the performance of previous systems pre-trained in a supervised way on several downstream tasks. We further discuss the important design considerations for both pre-training and fine-tuning.
Sangeeta Srivastava, Andros Tjandra, Anurag Kumar 0003, Chunxi Liu, Kritika Singh, Yatharth Saraf
ICASSP5
2022 Learning a Dual-Mode Speech Recognition Model VIA Self-Pruning
abstract
There is growing interest in unifying the streaming and full-context automatic speech recognition (ASR) networks into a single end-to-end ASR model to simplify the model training and deployment for both use cases. While in real-world ASR applications, the streaming ASR models typically operate under more storage and computational constraints - e.g., on embedded devices - than any server-side full-context models. Motivated by the recent progress in Omni-sparsity supernet training, where multiple subnetworks are jointly optimized in one single model, this work aims to jointly learn a compact sparse on-device streaming ASR model, and a large dense server non-streaming model, in a single supernet. Next, we present that, performing supernet training on both wav2vec 2.0 self-supervised learning and supervised ASR fine-tuning can not only substantially improve the large non-streaming model as shown in prior works, and also be able to improve the compact sparse streaming model.
Chunxi Liu, Yuan Shangguan, Haichuan Yang, Yangyang Shi, Raghuraman Krishnamoorthi, Ozlem Kalinli
SLT1
2021 Improving RNN Transducer Based ASR with Auxiliary Tasks
abstract
End-to-end automatic speech recognition (ASR) models with a single neural network have recently demonstrated state-of-the-art results compared to conventional hybrid speech recognizers. Specifically, recurrent neural network transducer (RNN-T) has shown competitive ASR performance on various benchmarks. In this work, we examine ways in which RNN-T can achieve better ASR accuracy via performing auxiliary tasks. We propose (i) using the same auxiliary task as primary RNN-T ASR task, and (ii) performing context-dependent graphemic state prediction as in conventional hybrid modeling. In transcribing social media videos with varying training data size, we first evaluate the streaming ASR performance on three languages: Romanian, Turkish and German. We find that both proposed methods provide consistent improvements. Next, we observe that both auxiliary tasks demonstrate efficacy in learning deep transformer encoders for RNN-T criterion, thus achieving competitive results -2.0%/4.2% WER on LibriSpeech test-clean/other - as compared to prior top performing models.
Chunxi Liu, Frank Zhang 0001, Suyoun Kim, Yatharth Saraf, Geoffrey Zweig
SLT1
2021 Dual Application of Speech Enhancement for Automatic Speech Recognition
abstract
In this work, we exploit speech enhancement for improving a re-current neural network transducer (RNN-T) based ASR system. We employ a dense convolutional recurrent network (DCRN) for complex spectral mapping based speech enhancement, and find it helpful for ASR in two ways: a data augmentation technique, and a preprocessing frontend. In using it for ASR data augmentation, we exploit a KL divergence based consistency loss that is computed between the ASR outputs of original and enhanced utterances. In using speech enhancement as an effective ASR frontend, we propose a three-step training scheme based on model pretraining and feature selection. We evaluate our proposed techniques on a challenging social media English video dataset, and achieve an average relative improvement of 11.2% with speech enhancement based data augmentation, 8.3% with enhancement based preprocessing, and 13.4% when combining both.
Ashutosh Pandey 0004, Chunxi Liu, Yatharth Saraf
SLT2
2021 Benchmarking LF-MMI, CTC And RNN-T Criteria For Streaming ASR
abstract
In this work, to measure the accuracy and efficiency for a latency-controlled streaming automatic speech recognition (ASR) application, we perform comprehensive evaluations on three popular training criteria: LF-MMI, CTC and RNN-T. In transcribing social media videos of 7 languages with training data 3K - 14K hours, we conduct large-scale controlled experimentation across each criterion using identical datasets and encoder model architecture. We find that RNN-T has consistent wins in ASR accuracy, while CTC models excel at inference efficiency. Moreover, we selectively examine various modeling strategies for different training criteria, including modeling units, encoder architectures, pre-training, etc. Given such large-scale real-world streaming ASR application, to our best knowledge, we present the first comprehensive benchmark on these three widely used training criteria across a great many languages.
Xiaohui Zhang 0007, Frank Zhang 0001, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, Geoffrey Zweig
SLT3
2020 DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks
abstract
Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-examine their input features in light of partial hypotheses we introduce intermediate model heads and loss function. We study this architecture in the context of deep Transformer networks, and we use an attention mechanism over both the previous layer activations and the input features. To train this model's intermediate output hypothesis, we apply the objective function at each layer right before feature re-use. We find that the use of such iterated loss significantly improves performance by itself, as well as enabling input feature re-use. We present results on both Librispeech, and a large scale video dataset, with relative improvements of 10 - 20% for Librispeech and 3.2 - 13% for videos.
Andros Tjandra, Chunxi Liu, Frank Zhang 0001, Xiaohui Zhang 0007, Yongqiang Wang 0005, Gabriel Synnaeve, Satoshi Nakamura 0001, Geoffrey Zweig
ICASSP2
2020 Transformer-Based Acoustic Modeling for Hybrid Speech Recognition
abstract
We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using limited right context in transformer models, which makes it possible for streaming applications. We demonstrate that on the widely used Librispeech benchmark, our transformer-based AM outperforms the best published hybrid result by 19% to 26% relative when the standard n-gram language model (LM) is used. Combined with neural network LM for rescoring, our proposed approach achieves state-of-the-art results on Librispeech. Our findings are also confirmed on a much larger internal dataset.
Yongqiang Wang 0005, Abdel-rahman Mohamed, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang 0007, Frank Zhang 0001, Christian Fügen, Geoffrey Zweig, Michael L. Seltzer
ICASSP4
2020 Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model
abstract
Videos uploaded on social media are often accompanied with textual descriptions. In building automatic speech recognition (ASR) systems for videos, we can exploit the contextual information provided by such video metadata. In this paper, we explore ASR lattice rescoring by selectively attending to the video descriptions. We first use an attention based method to extract contextual vector representations of video metadata, and use these representations as part of the inputs to a neural language model during lattice rescoring. Secondly, we propose a hybrid pointer network approach to explicitly interpolate the word probabilities of the word occurrences in metadata. We perform experimental evaluations on both language modeling and ASR tasks, and demonstrate that both proposed methods provide performance improvements by selectively leveraging the video metadata.
Da-Rong Liu, Chunxi Liu, Frank Zhang 0001, Gabriel Synnaeve, Yatharth Saraf, Geoffrey Zweig
INTERSPEECH2
2020 Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
abstract
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state-ofthe-art results.We then show that using wordpieces as modeling units combined with CTC training, we can greatly simplify the engineering pipeline compared to conventional frame-based cross-entropy training by excluding all the GMM bootstrapping, decision tree building and force alignment steps, while still achieving very competitive word-error-rate.Additionally, using wordpieces as modeling units can significantly improve runtime efficiency since we can use larger stride without losing accuracy.We further confirm these findings on two internal VideoASR datasets: German, which is similar to English as a fusional language, and Turkish, which is an agglutinative language.
Frank Zhang 0001, Yongqiang Wang 0005, Xiaohui Zhang 0007, Chunxi Liu, Yatharth Saraf, Geoffrey Zweig
INTERSPEECH4
2019 Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings
abstract
We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because they depend on having sufficient target-side text to train the attention and decoder networks. In this paper we address this shortcoming by pretraining our network parameters using only text-based data and transcribed speech from other languages. We analyze the relative contributions of both sources of data. Across 3 test languages, our text-based approach resulted in a 20% average relative improvement over a text-based augmentation technique without pretraining. Using transcribed speech from nearby languages gives a further 20-30% relative reduction in character error rate.
Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe 0001, Chunxi Liu, Najim Dehak, Sanjeev Khudanpur
INTERSPEECH4
2018 Automatic Speech Recognition and Topic Identification from Speech for Almost-Zero-Resource Languages
Matthew Wiesner, Chunxi Liu, Lucas Ondel Yang, Craig Harman, Vimal Manohar, Jan Trmal, Zhongqiang Huang, Najim Dehak, Sanjeev Khudanpur
INTERSPEECH2
2018 Low-Resource Contextual Topic Identification on Speech
abstract
In topic identification (topic ID) on real-world unstructured audio, an audio instance of variable topic shifts is first broken into sequential segments, and each segment is independently classified. We first present a general purpose method for topic ID on spoken segments in low-resource languages, using a cascade of universal acoustic modeling, translation lexicons to English, and English-language topic classification. Next, instead of classifying each segment independently, we demonstrate that exploring the contextual dependencies across sequential segments can provide large improvements. In particular, we propose an attention-based contextual model which is able to leverage the contexts in a selective manner. We test both our contextual and non-contextual models on four LORELEI languages, and on all but one our attention-based contextual model significantly outperforms the context-independent models.
Chunxi Liu, Matthew Wiesner, Shinji Watanabe 0001, Craig Harman, Jan Trmal, Najim Dehak, Sanjeev Khudanpur
SLT1
2017 An empirical evaluation of zero resource acoustic unit discovery
abstract
Acoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provided linguistic knowledge and transcribed speech are unavailable. Therefore, to further facilitate zero-resource AUD process, in this paper, we demonstrate acoustic feature representations can be significantly improved by (i) performing linear discriminant analysis (LDA) in an unsupervised self-trained fashion, and (ii) leveraging resources of other languages through building a multilingual bottleneck (BN) feature extractor to give effective cross-lingual generalization. Moreover, we perform comprehensive evaluations of AUD efficacy on multiple downstream speech applications, and their correlated performance suggests that AUD evaluations are feasible using different alternative language resources when only a subset of these evaluation resources can be available in typical zero resource applications.
Chunxi Liu, Jinyi Yang, Santosh Kesiraju, Alena Rott, Lucas Ondel Yang, Pegah Ghahremani, Najim Dehak, Lukás Burget, Sanjeev Khudanpur
ICASSP1
2017 Topic Identification for Speech Without ASR
abstract
Modern topic identification (topic ID) systems for speech use automatic speech recognition (ASR) to produce speech transcripts, and perform supervised classification on such ASR outputs. However, under resource-limited conditions, the manually transcribed speech required to develop standard ASR systems can be severely limited or unavailable. In this paper, we investigate alternative unsupervised solutions to obtaining tokenizations of speech in terms of a vocabulary of automatically discovered word-like or phoneme-like units, without depending on the supervised training of ASR systems. Moreover, using automatic phoneme-like tokenizations, we demonstrate that a convolutional neural network based framework for learning spoken document representations provides competitive performance compared to a standard bag-of-words representation, as evidenced by comprehensive topic ID evaluations on both single-label and multi-label classification tasks.
Chunxi Liu, Jan Trmal, Matthew Wiesner, Craig Harman, Sanjeev Khudanpur
INTERSPEECH1
2017 ASR for Under-Resourced Languages From Probabilistic Transcription
abstract
In many under-resourced languages it is possible to find text, and it is possible to find speech, but transcribed speech suitable for training automatic speech recognition (ASR) is unavailable. In the absence of native transcripts, this paper proposes the use of a probabilistic transcript: A probability mass function over possible phonetic transcripts of the waveform. Three sources of probabilistic transcripts are demonstrated. First, self-training is a well-established semisupervised learning technique, in which a cross-lingual ASR first labels unlabeled speech, and is then adapted using the same labels. Second, mismatched crowdsourcing is a recent technique in which nonspeakers of the language are asked to write what they hear, and their nonsense transcripts are decoded using noisy channel models of second-language speech perception. Third, EEG distribution coding is a new technique in which nonspeakers of the language listen to it, and their electrocortical response signals are interpreted to indicate probabilities. ASR was trained in four languages without native transcripts. Adaptation using mismatched crowdsourcing significantly outperformed self-training, and both significantly outperformed a cross-lingual baseline. Both EEG distribution coding and text-derived phone language models were shown to improve the quality of probabilistic transcripts derived from mismatched crowdsourcing.
Mark Hasegawa-Johnson, Preethi Jyothi, Daniel McCloy, Majid Mirbagheri, Giovanni M. Di Liberto, Amit Das 0007, Bradley Ekin, Chunxi Liu, Vimal Manohar, Hao Tang 0002, Edmund C. Lalor, Nancy F. Chen, Paul Hager, Tyler Kekona, Rose Sloan, Adrian K. C. Lee
IEEE ACM Trans. Audio Speech Lang. Process.8
2016 Context-dependent point process models for keyword search and detection-based ASR
abstract
The point process model (PPM) for keyword search (KWS) is a whole-word parametric approach that characterizes each query type by the timing of phonetic events observed during its production. In this paper, we first extend the PPM modeling framework to operate on context-dependent phonetic event patterns instead of monophone patterns considered in the past, which provides significant KWS improvements. Second, we use the context-dependent PPMs to drive a detection-based speech recognition architecture thats runs parallel word detectors covering the whole vocabulary and uses the independent detections to construct lattices that can be used for both KWS indexing and LVCSR decoding. This strategy produces significant improvements over the original PPM KWS framework and provides an encouraging first attempt at PPM-based LVCSR.
Chunxi Liu, Aren Jansen, Sanjeev Khudanpur
ICASSP1
2016 Adapting ASR for under-resourced languages using mismatched transcriptions
abstract
Mismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages, we use less than an hour of mismatched transcriptions to successfully adapt baseline multilingual models built with no access to native transcriptions in the target language. The adapted models provide up to 25% relative improvement in phone error rates on an unseen evaluation set.
Chunxi Liu, Preethi Jyothi, Hao Tang 0002, Vimal Manohar, Rose Sloan, Tyler Kekona, Mark Hasegawa-Johnson, Sanjeev Khudanpur
ICASSP1
2015 Deep contextual language understanding in spoken dialogue systems
abstract
We describe a unified multi-turn multi-task spoken language understanding (SLU) solution capable of handling multiple context sensitive classification (intent determination) and sequence labeling (slot filling) tasks simultaneously. The proposed architecture is based on recurrent convolutional neural networks (RCNN) with shared feature layers and globally normalized sequence modeling components. The temporal dependencies within and across different tasks are encoded succinctly as recurrent connections. The dialog system responses beyond SLU component are also exploited as effective external features. We show with extensive experiments on a number of datasets that the proposed joint learning framework generates state-of-the-art results for both classification and tagging, and the contextual modeling based on recurrent and external features significantly improves the context sensitivity of SLU models.
Chunxi Liu, Puyang Xu, Ruhi Sarikaya
INTERSPEECH1
2014 Low-resource open vocabulary keyword search using point process models
abstract
The point process model (PPM) for keyword search is a wholeword parametric modeling framework based on the timing of phonetic events rather than the evolution of frame-level phonetic likelihoods. Recent progress in PPM training and decoding algorithms has yielded state-of-the-art phonetic search performance in high-resource settings, both in terms of accuracy and computational efficiency. In this paper, we consider PPM application to low-resource settings where the amount of transcribed speech is severely limited and the pronunciation dictionary is incomplete. By using (i) state-of-the-art deep neural network acoustic models to generate phonetic events and (ii) grapheme-to-phoneme conversion to generate pronunciations for out-of-vocabulary (OOV) keywords, we find the PPM system reaches state-of-the-art OOV search performance at a small computational cost. Moreover, due to their complementary methodologies, combining PPM outputs with the LVCSR baseline produces average relative ATWV improvements of 7% and 50% for in-vocabulary and OOV keywords, respectively (16% overall).
Chunxi Liu, Aren Jansen, Guoguo Chen, Keith Kintzley, Jan Trmal, Sanjeev Khudanpur
INTERSPEECH1
2014 A keyword search system using open source software
abstract
Provides an overview of a speech-to-text (STT) and keyword search (KWS) system architecture build primarily on the top of the Kaldi toolkit and expands on a few highlights. The system was developed as a part of the research efforts of the Radical team while participating in the IARPA Babel program. Our aim was to develop a general system pipeline which could be easily and rapidly deployed in any language, independently on the language script and phonological and linguistic features of the language.
Jan Trmal, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur, Pegah Ghahremani, Xiaohui Zhang 0007, Vimal Manohar, Chunxi Liu, Aren Jansen, Dietrich Klakow, David Yarowsky, Florian Metze
SLT8
2014 Web video thumbnail recommendation with content-aware analysis and query-sensitive matching
Weigang Zhang, Chunxi Liu, Zhenjun Wang, Guorong Li, Qingming Huang, Wen Gao 0001
Multim. Tools Appl.2
2012 Motion Based Perceptual Distortion and Rate Optimization for Video Coding
abstract
Most conventional distortion metrics regard a video frame as a static image, and seldom exploit using the motion information of video frames in succession. Moreover, these methods usually calculate the visual distortion based on the independent spatial pixels. Recently, many researches show that the way people perceive the video signals is similar to the way filters process signals in the frequency domain. Therefore, in order to achieve better visual quality, we introduce a novel distortion measurement into the video coding system, which is consistent with human visual perception, and establish a perception-based rate-distortion optimization model. In this paper, we adopt Gabor filter family to decompose the video signals into frequency domain, and combine the video motion information to measure the perceptual distortion. We call it Motion tuned Distortion metric For Video coding (MDFV). After that we set up an MDFV based rate-distortion optimization model to select the best encoding mode. The experimental results show that the proposed approach is effective.
Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu, Ling-Yu Duan
ICME4
2012 An effective multi-clue fusion approach for web video topic detection
abstract
The efficient organization and navigation of web videos in the topic level could enhance the user experience and boost the user's understanding about the happened events. Due to the potential application prospects, topic detection attracts increasing research interests in the last decade. On one hand, the user concerned real world hot topic always leads to a massive discussion in the video sharing sites, such as YouTube, Youku, etc. On the other hand, the search volume of the topic related keywords are growing explosively in the search engine such as Google, Yahoo, etc. These keywords are the queries formulated by the users to search their concerned topics. They reflect the users' intention and could be used as a clue to find the hot topics. In this paper, different from the traditional topic detection methods, which mainly rely on data clustering, we propose a novel multi-clue fusion approach for web video topic detection. In our approach, firstly by utilizing the video related tag information, a maximum average score and a burstiness degree are proposed to extract the dense-bursty tag groups. Secondly, the near-duplicate keyframes (NDK) are extracted from the videos and fused with the extracted tag groups. After that, the hot search keywords from the search engine are used as guidance for topic detection. Finally, these clues are combined together to detect the topics hidden in the web video data. Experiment is conducted on the YouTube video data and the results demonstrate that the proposed method is effective.
Tianlong Chen 0003, Chunxi Liu, Qingming Huang
ACM Multimedia2
2011 Query sensitive dynamic web video thumbnail generation
abstract
With the fast rising of the video sharing websites, the online video becomes an important media for people to share messages, interests, ideas, beliefs, etc. In this paper, we propose a novel approach to dynamically generate the web video thumbnails according to user's query. Two issues are addressed: the video content representativeness of the selected video thumbnail, and the relationship between the selected video thumbnail and the user's query. For the first issue the reinforcement based algorithm is adopted to rank the frames in each video. For the second issue the relevance model based method is employed to calculate the similarity between the video frames and the query keywords. The final video thumbnail is generated by linear fusion of the above two scores. Compared with the existing web video thumbnails, which only reflect the preference of the video owner, the thumbnails generated in our approach not only consider the video content representativeness of the frame, but also reflect the intention of the video searcher. In order to show the effectiveness of the proposed method, experiments are conducted on the videos selected from the video sharing website. Experimental results and subjective evaluations demonstrate that the proposed method is effective and can meet the user's intention requirement.
Chunxi Liu, Qingming Huang, Shuqiang Jiang
ICIP1
2011 Visual perception based Lagrangian rate distortion optimization for video coding
abstract
In the conventional rate distortion optimization (RDO) video coding, the measure of distortion is mainly from the perspective of signal processing, while dose not fully take into account the characteristics of visual perception. People have concerns about not only the information of independent pixels, but also the temporal and spatial correlations between them. For different video content, human visual perception has different sensitivity. In this paper, in order to establish a RDO model which is more consistent with the human visual perception, we introduce the structural similarity and the content saliency information into the distortion metric. An adaptive Lagrange multiplier selection scheme is presented to allocate the bit resources more rationally by keeping the balance of the bit-rate and the visual quality. Experimental results show that the proposed method averagely reduces 10.14% bit-rate under the similar visual quality.
Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu
ICIP4
2011 News video story sentiment classification and ranking
abstract
In this paper, we present a novel approach for news video story sentiment analysis. Two research challenges are addressed: news video story sentiment classification and ranking. For classification, a graph based semi-supervised learning approach is utilized to classify the news stories into sentiment classes. Graph based semi-supervised learning is able to tackle the problem of lacking labeled data. After classification, two sentiment classes are obtained: positive and negative. In order to project the news videos into sentiment space, a multimodal approach by fusing the text sentiment and visual representation scores is adopted to rank the videos in each class. For sentiment representation, inter and intra sentiment class analysis is conducted based on affinity propagation clustering and PageRank algorithm. A user study is conducted to evaluate the video ranking performance. The experimental results on the selected topics are promising and demonstrate the proposed approach is effective.
Chunxi Liu, Li Su 0003, Qingming Huang, Shuqiang Jiang
ICME1
2010 Event based news video people classification and ranking using multimodality features
abstract
Existing research on news video analysis mainly concentrates on structure analysis, semantic concept detection, annotation and search. However, little work has been contributed to news video people community analysis, which is helpful for users to understand the event. In this paper, we propose a novel approach to classify the people appearing in the news video into different communities. In our approach, the people appearing in the news video are first identified by associating their faces with names. The faces are detected from the video frames, and the names are obtained from the text. Then, the people belonging to the same organization are clustered. After that, the relationships between these organizations are determined using sentiment analysis. The sentiment words are diverse in each news story and contain both positive and negative ones. However, we have news title, which is the summary of the story and the sentiment of which is clear, to help us to mine the relationships between the organizations. At last, social networks are built to classify those people/organizations into different classes, and the people/organizations are ranked in each community according to their influence. The main contributions of the paper are two folds: 1) we propose a novel approach to present the news video event according to communities; 2) we propose to use the sentiment analysis and social network to classify the news people/organizations. The experimental results on the selected news topics demonstrate that the proposed approach is effective.
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Changsheng Xu
ICME1
2010 The third eye: mining the visual cognition across multi-language communities
abstract
Existing research work in the multimedia domain mainly focuses on image/video indexing, retrieval, annotation, tagging, re-ranking, etc. However, little work has been contributed to people's visual cognition. In this paper, we propose a novel framework to mine people's visual cognition across multi-language communities. Two challenges are addressed: the visual cognition representation for a specific language community, and the visual cognition comparison between different language communities. We call it "the third eye", which means that through this way people with different backgrounds can better understand the cognition of each other, and can view the concept more objectively to avoid culture conflict. In this study, we utilize the image search engine to mine the visual cognition of the different communities. The assumption is that the image semantic distribution over the search results can reflect the visual cognition of the community. When a user submits a text query, it is first translated into different languages, and fed into the corresponding image search engine ports to retrieve images from these communities. After retrieval, the obtained images are categorized into different semantic clusters automatically. Finally, inter semantic cluster ranking is employed to rank the semantic clusters according to their relationship to the query, and intra cluster ranking is used to rank the images according to their representativeness. The visual cognition difference among these language communities is achieved by comparing the different community image distributions over these semantic clusters. The experimental results are promising and show that the proposed visual cognition mining approach is effective.
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Changsheng Xu
ACM Multimedia1
2009 A framework for flexible summarization of racquet sports video using multiple modalities
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Liyuan Xing, Qixiang Ye, Wen Gao 0001
Comput. Vis. Image Underst.1
2008 Naming faces in broadcast news video by image google
abstract
Naming faces is important for news videos browsing and indexing. Although some research efforts have been contributed to it, they only use the concurrent information between the face and name or employ some clues as features and use simple heuristic method or machine learning approach to finish the task. They use little extra knowledge about the names and faces. Different from previous work, in this paper we present a novel approach to name the faces by exploring extra knowledge obtained from image google. The behind assumption is that the faces of those important persons will turn out many times in the web images and could be retrieved from image google easily. Firstly, faces are detected in the video frames; and the name entities of candidate persons are extracted from the textual information by automatic speech recognition and close caption detection. Then, these candidate person names are used as queries to find the name related person images through image google. After that, the retrieved result is analyzed and some typical faces are selected through feature density estimation. Finally, the detected faces in the news video are matched with the faces selected from the result returned by image google to label each face. Experimental results on MSNBC news and CNN news demonstrate that the proposed approach is effective.
Chunxi Liu, Shuqiang Jiang, Qingming Huang
ACM Multimedia1
2006 Extracting Story Units in Sports Video Based on Unsupervised Video Scene Clustering
abstract
Many sports videos such as archery, diving and tennis have repetitive structure patterns. They are reliable clues to generate highlights, summarization and automatic annotation. In this paper, we present a novel approach to analyze these structure patterns in sports video to extract story units. First, an unsupervised scene clustering method for sports video is adopted to automatically categorize the video shots into several disparate scenes. Then, the clustering results are modeled by a transition matrix. Finally, the key scene shots are detected to analyze the structure patterns and extract the story units. Experimental results on several types of broadcast sports video demonstrate that our approach is effective
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Weigang Zhang
ICME1