Lie Lu

dblp:23/2390 · DBLP profile ↗
← Back
65ranked-venue papers
22as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 18 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Rethinking Transparent TCP Replacement: Practical Lessons from SMC-R in the Cloud
abstract
Transparent TCP acceleration has been a subject of extensive academic research for years, aiming to bypass kernel overhead without modifying legacy applications. Prior work, however, largely emphasizes performance under idealized settings and provides limited discussion of how costs and gains manifest in real production environments.
Wen Gu, Guangguan Wang, Lie Lu, Jinhu Li
SIGCOMM6
2025 CiTrus: Squeezing Extra Performance out of Low-data Bio-signal Transfer Learning
abstract
Transfer learning for bio-signals has recently become an important technique to improve prediction performance on downstream tasks with small bio-signal datasets. Recent works have shown that pre-training a neural network model on a large dataset (e.g. EEG) with a self-supervised task, replacing the self-supervised head with a linear classification head, and fine-tuning the model on different downstream bio-signal datasets (e.g., EMG or ECG) can dramatically improve the performance on those datasets. In this paper, we propose a new convolution-transformer hybrid model architecture with masked auto-encoding for low-data bio-signal transfer learning, introduce a frequency-based masked auto-encoding task, employ a more comprehensive evaluation framework, and evaluate how much and when (multimodal) pre-training improves fine-tuning performance. We also introduce a dramatically more performant method of aligning a downstream dataset with a different temporal length and sampling rate to the original pre-training dataset. Our findings indicate that the convolution-only part of our hybrid model can achieve state-of-the-art performance on some low-data downstream tasks. The performance is often improved even further with our full model. In the case of transformer-based models we find that pre-training especially improves performance on downstream datasets, multimodal pre-training often increases those gains further, and our frequency-based pre-training performs the best on average for the lowest and highest data regimes.
Eloy Geenjaar, Lie Lu
AAAI2
2025 Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
abstract
Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best music for a video can be a difficult and time-consuming task. To address this challenge, we propose a novel framework for automatically retrieving a matching music clip for a given video, and vice versa. Our approach leverages annotated music labels, as well as the inherent artistic correspondence between visual and music elements. Distinct from previous cross-modal music retrieval works, our method combines both self-supervised and supervised training objectives. We use self-supervised and label-supervised contrastive learning to train a joint embedding space between music and video. We show the effectiveness of our approach by using music genre labels for the supervised training component, and our framework can be generalized to other music annotations (e.g., emotion, instrument, etc.). Furthermore, our method enables fine-grained control over how much the retrieval process focuses on self-supervised vs. label information at inference time. We evaluate the learned embeddings through a variety of video-to-music and music-to-video retrieval tasks. Our experiments show that the proposed approach successfully combines self-supervised and supervised objectives and is effective for controllable music-video retrieval.
Shanti Stewart, Gouthaman KV, Lie Lu, Andrea Fanelli
ICASSP3
2025 RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text
Jiaben Chen, Xin Yan 0008, Siyuan Cen, Qinwei Ma, Haoyu Zhen, Kaizhi Qian, Lie Lu, Chuang Gan 0001
ICCV9
2025 XAttnMark: Learning Robust Audio Watermarking with Cross-Attention
abstract
The rapid proliferation of generative audio synthesis and editing technologies has raised significant concerns about copyright infringement, data provenance, and the spread of misinformation through deepfake audio. Watermarking offers a proactive solution by embedding imperceptible, identifiable, and traceable marks into audio content. While recent neural network-based watermarking methods like WavMark and AudioSeal have improved robustness and quality, they struggle to achieve both robust detection and accurate attribution simultaneously. This paper introduces the Cross-Attention Robust Audio Watermark (XAttnMark), which bridges this gap by leveraging partial parameter sharing between the generator and the detector, a cross-attention mechanism for efficient message retrieval, and a temporal conditioning module for improved message distribution. Additionally, we propose a psychoacoustic-aligned temporal-frequency masking loss that captures fine-grained auditory masking effects, enhancing watermark imperceptibility. Our approach achieves state-of-the-art performance in both detection and attribution, demonstrating superior robustness against a wide range of audio transformations, including challenging generative editing with strong editing strength. This work represents a significant step forward in protecting intellectual property and ensuring the authenticity of audio content in the era of generative AI.
Yixin Liu 0002, Lie Lu, Jihui Jin, Lichao Sun 0001, Andrea Fanelli
ICML2
2025 Transformation of audio embeddings into interpretable, concept-based representations
abstract
Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.
Alice Zhang 0002, Edison Thomaz, Lie Lu
IJCNN3
2025 Predicting subjective audiovisual quality of experience with oculomotor responses and perceptual video quality
abstract
Estimating user quality of experience (QoE) for audiovisual media is essential for service providers, but it is generally infeasible to monitor self-reporting within consumer applications. We examine how eye-tracking data correlates with viewers’ subjective quality of experience during audiovisual media consumption. Using measured fixation sequences, we implement a foveated visible differences predictor model (FovVideoVDP) to quantify just-objectionable-differences (JODs). Subjective QoEs correlated with fixation rate, potentially indicating increased engagement with higher bitrates. Foveated JODs aligned closely with mean opinion scores, comparing favorably to the most commonly used video quality metrics.
Ekin Tünçok, Neel Chaudhari, Nathan Swedlow, Stephanie Reeves, Lie Lu, Scott Daly, Daniel P. Darcy
QoMEX5
2025 Stateless and Proactive Routing for Dynamic Multicast With Deep Reinforcement Learning
abstract
Stateful multicast protocols manage multicast group memberships by maintaining state information about active groups and their members. They have seen limited adoption in the modern internet due to lack of scalability, simplicity, and flexibility. Although stateless multicast protocols, like BIER, eliminate extensive state management, they still face complex tree computation and limited scalability for concurrent requests. In this paper, we propose Hawkeye, a stateless multicast mechanism with deep reinforcement learning (DRL) for real-time responses to dynamic multicast requests with near-optimal multicast TE performance. This mechanism is suited for Software-Defined Networking (SDN) environment where the controller has a global view of the network and supports flexible configuration of network resources for traffic engineering. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership, and thus are able to build multicast trees proactively for upcoming requests. We develop a novel source aggregation mechanism to facilitate the convergence of the DRL agent under high volume of multicast requests. Moreover, to improve the practicality and robustness of Hawkeye, we design incremental deployment and single failure handling mechanisms, which take advantages of source aggregation and fit well with multicast routing. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye responds effectively to dynamic multicast requests. Itoffers rapid routing decisions, e.g., making routing decisions in under 5ms on a tested topology, and reduces path latency variation by up to 89.5%, with less than a 10% increase in bandwidth consumption compared to the offline theoretical minimum.
Qing Li 0006, Lie Lu, Dan Zhao 0003, Zeyu Luan, Yuan Yang 0001, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Shaoteng Liu, Dingding Chen
IEEE Trans. Netw.2
2024 Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
abstract
A "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indistinguishable transition between two shots. In this paper, we explore the ability to automatically find and create "audio match cuts" within videos and movies. We create a self-supervised audio representation for audio match cutting and develop a coarse-to-fine audio match pipeline that recommends matching shots and creates the blended audio. We further annotate a dataset for the proposed audio match cut task and compare the ability of multiple audio representations to find audio match cut candidates. Finally, we evaluate multiple methods to blend two matching audio candidates with the goal of creating a smooth transition. Project page and examples are available at: https://denfed.github.io/audiomatchcut/
Dennis Fedorishin, Lie Lu, Srirangaraj Setlur, Venu Govindaraju
ICASSP2
2023 High Quality Audio Coding with Mdctnet
abstract
We propose a neural audio generative model, MDCTNet, operating in the perceptually weighted domain of an adaptive modified discrete cosine transform (MDCT). The architecture of the model captures correlations in both time and frequency directions with recurrent layers (RNNs). An audio coding system is obtained by training MDCTNet on a diverse set of fullband monophonic audio signals at 48 kHz sampling, conditioned by a perceptual audio encoder. In a subjective listening test with ten excerpts chosen to be balanced across content types, yet critical for both codecs, the mean performance of the proposed system for 24 kb/s variable bitrate (VBR) is similar to that of Opus at twice the bitrate.
Grant A. Davidson, Mark Vinton, Per Ekstrand, Lars F. Villemoes, Lie Lu
ICASSP6
2023 Deepspace: Dynamic Spatial and Source CUE Based Source Separation for Dialog Enhancement
abstract
Dialog Enhancement (DE) is a feature which allows a user to increase the level of dialog in TV or movie content relative to nondialog sounds. When only the original mix is available, DE is "unguided," and requires source separation. In this paper, we describe the DeepSpace system, which performs source separation using both dynamic spatial cues and source cues to support unguided DE. Its technologies include spatio-level filtering (SLF) and deep-learning based dialog classification and denoising. Using subjective listening tests, we show that DeepSpace demonstrates significantly improved overall performance relative to state-of-the-art systems available for testing. We explore the feasibility of using existing automated metrics to evaluate unguided DE systems.
Aaron Master, Lie Lu, Jonas Samuelsson, Heidi-Maria Lehtonen, Scott Norcross, Nathan Swedlow, Audrey Howard
ICASSP2
2023 Hawkeye: A Dynamic and Stateless Multicast Mechanism with Deep Reinforcement Learning
abstract
Multicast traffic is growing rapidly due to the development of multimedia streaming. Lately, stateless multicast protocols, such as BIER, have been proposed to solve the excessive routing states problem of traditional multicast protocols. However, the high complexity of multicast tree computation and the limited scalability for concurrent requests still pose daunting challenges, especially under dynamic group membership. In this paper, we propose Hawkeye, a dynamic and stateless multicast mechanism with deep reinforcement learning (DRL) approach. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership and thus is able to build multicast trees proactively for upcoming requests. Moreover, an innovative source aggregation mechanism is designed to help the DRL agent converge when faced with a large amount of multicast requests, and relieve ingress routers from excessive routing states. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye adapts well to dynamic multicast: it reduces the variation of path latency by up to 89.5% with less than 12% additional bandwidth consumption compared with the theoretical optimum.
Lie Lu, Qing Li 0006, Dan Zhao 0003, Yuan Yang 0001, Zeyu Luan, Jianer Zhou, Yong Jiang 0001, Mingwei Xu 0001
INFOCOM1
2021 EPC-TE: Explicit Path Control in Traffic Engineering with Deep Reinforcement Learning
abstract
Segment Routing (SR) provides Traffic Engineering (TE) with Explicit Path Control (EPC) by steering data flows passing through a list of SR routers along a desired path. However, large-scale migration from a pure IP network to a full SR one requires prohibitive hardware replacement and software update. Therefore, network operators prefer to upgrade a subset of IP routers into SR routers during a transitional period. This paper proposes EPC-TE to optimize TE performance in hybrid IP/SR networks where partially deployed SR routers coexist with legacy IP routers. We propose a concept of key nodes to achieve EPC over desired paths and a criterion to select which IP routers to upgrade first under a pre-defined upgrading ratio. EPC-TE leverages Deep Reinforcement Learning (DRL) to inference the optimal traffic splitting ratio across multiple controllable paths between source-destination pairs. EPC-TE can achieve comparable TE performance as a full SR network with an upgrading ratio less than 30%. Extensive experimental results with real-world topologies show that EPC-TE significantly outperforms other baseline TE solutions in minimizing maximum link utilization.
Zeyu Luan, Lie Lu, Qing Li 0006, Yong Jiang 0001
GLOBECOM2
2012 Music/speech classification using high-level features derived from fmri brain imaging
abstract
With the availability of large amount of audio tracks through a variety of sources and distribution channels, automatic music/speech classification becomes an indispensable tool in social audio websites and online audio communities. However, the accuracy of current acoustic-based low-level feature classification methods is still rather far from satisfaction. The discrepancy between the limited descriptive power of low-level features and the richness of high-level semantics perceived by the human brain has become the 'bottleneck' problem in audio signal analysis. In this paper, functional magnetic resonance imaging (fMRI) which monitors the human brain's response under the natural stimulus of music/speech listening is used as high-level features in the brain imaging space (BIS). We developed a computational framework to model the relationships between BIS features and low-level features in the training dataset with fMRI scans, predict BIS features of testing dataset without fMRI scans, and use the predicted BIS features for music/speech classification in the application stage. Experimental results demonstrated the significantly improved performance of music/speech classification via predicted BIS features than that via the original low-level features.
Xi Jiang 0001, Xintao Hu, Lie Lu, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001
ACM Multimedia4
2010 Music rhythm characterization with application to workout-mix generation
abstract
In this paper, we present approaches to musical rhythm pattern extraction, rhythm-based music retrieval, and rhythm-synchronized music mixing. A probabilistic model is used to jointly estimate tempo and time signature as a basis for beat tracking and measure detection. A representative rhythm pattern is then extracted through clustering to characterize the rhythm of a song. Based on this, a probabilistic approach is used for retrieving songs with similar rhythmic patterns. These are then mixed rhythm-synchronously with transitions maintaining continuity and regularity of beats. We apply the presented methods into workout-mix generation, which aims at automatically selecting rhythmically similar music given a seed song and a user-defined tempo profile. Our probabilistic approaches achieve accuracies similar to best published results, but avoid manually tuned parameters and “fudge factors”.
Lie Lu, Christopher Weare, Frank Seide
ICASSP2
2009 Learning a music similarity measure on automatic annotations with application to playlist generation
abstract
This paper presents an approach to learn a better music similarity measure and presents an application to music playlist generation. Different from previous work, in our approach, automatically detected music attributes are used to represent each song. A set of kernels is employed in similarity measure, with each kernel measuring on a subset of music attributes and having a different importance weight. In automatic music playlist generation, a ranking method is presented, which considers multiple seed songs and possible outlier seed. Experiments show the effectiveness of the proposed approach, and the quality of the playlist generated based on automatic annotations is comparable to that based on manual annotations.
Linxing Xiao, Lie Lu, Frank Seide, Jie Zhou 0001
ICASSP2
2009 Text-Like Segmentation of General Audio for Content-Based Retrieval
abstract
Automatic detection of (semantically) meaningful audio segments, oraudioscenes, is an important step in high-level semantic inference from general audio signals, and can benefit various content-based applications involving both audio and multimodal (multimedia) data sets. Motivated by the known limitations of traditional low-level feature-based approaches, we propose in this paper a novel approach to discover audio scenes, based on an analysis ofaudioelementsandkeyaudioelements, which can be seen as equivalents to the words and keywords in a text document, respectively. In the proposed approach, an audio track is seen as a sequence of audio elements, and the presence of an audio scene boundary at a given time stamp is checked based on pair-wise measuring thesemanticaffinitybetween different parts of the analyzed audio stream surrounding that time stamp. Our proposed model for semantic affinity exploits the proven concepts from text document analysis, and is introduced here as a function of the distance between the audio parts considered, and the co-occurrence statistics and the importance weights of the audio elements contained therein. Experimental evaluation performed on a representative data set consisting of 5 h of diverse audio data streams indicated that the proposed approach is more effective than the traditional low-level feature-based approaches in solving the posed audio scene segmentation problem.
Lie Lu, Alan Hanjalic
IEEE Trans. Multim.1
2008 Unsupervised anchor space generation for similarity measurement of general audio
abstract
Reliably measuring similarity between audio clips is critical to many applications. As opposed to the conventional way of measuring audio similarity using low-level features directly, in this paper we consider the similarity computation using an anchor space. Each dimension of such a space corresponds to a semantic category (anchor). Mapping an audio clip onto this space results in a vector, which indicates the membership probability of this audio clip with respect to each semantic category. The more similar the mappings of two audio clips, the more similar they are. While an anchor space is typically generated in a supervised fashion, supervised approach is infeasible in many realistic scenarios where audio content semantics is too diverse or simply unknown a priori. We therefore propose an unsupervised approach to anchor space generation. There, spectral clustering is employed to cluster the audio clips with similar low-level features and then the obtained clusters are adopted as semantic categories. Using this semantic space for audio similarity computation shows a considerable accuracy improvement (7% on mAP) in an audio retrieval system, compared with the conventional low-level feature based approach.
Lie Lu, Alan Hanjalic
ICASSP1
2008 Mobile ringtone search through query by humming
abstract
In the context of voice-based mobile search, this paper presents a new approach to mobile ringtone search through query by humming: A user can call a service, hum a part of melody through the mobile phone, and obtain the ringtones or songs he or she is looking for. Correspondingly, we propose a method of query by humming tailored to this scenario. A robust front-end processing is first presented to deal with the mobile phone recording, which is distorted due to GSM codec, environment and wireless transmission. Then, a systematic probabilistic model and matching procedure inspired by Hidden Markov Model (HMM) is presented, by considering the alignment and error tolerance in the matching between query and songs. A rescoring heuristic is finally employed to further improve matching accuracy. Moreover, our system is evaluated on realistic mobile recordings from the field. Experiments show our approach can achieve 83% accuracy on a database with 3000 songs in this realistic scenario.
Lie Lu, Frank Seide
ICASSP1
2008 Audio tonality mode classification without tonic annotations
abstract
Traditional tonality mode (major or minor) classification or audio key finding algorithms often rely on tonic annotations (key names) of the training songs. However, unlike classical music whose keys are usually explicitly labeled in their titles, the keys of numerous popular music are hard to obtain. In contrast, it is much easier to only label the mode for each song. With only modes labeled, traditional approaches to key or mode classification cannot be directly applied, due to the lack of the reference point to transpose and align the chroma features with different keys. In this paper, we present an alignment approach to transpose chroma features within each mode to a reference (but unknown) tonic. Then several methods, including Single Profile Correlation, Multiple Profile Correlation and Support Vector Machine, are exploited to address mode learning and classification. Experimental results show the feasibility of the proposed approach.
Zhiyao Duan, Lie Lu, Changshui Zhang
ICME2
2008 Mobile Search With Multimodal Queries
abstract
The popularity of mobile devices, such as PDAs and SmartPhones, has grown rapidly over the last couple of years. Though most users still perform searches using desktop computers, it is expected that more and more people will also search the Web while they are on the move. In addition to text-based keyword queries, mobile devices can support richer and hybrid queries such as images, audio, video, and their combinations. In this paper, we will discuss mobile search systems that support image queries and audio queries, covering typical designs for mobile visual and audio search, as well as the opportunities and challenges. Specifically, we will present an in-depth study of two real systems we have developed: product image categorization and mobile ringtone search, which use image queries and audio queries, respectively. Experimental results on real-life data demonstrate their effectiveness and efficiency.
Xing Xie 0001, Lie Lu, Menglei Jia, Hua Li 0001, Frank Seide, Wei-Ying Ma
Proc. IEEE2
2008 Co-clustering for Auditory Scene Categorization
abstract
Auditory scenes are temporal audio segments with coherent semantic content. Automatically classifying and grouping auditory scenes with similar semantics into categories is beneficial for many multimedia applications, such as semantic event detection and indexing. For such semantic categorization, auditory scenes are first characterized with either low-level acoustic features or some mid-level representations like audio effects, and then supervised classifiers or unsupervised clustering algorithms are employed to group scene segments into various semantic categories. In this paper, we focus on the problem of automatically categorizing audio scenes in unsupervised manner. To achieve more reasonable clustering results, we introduce the co-clustering scheme to exploit potential grouping trends among different dimensions of feature spaces (either low-level or mid-level feature spaces), and provide more accurate similarity measure for comparing auditory scenes. Moreover, we also extend the co-clustering scheme with a strategy based on the Bayesian information criterion (BIC) to automatically estimate the numbers of clusters. Evaluation performed on 272 auditory scenes extracted from 12-h audio data shows very encouraging categorization results. Co-clustering achieved a better performance compared to some traditional one-way clustering algorithms, both based on the low-level acoustic features and on the mid-level audio effect representations. Finally, we present our vision regarding the applicability of this approach on general multimedia data, and also show some preliminary results on content-based image clustering.
Rui Cai 0002, Lie Lu, Alan Hanjalic
IEEE Trans. Multim.2
2008 Audio Keywords Discovery for Text-Like Audio Content Analysis and Retrieval
abstract
Inspired by classical text document analysis employing the concept of (key) words, this paper presents an unsupervised approach to discover (key) audio elements in general audio documents. The (key) audio elements can be considered the equivalents of the text (key) words, and enable content-based audio analysis and retrieval following the analogy to the proven text analysis theories and methods. Since general audio signals usually show complicated and strongly varying distribution and density in the feature space, we propose an iterative spectral clustering method with context-dependent scaling factors to decompose an audio data stream into audio elements. Using this clustering method, temporal signal segments with similar low-level features are grouped into natural clusters that we adopt as audio elements. To detect those audio elements that are most representative for the semantic content, that is, the key audio elements, two cases are considered. First, if only one audio document is available for analysis, a number of heuristic importance indicators are defined and employed to detect the key audio elements. For the case that multiple audio documents are available, more sophisticated measures for audio element importance, including expected term frequency (ETF), expected inverse document frequency (EIDF), expected term duration (ETD) and expected inverse document duration (EIDD), are proposed. Our experiments showed encouraging results regarding the quality of the obtained (key) audio elements and their potential applicability for content-based audio document analysis and retrieval.
Lie Lu, Alan Hanjalic
IEEE Trans. Multim.1
2006 Audio Elements Based Auditory Scene Segmentation
abstract
Auditory scene segmentation is an important step in the process of high-level semantic inference from audio data streams, and in particular, a prerequisite for auditory scene categorization. In this paper, we analyze the limits of previous works on auditory scene segmentation, and then propose a novel method that, conceptually, is inspired by the ideas used in text and video scene segmentation, and is based on an analysis of audio elements and key audio elements, which can be seen as equivalents to the words and keywords in a text document, respectively. Experiments performed on 1.5 hours of audio data indicate that the proposed approach is promising.
Lie Lu, Rui Cai 0002, Alan Hanjalic
ICASSP (5)1
2006 Towards optimal audio "keywords" detection for audio content analysis and discovery
abstract
Natural semantic sound clusters in an audio document, also referred to as audio elements, can be seen as an analogy to words in a text document. Based on the obtained set of audio elements, the key audio elements, or audio "keywords", can be detected, which are most prominent in characterizing the content of audio data. As such, they can be of great use for automatic audio content analysis and discovery. Motivated by the limitations of the existing methods for key audio element detection, we propose in this paper a novel unsupervised approach to audio elements weighting using multiple audio documents, analog to word weighting in text document analysis. In our approach, dominant feature vectors (DFV) are first extracted from each audio element, and used to measure the audio elements similarity, based on which the occurrence probability of one audio element in different audio documents can be estimated. Then, four factors, including expected term frequency, expected inverse document frequency, expected term duration, and expected inverse document duration, are calculated and combined to give the importance weight of each audio element. Evaluation of the obtained audio "keywords" and their usability for auditory scene segmentation and audio document clustering, performed on 5 hours of diverse audio data, shows highly promising results.
Lie Lu, Alan Hanjalic
ACM Multimedia1
2006 A flexible framework for key audio effects detection and auditory context inference
abstract
Key audio effects are those special effects that play critical roles in human's perception of an auditory context in audiovisual materials. Based on key audio effects, high-level semantic inference can be carried out to facilitate various content-based analysis applications, such as highlight extraction and video summarization. In this paper, a flexible framework is proposed for key audio effect detection in a continuous audio stream, as well as for the semantic inference of an auditory context. In the proposed framework, key audio effects and the background sounds are comprehensively modeled with hidden Markov models, and a Grammar Network is proposed to connect various models to fully explore the transitions among them. Moreover, a set of new spectral features are employed to improve the representation of each audio effect and the discrimination among various effects. The framework is convenient to add or remove target audio effects in various applications. Based on the obtained key effect sequence, a Bayesian network-based approach is proposed to further discover the high-level semantics of an auditory context by integrating prior knowledge and statistical learning. Evaluations on 12 h of audio data indicate that the proposed framework can achieve satisfying results, both on key audio effect detection and auditory context inference.
Rui Cai 0002, Lie Lu, Alan Hanjalic, HongJiang Zhang, Lianhong Cai
IEEE Trans. Speech Audio Process.2
2006 Automatic mood detection and tracking of music audio signals
abstract
Music mood describes the inherent emotional expression of a music clip. It is helpful in music understanding, music retrieval, and some other music-related applications. In this paper, a hierarchical framework is presented to automate the task of mood detection from acoustic music data, by following some music psychological theories in western cultures. The hierarchical framework has the advantage of emphasizing the most suitable features in different detection tasks. Three feature sets, including intensity, timbre, and rhythm are extracted to represent the characteristics of a music clip. The intensity feature set is represented by the energy in each subband, the timbre feature set is composed of the spectral shape features and spectral contrast features, and the rhythm feature set indicates three aspects that are closely related with an individual's mood response, including rhythm strength, rhythm regularity, and tempo. Furthermore, since mood is usually changeable in an entire piece of classical music, the approach to mood detection is extended to mood tracking for a music piece, by dividing the music into several independent segments, each of which contains a homogeneous emotional expression. Preliminary evaluations indicate that the proposed algorithms produce satisfactory results. On our testing database composed of 800 representative music clips, the average accuracy of mood detection achieves up to 86.3%. We can also on average recall 84.1% of the mood boundaries from nine testing music pieces.
Lie Lu, Dan Liu 0001, HongJiang Zhang
IEEE Trans. Speech Audio Process.1
2006 Photo2Video - A System for Automatically Converting Photographic Series Into Video
abstract
A novel method for browsing and sharing a single and series of photographs is presented, which can be regarded as a system exploring a new medium type between photograph and video. The scheme exploits the rich content embedded in a single photograph, as well as in a photographic series. Based on studying the typical process of a viewer's attention to variations on objects or regions in an image, a photograph can be converted into a motion clip by simulating camera motions on it. For a selected photographic series, an appropriate set of key-frames are determined for each photograph based on content analyses. And then camera motion pattern is selected for each photograph to generate a corresponding motion photograph clip. Finally, the final output video is rendered by connecting a series of motion photograph clips with specific transitions based on the content of the images on either side of the transition. Also, each motion photograph clip is aligned with the selected incidental music based on music content analysis. As the system, named Photo2Video, generates motion photographs in a fully automatic or semi-automatic manner, it can be used for increasing efficiency in many applications, such as automatic walkthroughs of photograph galleries, motion photographs on website, electronic greeting cards, and personalized Karaoke
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2005 Unsupervised auditory scene categorization via key audio effects and information-theoretic co-clustering
abstract
Automatic categorization of auditory scenes is very useful in various content-based multimedia applications, such as video indexing and context-aware computing. An unsupervised approach is proposed to group auditory scenes with similar semantics. In our approach, auditory scenes are described by the key audio effects they contain. In order to exploit the relationships between different audio effects and provide a more accurate similarity measure for auditory scene categorization, co-clustering is used to group auditory scenes and key audio effects simultaneously. In addition, a Bayesian information criterion (BIC) is used to select cluster numbers automatically for both the key effects and the auditory scenes. Evaluation on 272 auditory scenes extracted from 12-hour audio data shows very encouraging results.
Rui Cai 0002, Lie Lu, Lianhong Cai
ICASSP (2)2
2005 Towards a unified framework for content-based audio analysis
abstract
Audio content analysis is helpful in many multimedia applications. We present a unified framework for content analysis of composite audio. The framework is designed to extract relevant information from different available audio modalities and to discover high-level semantics conveyed by the data. We also demonstrate an implementation of the proposed framework for the detection of scenes and events in various TV shows and movies, in which key audio effects are first extracted as a midlevel representation, and then a Bayesian network is used for high-level semantics inference. Experiments on 12-hour audio data indicate that the proposed framework has a satisfying performance.
Lie Lu, Rui Cai 0002, Alan Hanjalic
ICASSP (2)1
2005 Robust learning-based TV commercial detection
abstract
A robust learning-based TV commercial detection approach is proposed in this paper. Firstly, a set of basic features that facilitate distinguishing commercials from general program are analyzed. Then, a series of context-based features, which are more effective for identifying commercials, are derived from these basic features. Next, each shot is classified as commercial or general program based on these features by a pre-trained SVM classifier. And last, the detection results are further refined by scene grouping and some heuristic rules. Experiments on around 10-hour TV recordings of various genres show that the proposed scheme is able to identify commercial blocks with relatively high detection accuracy.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ICME2
2005 Perceptual Visualization of a Music Collection
abstract
Music visualization provides users with a new interface to browse, search, and navigate their personal digital music collection. Although there are several previous works on visualizing music collection based on some "surface" musical metadata, such as artist, album and genre, there has been few works on content-based or perception-based visualizations. In this paper, we developed an algorithm to automatically estimate human perceptions on rhythm and timbre of a music clip. Then, based on these two values, each music clip is mapped into a 2D (timbre-rhythm) space. Thus, a 2D perception-based visualization is built. Experimental evaluation indicates that this kind visualization is efficiently helpful in many cases of music management manipulations, such as music navigation, similar music search and music play list generation
Lie Lu
ICME2
2005 Web object indexing using domain knowledge
abstract
A web object is defined to represent any meaningful object embedded in web pages (e.g. images, music) or pointed to by hyperlinks (e.g. downloadable files). In many cases, users would like to search for information of a certain 'object', rather than a web page containing the query terms. To facilitate web object searching and organizing, in this paper, we propose a novel approach to web object indexing, by discovering its inherent structure information with existed domain knowledge. In our approach, first, Layered LSI spaces are built for a better representation of the hierarchically structured domain knowledge, in order to emphasize the specific semantics and term space in each layer of the domain knowledge. Meanwhile, the web object representation is constructed by hyperlink analysis, and further pruned to remove the noises. Then an optimal matching between the web object and the domain knowledge is performed, in order to pick out the structure attributes of the web object from the knowledge. Finally, the obtained structure attributes are used to re-organize and index the web objects. Our approach also indicates a new promising way to use trust-worthy Deep Web knowledge to help organize dispersive information of Surface Web.
Muyuan Wang, Zhiwei Li 0006, Lie Lu, Wei-Ying Ma, Naiyao Zhang
KDD3
2005 Unsupervised content discovery in composite audio
abstract
Automatically extracting semantic content from audio streams can be helpful in many multimedia applications. Motivated by the known limitations of traditional supervised approaches to content extraction, which are hard to generalize and require suitable training data, we propose in this paper an unsupervised approach to discover and categorize semantic content in a composite audio stream. In our approach, we first employ spectral clustering to discover natural semantic sound clusters in the analyzed data stream (e.g. speech, music, noise, applause, speech mixed with music, etc.). These clusters are referred to as audio elements. Based on the obtained set of audio elements, the key audio elements, which are most prominent in characterizing the content of input audio data, are selected and used to detect potential boundaries of semantic audio segments denoted as auditory scenes. Finally, the auditory scenes are categorized in terms of the audio elements appearing therein. Categorization is inferred from the relations between audio elements and auditory scenes by using the information-theoretic co-clustering scheme. Evaluations of the proposed approach performed on 4 hours of diverse audio data indicate that promising results can be achieved, both regarding audio element discovery and auditory scene categorization.
Rui Cai 0002, Lie Lu, Alan Hanjalic
ACM Multimedia2
2005 Automated rich presentation of a semantic topic
abstract
To have a rich presentation of a topic, it is not only expected that many relevant multimodal information, including images, text, audio and video, could be extracted; it is also important to organize and summarize the related information, and provide users a concise and informative storyboard about the target topic. It facilitates users to quickly grasp and better understand the content of a topic. In this paper, we present a novel approach to automatically generating a rich presentation of a given semantic topic. In our proposed approach, the related multimodal information of a given topic is first extracted from available multimedia databases or websites. Since each topic usually contains multiple events, a text-based event clustering algorithm is then performed with a generative model. Other media information, such as the representative images, possibly available video clips and flashes (interactive animates), are associated with each related event. A storyboard of the target topic is thus generated by integrating each event and its corresponding multimodal information. Finally, to make the storyboard more expressive and attractive, an incidental music is chosen as background and is aligned with the storyboard. A user study indicates that the presented system works quite well on our testing examples.
Lie Lu
ACM Multimedia1
2005 Unsupervised speaker segmentation and tracking in real-time audio content analysis
Lie Lu, HongJiang Zhang
Multim. Syst.1
2005 A generic framework of user attention model and its application in video summarization
abstract
Due to the information redundancy of video, automatically extracting essential video content is one of key techniques for accessing and managing large video library. In this paper, we present a generic framework of a user attention model, which estimates the attentions viewers may pay to video contents. As human attention is an effective and efficient mechanism for information prioritizing and filtering, user attention model provides an effective approach to video indexing based on importance ranking. In particular, we define viewer attention through multiple sensory perceptions, i.e. visual and aural stimulus as well as partly semantic understanding. Also, a set of modeling methods for visual and aural attentions are proposed. As one of important applications of user attention model, a feasible solution of video summarization, without fully semantic understanding of video content as well as complex heuristic rules, is implemented to demonstrate the effectiveness, robustness, and generality of the user attention model. The promising results from the user study on video summarization indicate that the user attention model is an alternative way to video understanding.
Yufei Ma 0006, Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
IEEE Trans. Multim.3
2004 Improve audio representation by using feature structure patterns
abstract
Although statistical characteristics of audio features are widely used for audio representation in most current audio analysis systems and have been proved to be effective, they only utilize the average feature variations over time, and thus lead to ambiguities in some cases. Structure patterns, which describe the representative structure characteristics of both temporal and spectral features, are proposed to improve audio representation. In this paper, three structure patterns, including energy envelope pattern, sub-band spectral shape pattern and harmonicity prominence pattern, are proposed or refined, as successive development of our previous work. Evaluations on a content-based audio retrieval system with more than 1500 clips showed very encouraging results.
Rui Cai 0002, Lie Lu, HongJiang Zhang, Lianhong Cai
ICASSP (4)2
2004 Repeating pattern discovery from acoustic musical signals
abstract
Music pieces are typically repetitive. The automatic extraction of repeating patterns is useful for music summary, indexing and retrieval. An effective approach for repeating pattern discovery is proposed. In order to represent the melody similarity more accurately, a constant Q transform is used for feature extraction and a novel similarity measure between musical features is proposed. From the self-similarity matrix of the music, an adaptive method is used to extract all the significant repeating patterns. Experiments on pop music indicate the approach is promising.
Muyuan Wang, Lie Lu, HongJiang Zhang
ICME2
2004 P-Karaoke: personalized karaoke system
abstract
In this demonstration, a personalized Karaoke system, P-Karaoke, is proposed. In the P-Karaoke system, personal home videos and photographs, which are automatically selected from users' multimedia database according to their content, users' preferences or music, are utilized as the background videos of the Karaoke. The selected video clips, photographs, music and lyrics are well aligned to compose a Karaoke video, connecting by specific content-based transitions.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ACM Multimedia2
2004 Automatic music video generation based on temporal pattern analysis
abstract
Music video (MV) is a short film meant to present a visual representation of a popular music song. In this paper, we present a system that automatically generates MV-like videos from personal home videos based on observations that generally there are obvious repetitive visual and aural patterns in MVs. Based on a set of video and music analysis algorithms, the automatic music video (AMV) generation system automatically extracts temporal structures of the video and music, as well as repetitive patterns in the music. And then, according to the structure and patterns, a set of highlight segments from the raw home video footage are selected, aiming at matching the visual content with the aural structure and pattern. And last, the output music video is rendered by connecting the selected highlight video segments with appropriate transition effects, accompanied with the music. Experiments show that the results are compelling and promising.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ACM Multimedia2
2004 Automatically converting otograic series into video
abstract
In this paper, we proposed a novel way to browse a series of otogras, which can be regarded as a system exploring the new medium between otogra and video. The scheme exploits the rich content embedded in a single otogra and otograic series. Based on studying the process of a viewer's attention variation on objects or regions of an image, a otogra can be converted into a motion clip. A system named oto2Video was developed to automatically convert a otograic series into a video by simulating camera motions, set to incidental music of the user's choice. For a selected otograic series, an appropriate set of key-frames are determined for each otogra based on soisticated content analytical results. Then camera motion pattern (both the key-frame sequencing scheme and trajectory/speed control strategy) is selected for each otogra to generate a corresponding motion otogra clip. And last, the final output video is rendered by connecting a series of motion otogra clips with specific transitions based on the content of the images on either side, as well as each motion otogra clip is aligned with the selected incidental music based on music content analysis.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ACM Multimedia2
2004 Audio textures: theory and applications
abstract
In this paper, we introduce a new audio medium, called audio texture, as a means of synthesizing long audio stream according to a given short example audio clip. The example clip is first analyzed to extract its basic building patterns. An audio stream of arbitrary length is then synthesized using a sequence of extracted building patterns. The patterns can be varied in the synthesis process to add variations to the generated sound to avoid simple repetition. Audio textures are useful in applications such as background music, lullabies, game music, and screen saver sounds. We also extend this idea to audio texture restoration, or constrained audio texture synthesis for restoring the missing part in an audio clip. It is also useful in many applications such as error concealment for audio/music delivery with packets loss on the Internet. Novel methods are proposed for unconstrained and constrained audio texture synthesis. Preliminary results are provided for evaluation.
Lie Lu, Wenyin Liu, HongJiang Zhang
IEEE Trans. Speech Audio Process.1
2004 Optimization-based automated home video editing system
abstract
In this paper, we present an optimization-based system that automates home video editing. This system automatically selects suitable or desirable highlight segments from a set of raw home videos and aligns them with a given piece of incidental music to create an edited video segment to a desired length based on the content of the video and incidental music. We developed an approach for extracting temporal structure and determining the importance of a video segment in order to facilitate the selection of highlight segments. Additionally we extract a temporal structure, beats, and tempos from the incidental music. In order to create more professional-looking results, the selected highlight segments satisfy a set of editing rules and are matched to the content of the incidental music. This task is formulated as a nonlinear 0-1 programming problem and the rules, which are adjustable and increasable, are embedded as constraints. The output video is rendered by connecting the selected highlight video segments with transition effects and the incidental music. Under this framework, we can choose the best-matched music for a given video and support different output styles.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2003 Audio restoration by constrained audio texture synthesis
abstract
Audio texture, a new audio medium, is used to synthesize long audio streams according to a given short example audio clip. In this paper, we extend this idea to audio texture restoration, or constrained audio texture synthesis for restoring those missing parts in an audio clip. It is useful in many applications such as audio restoration and audio reconstruction. It can also be used in error concealment for audio/music delivery with packet loss on the Internet. A novel method is proposed for constrained audio texture synthesis. Preliminary results are provided for evaluation.
Lie Lu, Wenyin Liu, HongJiang Zhang
ICASSP (5)1
2003 Speech segmentation without speech recognition
abstract
In this paper, we presented a semantic speech segmentation approach, in particular sentence segmentation, without speech recognition. In order to get phoneme level information without word recognition information, a novel vowel/consonant/pause (V/C/P) classification is proposed. An adaptive pause detection method is also presented to adapt to various backgrounds and environments. Three feature sets, which include pause, rate of speech and prosody, are used to discriminate the sentence boundary. Experiments on broadcasting news indicate that the performance of the proposed algorithm is satisfying.
Lie Lu, HongJiang Zhang
ICASSP (1)2
2003 UBM-based real-time speaker segmentation for broadcasting news
abstract
This paper addresses the problem of real-time speaker change detection in broadcast news, in which no prior knowledge on speakers is assumed. Our speaker segmentation is a "coarse to refine" process, which consists of two stages: pre-segmentation and refinement. In the pre-segmentation stage, a new approach based on Gaussian mixture model-universal background model (GMM-UBM) is proposed to categorize feature vectors into three sets, i.e. reliable speaker-related set, doubtful speaker-related set and unreliable speaker-related set, in order to enhance the effect of the reliable speaker-related feature vectors. Then potential speaker change boundaries are detected based on a novel distance measure. In the refinement stage, incremental speaker adaptation (ISA), which is suitable for real-time requirement, is proposed to obtain considerably precise speaker models so that the potential speaker change boundaries can be confirmed and refined. Experimental results demonstrate that our approach yields satisfactory performance.
Ting-Yao Wu, Lie Lu, Ke Chen 0001, HongJiang Zhang
ICASSP (2)2
2003 Highlight sound effects detection in audio stream
abstract
This paper addresses the problem of highlight sound effects detection in audio stream, which is very useful in fields of video summarization and highlight extraction. Unlike researches on audio segmentation and classification, in this domain, it just locates those highlight sound effects in audio stream. An extensible framework is proposed and in current system three sound effects are considered: laughter, applause and cheer, which are tied up with highlight events in entertainments, sports, meetings and home videos. HMMs are used to model these sound effects and a log-likelihood scores based method is used to make final decision. A sound effect attention model is also proposed to extend general audio attention model for highlight extraction and video summarization. Evaluations on a 2-hours audio database showed very encouraging results.
Rui Cai 0002, Lie Lu, HongJiang Zhang, Lianhong Cai
ICME2
2003 Audio restoration by constrained audio texture synthesis
abstract
Audio texture, a new audio medium, is used to synthesize long audio stream according to a given short example audio clip. In this paper, we extend this idea to audio texture restoration, or constrained audio texture synthesis for restoring those missing parts in an audio clip. It is useful in many applications such as audio restoration and audio reconstruction. It can also be used in error concealment for audio/music delivery with packet loss on the Internet. A novel method is proposed for constrained audio texture synthesis. Preliminary results are provided for evaluation.
Lie Lu, Wenyin Liu, HongJiang Zhang
ICME1
2003 Speech segmentation without speech recognition
abstract
In this paper, we presented a semantic speech segmentation approach, in particular sentence segmentation, without speech recognition. In order to get phoneme level information without word recognition information, a novel vowel/consonant/pause (V/C/P) classification is proposed. An adaptive pause detection method is also presented to adapt to various background and environment. Three feature sets, which include pause, rate of speech and prosody, are used to discriminate the sentence boundary. Experiments on broadcasting news indicate that the performance of proposed algorithm is satisfying.
Lie Lu, HongJiang Zhang
ICME2
2003 UBM-based incremental speaker adaptation
abstract
This paper addresses a novel algorithm of incremental speaker adaptation (ISA) based on universal background model (UBM) for saving storage and real-time processing. This algorithm can be seen as an extension of traditional speaker adaptation. It consists of two steps, adaptation and combination. It not only considers the speaker's characteristics in limited training data, but also prohibits over-fitting of the updated model. The incremental adaptation algorithm needs little storage and meets the requirement of real-time processing. In order to evaluate the efficiency and effectivity of the proposed approach, a real-time speaker segmentation system for broadcasting news is built. Experiment results demonstrate that our approach yields real time operation and achieves satisfactory performance.
Ting-Yao Wu, Lie Lu, Ke Chen 0001, HongJiang Zhang
ICME2
2003 Using structure patterns of temporal and spectral feature in audio similarity measure
abstract
Although statistical characteristics of audio features are widely used for similarity measure in most of current audio analysis systems and have been proved to be effective, they only utilized the averaged feature variations over time, and thus lead to inaccuracy in some cases. In this paper, structure pattern, which describes the representative structure characteristics of both temporal and spectral features, is proposed to improve the similarity measure for audio effects. Three kind structure patterns are proposed and utilized in current work, including energy contour pattern, harmonicity pattern and pitch contour pattern. Evaluations on a content-based audio retrieval system indicate that structure patterns can improve the performance pretty much.
Rui Cai 0002, Lie Lu, HongJiang Zhang
ACM Multimedia2
2003 AVE: automated home video editing
abstract
In this paper, we present a system that automates home video editing. This system automatically extracts a set of highlight segments from a set of raw home videos and aligns them with user supplied incidental music based on the content of the video and incidental music. We developed an approach for extracting temporal structure and determining the importance of a video segment in order to facilitate the selection of highlight segments. Additionally we extract temporal structure, beats and tempos from the incidental music. In order to create more professional-looking results, the selected highlight segments satisfy a set of editing rules and are matched to the content of the incidental music. This task is formulated as a non-linear 0-1 programming problem and the rules are embedded as constraints. The output video is rendered by connecting the selected highlight video segments with transition effects and the incidental music.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ACM Multimedia2
2003 Photo2Video
abstract
To exploit rich content embedded in a single photograph, a system named Photo2Video was developed to automatically convert a photographic series into a video by simulating camera motions, set to incidental music of the user's choice. For a chosen photographic series, an appropriate camera motion pattern is selected for each photograph to generate a corresponding motion photograph clip. Then, the final output video is rendered by connecting a series of motion photograph clips with specific transitions, and aligning with the selected incidental music. Photo2Video provides a novel way to browse a series of images and can be regarded as a system exploring the new medium between photograph and video.
Xian-Sheng Hua 0001, Lie Lu, HongJiang Zhang
ACM Multimedia2
2003 Automated extraction of music snippets
abstract
Similar to image and video thumbnail, music snippet is defined as the most representative or highlight excerpt of a music clip, and can be used efficiently for fast browsing large number of music files. Music snippet is usually a part of the repeated melody, main theme or chorus. In this paper, we present an approach to extracting music snippet automatically. In our approach, the most salient segment of the music is firstly detected based on its occurrence frequency and energy information. Meanwhile, the boundaries of musical phrases are also detected based on the estimated phrase length and phrase boundary confidence of each frame. These boundaries are used to ensure that an extracted snippet does not break musical phrases. Finally, the musical phrases including the most salient segment are extracted as music snippet. User study indicates that the proposed algorithm works very well on our music database.
Lie Lu, HongJiang Zhang
ACM Multimedia1
2003 Universal Background Models for Real-time Speaker Change Detection
Ting-Yao Wu, Lie Lu, Ke Chen 0001, HongJiang Zhang
MMM2
2003 Content-based audio classification and segmentation by using support vector machines
Lie Lu, HongJiang Zhang, Stan Z. Li
Multim. Syst.1
2002 Audio textures
abstract
In this paper, we introduce a new audio medium, called audio texture, as a means of synthesizing long audio stream according to a given short example audio clip. The example clip is analyzed, and basic building patterns are extracted. Then an audio stream of arbitrary length is synthesized using a sequence of extracted building patterns. The patterns can be varied in the synthesis process to add variations to the generated sound. Audio textures are useful in applications such as background music, lullabies, game music, and screen saver sounds. A method is proposed for implementing audio textures. Preliminary results of audio textures are provided at our website for evaluation.
Lie Lu, Stan Z. Li, Wenyin Liu, HongJiang Zhang
ICASSP1
2002 Music type classification by spectral contrast feature
abstract
Automatic music type classification is very helpful for the management of digital music databases. In this paper, the octave-based spectral contrast feature is proposed to represent the spectral characteristics of a music clip. It represented the relative spectral distribution instead of average spectral envelope. Experiments show that the octave-based spectral contrast feature performs well in music type classification. Another comparison experiment demonstrates that the octave-based spectral contrast feature has a better discrimination among different music types than mel-frequency cepstral coefficients (MFCC), which is often used in previous music type classification systems.
Dan-Ning Jiang, Lie Lu, HongJiang Zhang, Jianhua Tao 0001, Lianhong Cai
ICME (1)2
2002 Speaker change detection and tracking in real-time news broadcasting analysis
abstract
This paper addresses the problem of real time speaker change detection and speaker tracking in broadcasted news video analysis. In such a case, both speaker identities and number of speakers are assumed unknown. A two-step speaker change detection algorithm, including potential change detection and refinement, is proposed. Speaker tracking is performed based on the results of speaker change detection. A Bayesian Fusion method is used to fuse multiple audio features to get a more reliable result. The algorithm has low complexity and runs in real-time with a very limited delay in analysis. Our experiments show that the algorithms produce very satisfactory results.
Lie Lu, HongJiang Zhang
ACM Multimedia1
2002 A user attention model for video summarization
abstract
Automatic generation of video summarization is one of the key techniques in video management and browsing. In this paper, we present a generic framework of video summarization based on the modeling of viewer's attention. Without fully semantic understanding of video content, this framework takes advantage of computational attention models and eliminates the needs of complex heuristic rules in video summarization. A set of methods of audio-visual attention model features are proposed and presented. The experimental evaluations indicate that the computational attention based approach is an effective alternative to video semantic analysis for video summarization.
Yufei Ma 0006, Lie Lu, HongJiang Zhang, Mingjing Li
ACM Multimedia2
2002 Content analysis for audio classification and segmentation
abstract
We present our study of audio content analysis for classification and segmentation, in which an audio stream is segmented according to audio type or speaker identity. We propose a robust approach that is capable of classifying and segmenting an audio stream into speech, music, environment sound, and silence. Audio classification is processed in two steps, which makes it suitable for different applications. The first step of the classification is speech and nonspeech discrimination. In this step, a novel algorithm based on K-nearest-neighbor (KNN) and linear spectral pairs-vector quantization (LSP-VQ) is developed. The second step further divides nonspeech class into music, environment sounds, and silence with a rule-based classification scheme. A set of new features such as the noise frame ratio and band periodicity are introduced and discussed in detail. We also develop an unsupervised speaker segmentation algorithm using a novel scheme based on quasi-GMM and LSP correlation analysis. Without a priori knowledge, this algorithm can support the open-set speaker, online speaker modeling and real time segmentation. Experimental results indicate that the proposed algorithms can produce very satisfactory results.
Lie Lu, HongJiang Zhang, Hao Jiang 0007
IEEE Trans. Speech Audio Process.1
2001 Content-Based Audio Segmentation Content-Based Audio Segmentation
Lie Lu, Stan Z. Li, HongJiang Zhang
ICME1
2001 A Newapproach To Query By Humming In Music Retrieval
abstract
In this paper, we present a method for querying desired songs from music database by humming a tune. Since errors are inevitable in humming, tolerance should be considered. In order to suit or adapt to people's humming habit, a new melody representation and new hierarchical matching method are proposed in this paper. Query processing and database processing algorithm are also described in detail. The proposed approach achieves high performance in experimental evaluation. For 88% queries, the correct song can be retrieved among the top 10 matches.
Lie Lu, Hong You, HongJiang Zhang
ICME1
2001 A robust audio classification and segmentation method
abstract
In this paper, we present a robust algorithm for audio classification that is capable of segmenting and classifying an audio stream into speech, music, environment sound and silence. Audio classification is processed in two steps, which makes it suitable for different applications. The first step of the classification is speech and non-speech discrimination. In this step, a novel algorithm based on KNN and LSP VQ is presented. The second step further divides non-speech class into music, environment sounds and silence with a rule based classification scheme. Some new features such as the noise frame ratio and band periodicity are introduced and discussed in detail. Our experiments in the context of video structure parsing have shown the algorithms produce very satisfactory results.
Lie Lu, Hao Jiang 0007, HongJiang Zhang
ACM Multimedia1