Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ying Li 0121

dblp:22/1805-121 · DBLP profile ↗
← Back
32ranked-venue papers
25as first author
0since 2021 · last 2016
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 20 first-authorArtificial intelligence and machine learning · 6 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 71% Audio and music processing · 29%
Artificial intelligence
1 paper
Face, body and person analysis · 77% Trustworthy machine learning · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 77% Medical and health informatics · 23%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › video content analysis
instructional video analysis
0.132006
Instructional Video Content Analysis Using Audio Information · IEEE Trans. Speech Audio Process. 2006
Analyzing discussion scene contents in instructional videos · ACM Multimedia 2004
Atomic topical segments detection for instructional videos · ACM Multimedia 2006
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.112010
Robust Symbolic Dual-View Facial Expression Recognition With Skin Wrinkles: Local Versus Global Approach · IEEE Trans. Multim. 2010
Audio and music processing › audio analysis
audio content analysis
0.112006
Instructional Video Content Analysis Using Audio Information · IEEE Trans. Speech Audio Process. 2006
Multimedia analysis and retrieval › video content analysis
topic segmentation
0.112006
Atomic topical segments detection for instructional videos · ACM Multimedia 2006
Multimedia analysis and retrieval › video content analysis
video structure analysis
0.112006
Atomic topical segments detection for instructional videos · ACM Multimedia 2006
Audio and music processing › speaker diarization
speaker clustering
0.012004
Analyzing discussion scene contents in instructional videos · ACM Multimedia 2004
Multimedia analysis and retrieval
video content analysis
0.012004
Analyzing discussion scene contents in instructional videos · ACM Multimedia 2004
Machine learning › Trustworthy machine learning
robustness
0.012010
Robust Symbolic Dual-View Facial Expression Recognition With Skin Wrinkles: Local Versus Global Approach · IEEE Trans. Multim. 2010
Distributed systems › publish/subscribe systems
publish/subscribe middleware
0.012007
SMILE: distributed middleware for event stream processing · IPSN 2007
Medical and health informatics
distributed learning
0.012006
MAGICAL demonstration: system for automated metadata generation for instructional content · ACM Multimedia 2006
Audio and music processing
audio classification
0.012006
Instructional Video Content Analysis Using Audio Information · IEEE Trans. Speech Audio Process. 2006
Multimedia analysis and retrieval › audio-visual learning
audio-visual video analysis
0.012005
Creating MAGIC: system for generating learning object metadata for instructional content · ACM Multimedia 2005

Methods — techniques the papers use, named apart from their topics

continuous query · 0.1mode-based clustering · 0.1symbolic dual-view representation · 0.1local and global feature comparison · 0.1dataflow network · 0.1data flow network · 0.1video analysis · 0.1text analysis · 0.1metadata extraction · 0.1gaussian mixture model · 0.1document analysis · 0.1audiovisual cue fusion · 0.1text categorization · 0.1audiovisual analysis · 0.1
YearPublicationVenuePosition
2016 A Case Study of Mobile User Behaviors Using Spatio-temporal Data
abstract
Increasing use of mobile apps which capture location information has led to wide availability of spatio-temporal data. This paper details our recent efforts on using such data to understand mobile user behaviors in terms of their interaction with apps. Specifically, we aim to mine the association between users' app open (AO) behaviors and their waiting times associated with some transport modes. Here, the transport mode is derived based on speed information measured from a user's location update data, without using any additional map data. One particular case study that we conducted and report here is to understand if users tend to access the app more often while they are waiting at airports. Using a two-week period of a particular iPhone app data from a major U.S. Retailer, the study shows that the app open rate (AR) of air travelers measured during their airport-dwelling time is 8x higher than their AR at other locations. Moreover, for the same group of travelers who have AOs at both airports and other locations, their AR at airports is 45x higher than that at other locations and times. Findings drawn from this study can be applied to assist the definition of geofences by retailers to improve targeting of customers at the right locations, and consequently improve the success of marketing campaigns.
Ying Li 0121, Wesley M. Gifford, Anshul Sheopuri
MDM1
2015 Applying image analysis to assess food aesthetics and uniqueness
abstract
This paper describes our latest work on assessing the aesthetics of food dishes by analyzing its visual appearance. Specifically, our framework builds upon work in the areas of color science, psychology and statistics to enable users to computationally assess whether the image of a dish is visually appealing. Moreover, we further assess the uniqueness of a dish by comparing its color composition against a dish image repository. The assessment engine has been recently integrated into a Google Glass Application which brings a new dimension of wearable computing to this work. Our preliminary experiments with the assessment system have shown encouraging results, and the initial deployment of the Google Glass Application has raised quite some interests.
Ying Li 0121, Anshul Sheopuri
ICIP1
2015 Creative design of color palettes for product packaging
abstract
This paper describes our latest work on assisting CPG (Consumer Packaged Goods) companies with their product packaging designs by providing color palettes that are visually appealing, novel and consistent with desired marketing messages for a particular brand and product. Specifically, we start by mining a large collections of images of different products and brands to learn about all the colors and color combinations that frequently appear among them. Meanwhile, a color-message graph is constructed to represent messages conveyed by different colors as well as to capture the interrelationship among them. Knowledge from both color psychology and information sources like Thesaurus are extensively exploited in this case. Now, given a particular product and brand to be designed for its packaging, along with the company's desired marketing message, we apply a computational method to generate quintillions of novel color palettes that can be used for the design. This process will leverage existing palettes used by same products of different brands or different products of the same brand, take in optional color preferences from users, identify then utilize the right colors to convey the desired marketing message. Finally, we rank the palettes based on assessment of their visual aesthetics, novelty and the way that different messages of the same palette interact with each other, so as to guide human designers to choose the right ones. Our initial demonstrations of this work to colleagues of subject matter have received very positive feedback. We are now exploring opportunities to collaborate with them to validate this technology in a controlled experimental setting.
Ying Li 0121, Anshul Sheopuri
ICME1
2014 Exploring Application Domains for Computational Creativity
Ashish Jagmohan, Ying Li 0121, Anshul Sheopuri, Dashun Wang, Lav R. Varshney
ICCC2
2014 Rail Component Detection, Optimization, and Assessment for Automatic Rail Track Inspection
abstract
In this paper, we present a real-time automatic vision-based rail inspection system, which performs inspections at 16 km/h with a frame rate of 20 fps. The system robustly detects important rail components such as ties, tie plates, and anchors, with high accuracy and efficiency. To achieve this goal, we first develop a set of image and video analytics and then propose a novel global optimization framework to combine evidence from multiple cameras, Global Positioning System, and distance measurement instrument to further improve the detection performance. Moreover, as the anchor is an important type of rail fastener, we have thus advanced the effort to detect anchor exceptions, which includes assessing the anchor conditions at the tie level and identifying anchor pattern exceptions at the compliance level. Quantitative analysis performed on a large video data set captured with different track and lighting conditions, as well as on a real-time field test, has demonstrated very encouraging performance on both rail component detection and anchor exception detection. Specifically, an average of 94.67% precision and 93% recall rate has been achieved for detecting all three rail components, and a 100% detection rate is achieved for compliance-level anchor exception with three false positives per hour. To our best knowledge, our system is the first to address and solve both component and exception detection problems in this rail inspection area.
Ying Li 0121, Hoang Trinh, Norman Haas, Charles Otto, Sharath Pankanti
IEEE Trans. Intell. Transp. Syst.1
2012 Anomalous tie plate detection for railroad inspection
Ying Li 0121, Sharath Pankanti
ICPR1
2012 Enhanced rail component detection and consolidation for rail track inspection
abstract
For safety purposes, railroad tracks need to be inspected on a regular basis for physical defects or design noncompliances. Such track defects and non-compliances, if not detected in a timely manner, may eventually lead to grave consequences such as train derailments. In this paper, we present a real-time automatic vision-based rail inspection system, with main focus on anchors - an important rail component type, and anchor-related rail defects, or exceptions. Our system robustly detects important rail components including ties, tie plates, anchors with high accuracy and efficiency. Detected objects are then consolidated across video frames and across camera views to map to physical rail objects, by combining the video data streams from all camera views with GPS information and speed information from the distance measuring instrument (DMI). After these rail components are detected and consolidated, further data integration and analysis is followed to detect sequence-level track defects, or exceptions. Quantitative analysis performed on a real online field test conducted on different track conditions demonstrates that our system achieves very promising performance in terms of rail component detection, anchor condition assessment, and compliance-level exception detection. We also show that our system outperforms another advanced rail inspection system in anchor detection.
Hoang Trinh, Norman Haas, Ying Li 0121, Charles Otto, Sharath Pankanti
WACV3
2011 Intelligent headlight control using learning-based approaches
abstract
This paper describes our recent work on developing an intelligent headlight control system using machine learning-based approaches. Specifically, such a system aims to automatically control a vehicle's beam state (high beam or low beam) during a night-time drive based on the detection of oncoming/overtaking/leading traffics as well as urban areas from the videos captured by a camera. Two machine learning-based approaches, namely, support vector machine (SVM) and AdaBoost, have been applied to accomplish this task. The architect of each approach, as well as its detailed processing modules, will be elaborated in the paper. The system has been extensively tested both online and offline to validate the robustness and effectiveness of the two proposed approaches. A detailed performance study along with some comparisons between the two approaches will be reported at the end.
Ying Li 0121, Norman Haas, Sharath Pankanti
Intelligent Vehicles Symposium1
2011 Component-based track inspection using machine-vision technology
abstract
In this paper, we present our latest research engagement with a railroad company to apply machine vision technologies to automate the inspection and condition monitoring of railroad tracks. Specifically, we have proposed a complete architecture including imaging setup for capturing multiple video streams, important rail component detection such as tie plate, spike, anchor and joint bar bolt, defect identification such as raised spikes, defect severity analysis and temporal condition analysis, and long-term predictive assessment. This paper will particularly present various video analytics that we have developed to detect rail components, which form the building block of the entire framework. Our preliminary performance study has achieved an average of 98.2% detection rate, 1.57% false positive rate and 1.78% false negative rate on the component detection. Finally, with the lack of sufficient representative data and annotations to evaluate system performance on exception detection at both sequence and compliance levels, we proposed a mathematical modeling approach to calculate the probabilities of detecting such exceptions. Such analysis shows that there is still big room for us to improve our approaches in order to achieve desired false positive rate and miss detection rate at the sequence level.
Ying Li 0121, Charles Otto, Norman Haas, Yuichi Fujiki, Sharath Pankanti
ICMR1
2011 A performance study of an intelligent headlight control system
abstract
In this paper, we first present the architecture of an intelligent headlight control (IHC) system that we developed in our earlier work. This IHC system aims to automatically control a vehicle's beam state (high beam or low beam) during a night-time drive. A three-level decision framework built around a support vector machine (SVM) learning engine is then briefly discussed. Next, we switch our focus to the study of system performance by varying the SVM feature set, as well as by exploiting various SVM training options and adjustments through a set of experiments. We believe that what we learned from this performance study can provide readers useful guidelines on extracting effective SVM features within the IHC problem domain, as well as on training an effective SVM learning engine for more generalized applications.
Ying Li 0121, Sharath Pankanti
WACV1
2010 Real-Time Traffic Sign Detection: An Evaluation Study
abstract
This paper presents an experimental evaluation of three different traffic sign detection approaches, which detect or localize various types of traffic signs from real-time videos. Specifically, the first approach exploits geometric features to identify traffic signs, while the other two are developed based on SVM (Support Vector Machine) and AdaBoost learning mechanisms. We describe each of the three approaches, conduct a detailed comparison among them, and examine their pros and cons. Our conclusions should lead to useful guidelines for developing a real-time traffic sign detector.
Ying Li 0121, Sharath Pankanti, Weiguang Guan
ICPR1
2010 Robust Symbolic Dual-View Facial Expression Recognition With Skin Wrinkles: Local Versus Global Approach
abstract
Simple cartoon facial expressions can be represented by emoticons, that is, a special sequence of symbols. This inspires us that a sketch of facial feature contour may be adequate to recognize expressions. Metrics of such sketches are easier to be calibrated under varying illumination and head pose. While skin wrinkles such as nasolabial folds, eye pouches, dimples, forehead, and chin furrows are not salient facial features, they may convey crucial subtle signals about an individual's emotion. Our experiments have shown that the side-view profile plus skin wrinkles can correctly differentiate nearly 70% expressions, and it contributes to the increase of overall recognition rate. Finally, we compare the accuracy and robustness of various local and global processing schemes, especially under the condition of partial occlusion.
Yizhen Huang, Ying Li 0121, Na Fan 0001
IEEE Trans. Multim.2
2008 Automatically constructing blue pages for characters in instructional videos
abstract
This paper presents our recent work on automatically constructing blue pages for main characters in instructional videos. Specifically, a blue page is a personal profile which contains various types of information such as name, affiliation, portrait and voice that are specific to each individual. To accomplish this, we first extract various types of personal identity information from the video by analyzing multiple media cues; then we carefully correlate them with each other w.r.t. individual video characters based on an advanced video context analysis. When additional information sources are available, the constructed blue pages could be further enriched with advanced data mining and information extraction techniques. To validate the proposed ideas, we have carried out some preliminary experiments on a set of instructional videos with acceptable results obtained.
Ying Li 0121, Youngja Park
ICME1
2007 SMILE: distributed middleware for event stream processing
abstract
In this paper, we describe the SMILE (Smart MIddleware, Light Ends) system which is one of the earliest systems built in the area of distributed event stream processing. SMILE unites the "publish-subscribe" model of messaging middleware with the "continuous query" model of database systems. In SMILE, information producers, which may be sensors, applications, or databases, generate streams of events, such as RFID data, news items or stock trades; consumers specify stateful subscriptions to derived views, such as "individuals trading top 5 total volume of stock Y within x minutes before a major news story about the same company Y"; the SMILE system constructs and deploys a dataflow network of computations which process events from producers, compute derived views and deliver continuous and timely updates of subscribed views to consumers. Research challenges addressed by SMILE include: rigorously defining the semantics for correct state delivery; implementing this semantics in the presence of failure; distributing computations optimally over a network of multiple servers; scheduling and controlling message flow in this network; smoothly integrating this technology with other components, e.g., databases, services, user interfaces. SMILE is applicable to any "sense and respond" scenario that requires monitoring and processing of events as they are generated. Applications span multiple industries, e.g., monitoring financial opportunities and detecting potential fraud in financial industry, systems management and alerts in data centers, inventory management in RFID applications, and business performance management. In this paper, we describe the SMILE system, its architecture, services supported, and implementation details, and its use in multiple application scenarios.
Robert E. Strom, Chitra Dorai, Gerry Buttner, Ying Li 0121
IPSN4
2007 Applying Image Analysis to Auto Insurance Triage: A Novel Application
abstract
For the auto insurance claims process, improvements in the First Notice of Loss and rapidity in the investigation and evaluation of claims could drive significant values by reducing loss adjustment expense. This paper proposes a novel application where advanced technologies in image analysis and pattern recognition are applied to automatically identify and characterize automobile damage. Success in this will allow some cases to proceed without human adjusters, while others to proceed more efficiently, thus ultimately shortening the time between the first Notice of Loss and the final payout. To investigate its feasibility, we built a prototype system which automatically identifies the damaged area(s) based on the comparison of before-and after-accident automobile images. Performance of the prototype system has been evaluated on images taken from forty scaled model cars under reasonably controlled environments, and encouraging results were obtained. It is our belief that, with the advancement of image analysis and pattern recognition technologies, the proposed idea could evolve into a very promising application area where the auto insurance industry could significantly benefit.
Ying Li 0121, Chitra Dorai
MMSP1
2006 MAGICAL demonstration: system for automated metadata generation for instructional content
abstract
The "Tools for Automatic Generation of Learning Object Metadata" project addresses the requirement of developing advanced distributed learning delivery architecture and services for a large US government agency. We have developed a Webbased system called MAGIC (Metadata Automated Generation for Instructional Content) to assist content authors and course developers in generating metadata for learning objects and information assets to enable wider reuse of these objects across departments and organizations. Using the MAGIC system, content authors review and edit automatically-generated metadata sufficient to register and describe their assets for use and discovery in current and future distributed learning applications complying with the ADL SCORM standard. Course developers can use the system to assist in the conversion of existing courses to SCORM format or in developing new SCORM courses. The MAGIC system includes software tools to analyze and extract descriptive metadata from instructional videos, training documents, and other information assets. The tools generate some of the most critical SCORM metadata completely automatically. Benefits of MAGIC include easier reuse and repurposing, improved interoperability, and more timely registration of content for use by course developers. In this paper, we describe the system architecture, analysis tools developed, and services supported. A live demonstration of the system illustrating several use cases of the system will be presented at the conference, with a discussion of results from user studies and evaluation of the system.
Chitra Dorai, Robert G. Farrell, Amy Katriel, Galina Kofman, Ying Li 0121, Youngja Park
ACM Multimedia5
2006 Atomic topical segments detection for instructional videos
abstract
This paper presents our latest work on structuring instructional videos into units of atomic topical segment so as to facilitate topic-based video browsing and offer efficient video authoring. Specifically, we developed a comprehensive text analysis component to first extract informative text cues such as keyword synonym set and sentence boundary information, from a video's transcript. These text cues are then applied with various audiovisual cues such as silence/music break and speech similarity, to identify topical segments. Early experiments carried out on collections of real data from targeted user communities have yielded good results, and the user feedback on using the generated topical segment information is very encouraging.
Ying Li 0121, Youngja Park, Chitra Dorai
ACM Multimedia1
2006 Extracting Salient Keywords from Instructional Videos Using Joint Text, Audio and Visual Cues
Youngja Park, Ying Li 0121
HLT-NAACL2
2006 Instructional Video Content Analysis Using Audio Information
abstract
Automatic media content analysis and understanding for efficient topic searching and browsing are current challenges in the management of e-learning content repositories. This paper presents our current work on analyzing and structuralizing instructional videos using pure audio information. Specifically, an audio classification scheme is first developed to partition the sound-track of an instructional video into homogeneous audio segments where each segment has a unique sound type such as speech or music. We then apply a statistical approach to extract discussion scenes in the video by modeling the instructor with a Gaussian mixture model (GMM) and updating it on the fly. Finally, we categorize obtained discussion scenes into either two-speaker or multispeaker discussions using an adaptive mode-based clustering approach. Experiments carried out on four training videos and five IBM MicroMBA class videos have yielded encouraging results. It is our belief that by detecting and identifying various types of discussions, we are able to better understand and annotate the learning media content and subsequently facilitate its content access, browsing, and retrieval
Ying Li 0121, Chitra Dorai
IEEE Trans. Speech Audio Process.1
2005 An overview of technologies for e-meeting and e-lecture
abstract
Over the past few years, with the rapid adoption of broadband communication and advances in multimedia content capture and delivery, Web-based meetings and lectures, also referred to as e-meeting and e-lecture, have become popular among businesses and academic institutions because of their cost savings and capabilities in providing self-paced education and convenient content access and retrieval. In fact, the technological achievements in capture, analysis, access, and delivery of e-meeting and e-lecture media have already resulted in several working systems that are currently of regular usage. This paper gives an overview of existing work as well as state-of-the-art in these two research areas which are bound to affect the way we teach, learn, and collaborate.
Berna Erol, Ying Li 0121
ICME2
2005 Video Frame Identification for Learning Media Content Understanding
abstract
This paper presents our latest work on identifying frame content types for understanding learning media content. In particular, we categorize frames into six classes namely, slide, Web-page, instructor, audience, picture-in-picture and miscellaneous, which make up salient narrative modes in learning videos. Various image and video analysis approaches are explored to achieve this task. Preliminary experiments carried out on three recorded seminars have yielded encouraging results. The identification of fine-grained visual content types can assist us in content understanding, access, browsing and searching of generic learning videos
Ying Li 0121, Chitra Dorai
ICME1
2005 Creating MAGIC: system for generating learning object metadata for instructional content
abstract
This paper presents our latest work on building a system called MAGIC (Metadata Automated Generation for Instructional Content) that will automatically identify segments and generate critical metadata conforming with the SCORM (Sharable Content Object Reference Model) standard for instructional content. Various content analytics engines are utilized to automatically generate key metadata, which include audiovisual analysis modules that recognize semantic sound categories and identify narrators and informative text segments; text analysis modules that extract title, keywords and summary from text documents; and a text categorizer that classifies a document according to a pre-generated taxonomy. With MAGIC, instructional content developers can generate and edit SCORM metadata to richly describe their content asset for use in distributed learning applications. Experimental results obtained from collections of real data from targeted user communities will be presented.
Ying Li 0121, Chitra Dorai, Robert G. Farrell
ACM Multimedia1
2004 SVM-based audio classification for instructional video analysis
abstract
Automatic content analysis and annotation for efficient search and browsing of topics in instructional videos are current challenges in the management of e-learning content repositories. This paper presents our current work on classifying the soundtrack of instructional videos into seven distinct audio classes using the support vector machine (SVM) technology. The classification results are then used to partition a video into homogeneous audio segments, which forms the fundamental basis for its higher-level content analysis and exploration. Initial experiments carried out on three education and four training videos totalling 185 minutes have yielded an average 97.9% classification accuracy. The performance comparisons between the SVM-based, the decision tree (DT)-based and the threshold-based audio classification schemes further demonstrates the superiority of the proposed scheme.
Ying Li 0121, Chitra Dorai
ICASSP (5)1
2004 Detecting discussion scenes in instructional videos
abstract
This paper addresses the problem of detecting discussion scenes in instructional videos using statistical approaches. Specifically, given a series of speech segments separated from the audio tracks of educational videos, we first model the instructor using a Gaussian mixture model (GMM), then a four-state transition machine is designed to extract discussion scenes in real-time, based on detected instructor-student speaker change points. Meanwhile, we keep updating the GMM model to accommodate the instructor's voice variation along time. Promising experimental results have been achieved on five educational (IBM MicroMBA program) videos, and very interesting instruction/teaching patterns have been observed. The extracted scene information would facilitate the semantic indexing and structuralization of instructional video content.
Ying Li 0121, Chitra Dorai
ICME1
2004 Analyzing discussion scene contents in instructional videos
abstract
This paper describes our current effort on analyzing the contents of discussion scenes in instructional videos based on a clustering technique. Specifically, given a discussion scene pre-detected from an education or training video, we first apply a mode-based clustering approach to group all speech segments into an optimal number of clusters where each cluster contains speech from one speaker; we then analyze the discussion patterns in the scene, and subsequently classify it into either a 2-speaker or multi-speaker discussion. Encouraging classification results have been achieved on 122 discussion scenes detected from five IBM MicroMBA videos. Moreover, we have also observed fairly good performance on the speaker clustering scheme, which demonstrates the superiority of the proposed clustering approach. Undoubtedly, the discussion scene information output from this analysis scheme would facilitate the content browsing, searching and understanding of instructional videos.
Ying Li 0121, Chitra Dorai
ACM Multimedia1
2004 Adaptive speaker identification with audiovisual cues for movie content analysis
Ying Li 0121, Shri Narayanan, C.-C. Jay Kuo
Pattern Recognit. Lett.1
2004 Content-based movie analysis and indexing based on audiovisual cues
abstract
A content-based movie parsing and indexing approach is presented; it analyzes both audio and visual sources and accounts for their interrelations to extract high-level semantic cues. Specifically, the goal of this work is to extract meaningful movie events and assign them semantic labels for the purpose of content indexing. Three types of key events, namely, 2-speaker dialogs, multiple-speaker dialogs, and hybrid events, are considered. Moreover, speakers present in the detected movie dialogs are further identified based on the audio source parsing. The obtained audio and visual cues are then integrated to index the movie content. Our experiments have shown that an effective integration of the audio and visual sources can lead to a higher level of video content understanding, abstraction and indexing.
Ying Li 0121, Shri Narayanan, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.1
2003 Audiovisual-based adaptive speaker identification
abstract
An adaptive speaker identification system is presented in this paper, which aims to recognize speakers in feature films by exploiting both audio and visual cues. Specifically, the audio source is first analyzed to identify speakers using a likelihood-based approach. Meanwhile, the visual source is parsed to recognize talking faces using face detection/recognition and mouth tracking techniques. These two information sources are then integrated under a probabilistic framework for improved system performance. Moreover, to account for speakers' voice variations along time, we update their acoustic models on the fly by adapting to their newly contributed speech data. An average of 80% identification accuracy has been achieved on two test movies. This shows a promising future for the proposed audiovisual-based adaptive speaker identification approach.
Ying Li 0121, Shri Narayanan, C.-C. Jay Kuo
ICASSP (5)1
2003 Audiovisual-based adaptive speaker identification
abstract
An adaptive speaker identification system is presented in this paper, which aims to recognize speakers in feature films by exploiting both audio and visual cues. Specifically, the audio source is first analyzed to identify speakers using a likelihood-based approach. Meanwhile, the visual source is parsed to recognize talking faces using face detection/recognition and mouth tracking techniques. These two information sources are then integrated under a probabilistic framework for improved system performance. Moreover, to account for speakers' voice variations along time, we update their acoustic models on the fly by adapting to their newly contributed speech data. An average of 80% identification accuracy has been achieved on two test movies. This shows a promising future of the proposed audiovisual-based adaptive speaker identification approach.
Ying Li 0121, Shri Narayanan, C.-C. Jay Kuo
ICME1
2002 Identification of speakers in movie dialogs using audiovisual cues
abstract
The problem of identifying speakers from a movie dialog scene is addressed in this paper. While most previous work on speaker identification has been carried out using pure audio data, more robust results could be obtained by integrating the knowledge from multiple media sources such as visual and audio information when they are available. In this work, we first identify and isolate speech segments from background by applying video shot detection, audio classification and adaptive silence detection techniques, then a decision is made based on the calculated likelihood between the incoming speech data and pre-trained speaker/background models. Moreover, to verify the effectiveness of the adaptive silence detector, we have compared it with a statistically trained silence model. Experimental results show that the proposed algorithm can achieve approximately 84% identification accuracy by integrating multiple media cues.
Ying Li 0121, Shri Narayanan, C.-C. Jay Kuo
ICASSP1
2001 Semantic Video Content Abstraction Based On Multiple Cues
abstract
This research addresses the problem of automatically extracting video’s semantic structure and summarizing it in a hierarchical manner. Multiple media cues are employed in this procedure including visual, audio and text information. The generated hierarchy can provide us a compact yet meaningful abstraction of the video data similar to the conventional table-of-contents, which will facilitate user’s access to multimedia contents including browsing and retrieval. Preliminary experiments of integrating different media for hierarchically representing video semantics have yielded encouraging results.
Ying Li 0121, Wei Ming, C.-C. Jay Kuo
ICME1
2001 Video classification in user profile generation for personalized broadcast services
Ying Li 0121, C.-C. Jay Kuo
VCIP1