John R. Kender

dblp:l/JohnRKender · DBLP profile ↗
← Back
87ranked-venue papers
19as first author
0since 2021 · last 2015
0009-0009-2522-053XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 16 first-authorArtificial intelligence and machine learning · 34 · 12 first-authorDatabases, data management, data science and information retrieval · 9Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
24 papers
Multimedia analysis and retrieval · 90% Image and video processing · 6% Geometric modeling and processing · 2%
Databases, data mining, and information retrieval
9 papers
Data mining · 55% Web and social media mining · 25% Information retrieval · 19%
Human-computer interaction and pervasive computing
3 papers
Wearable and physiological sensing · 40% Collaborative and social computing · 40% Interaction techniques and input · 11%
Artificial intelligence
8 papers
Face, body and person analysis · 37% Segmentation and scene understanding · 28% Robot navigation and mapping · 20%

Topics — the 30 heaviest of 78, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval
video retrieval
0.222011
Visual memes in social media: tracking real-world news in YouTube videos · ACM Multimedia 2011
Evaluation of video browser features and user interaction with VAST MM · ACM Multimedia 2008
Data mining
pattern mining
0.122007
Optimizing Frequency Queries for Data Mining Applications · ICDM 2007
High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006
Information retrieval › multimedia analysis and retrieval
video browsing
0.112011
VastMM-Tag: a semantic tagging browser for unstructured videos · ACM Multimedia 2011
Multimedia analysis and retrieval › video annotation
semantic video annotation
0.112011
VastMM-Tag: a semantic tagging browser for unstructured videos · ACM Multimedia 2011
Multimedia analysis and retrieval
near-duplicate detection
0.112010
Video genetics: a case study from YouTube · ACM Multimedia 2010
Multimedia analysis and retrieval › near-duplicate detection
near-duplicate keyframe detection
0.112010
Video genetics: a case study from YouTube · ACM Multimedia 2010
Multimedia analysis and retrieval
video analysis
0.112010
Video genetics: a case study from YouTube · ACM Multimedia 2010
Multimedia analysis and retrieval
video classification
0.112010
Video genetics: a case study from YouTube · ACM Multimedia 2010
Multimedia analysis and retrieval
video indexing
0.122011
Augmented segmentation and visualization for presentation videos · ACM Multimedia 2005
Selecting the best faces to index presentation videos · ACM Multimedia 2011
Multimedia analysis and retrieval › video summarization
key frame extraction
0.122005
Augmented segmentation and visualization for presentation videos · ACM Multimedia 2005
Optimization Algorithms for the Selection of Key Frame Sequences of Variable Length · ECCV (4) 2002
Multimedia analysis and retrieval
video summarization
0.122005
Augmented segmentation and visualization for presentation videos · ACM Multimedia 2005
Video Summaries through Mosaic-Based Shot and Scene Clustering · ECCV (4) 2002
Data mining › predictive modeling
classification
0.112008
Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008
Data mining › predictive modeling › classification
pattern classification
0.112008
Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008
Data mining › text mining
text classification
0.112008
Classifying High-Dimensional Text and Web Data Using Very Short Patterns · ICDM 2008
Multimedia analysis and retrieval › multimedia browsing
video browsing
0.112008
Evaluation of video browser features and user interaction with VAST MM · ACM Multimedia 2008
Data mining › pattern mining › itemset mining
frequent itemset mining
0.112007
Optimizing Frequency Queries for Data Mining Applications · ICDM 2007
Data mining › pattern mining
support counting
0.112007
Optimizing Frequency Queries for Data Mining Applications · ICDM 2007
Web and social media mining
social media analysis
0.122011
Visual memes in social media: tracking real-world news in YouTube videos · ACM Multimedia 2011
Video genetics: a case study from YouTube · ACM Multimedia 2010
Information retrieval › ranking › graph-based ranking
pagerank
0.112015
Tracking Cultural Differences in News Video Creation · ACM Multimedia 2015
Information retrieval
ranking
0.112015
Tracking Cultural Differences in News Video Creation · ACM Multimedia 2015
Data mining
clustering
0.112006
High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006
Data mining › clustering
document clustering
0.112006
High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006
Data mining › pattern mining
itemset mining
0.112006
High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets · ICDM 2006
Multimedia analysis and retrieval
gesture analysis
0.112014
Correlating Speaker Gestures in Political Debates with Audience Engagement Measured via EEG · ACM Multimedia 2014
Multimedia analysis and retrieval
news story tracking
0.112005
Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment · CVPR (1) 2005
Multimedia analysis and retrieval
video annotation
0.112005
Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment · CVPR (1) 2005
Web and social media mining › social media analysis
social media content analysis
0.012013
Tracking Large-Scale Video Remix in Real-World Events · IEEE Trans. Multim. 2013
Image and video processing
image segmentation
0.022000
A Comparative Technique and Performance Results on Novel Learned Snakes in Two Dissimilar Medical Domains · CVPR 2000
Sectored Snakes: Evaluating Learned-Energy Segmentations · ICCV 1998
User interface design and tools › multimedia interface design
video browsing interface
0.012011
VastMM-Tag: a semantic tagging browser for unstructured videos · ACM Multimedia 2011
Mathematical optimization
combinatorial optimization
0.012002
Optimization Algorithms for the Selection of Key Frame Sequences of Variable Length · ECCV (4) 2002

Methods — techniques the papers use, named apart from their topics

spearman rank correlation · 0.4electroencephalography · 0.4machine learning semantic labeling · 0.4boolean algebra over tags · 0.4predictive modeling · 0.3influence measures · 0.2graph model · 0.2face tracking · 0.2face detection · 0.2temporal modeling · 0.2pagerank algorithm · 0.2latent topic models · 0.2latent topic model · 0.2visual meme detection · 0.1lempel-ziv encoding · 0.1user study · 0.1learned energy · 0.1active contour model · 0.1
YearPublicationVenuePosition
2015 Tracking Cultural Differences in News Video Creation
abstract
Many videos on the Web are created in different countries about the same international event. Their specialized video content, as well as their viewing and reposting rates, reflect different cultural interests. Effectively tracking cross-cultural visual memes of the same event, in online video repositories of different cultures, can provide users with a more comprehensive understanding of an international event. We propose a new way to use the PageRank algorithm to model cross-cultural visual meme influence, which more accurately captures the rates at which visual memes are re-posted in a specified time period in a specified culture.
Chun-Yu Tsai, John R. Kender
ACM Multimedia2
2014 Highly Efficient Multimedia Event Recounting from User Semantic Preferences
abstract
We present the design of a video event recounting system that takes YouTube-like videos, and identifies a minimal set of video segments and textual keyword descriptions in order to convince a user, in a time efficient manner, that the video contains an instance of a user-specificed human activity. The system is based on extensive user studies that have lead to nine design principles about human preferences and limits in semantic understanding. The processing pipeline locates the presence of user query keywords within the video, segments the video according to a model of human short-term memory for semantic similarities, selects those segments that best contain query terms, and abbreviates both the video and textual presentation. Speed-ups of a factor of 6 over simple video viewing time are achievable, without loss of semantic accuracy. In the 2013 Trecvid Multimedia Event Recounting competition, this system placed first in time efficiency, while remaining above average in description accuracy.
Chun-Yu Tsai, Michelle L. Alexander, Nnenna Okwara, John R. Kender
ICMR4
2014 Correlating Speaker Gestures in Political Debates with Audience Engagement Measured via EEG
abstract
We hypothesize that certain speaker gestures can convey significant information that are correlated to audience engagement. We propose gesture attributes, derived from speakers' tracked hand motions to automatically quantify these gestures from video. Then, we demonstrate a correlation between gesture attributes and an objective method of measuring audience engagement: electroencephalography (EEG) in the domain of political debates. We collect 47 minutes of EEG recordings from each of 20 subjects watching clips of the 2012 U.S. Presidential debates. The subjects are examined in aggregate and in subgroups according to gender and political affiliation. We find statistically significant correlations between gesture attributes (particularly extremal pose) and our feature of engagement derived from EEG both with and without audio. For some stratifications, the Spearman rank correlation reaches as high as rho = 0.283 with p < 0.05, Bonferroni corrected. From these results, we identify those gestures that can be used to measure engagement, principally those that break habitual gestural patterns.
John R. Zhang, Jason Sherwin, Jacek Dmochowski, Paul Sajda, John R. Kender
ACM Multimedia5
2013 Recognizing and tracking clasping and occluded hands
abstract
We present a purely algorithmic method for distinguishing when two hands are visually merged together and tracking their positions by propagating tracking information from anchor frames in single-camera video without depth information. We demonstrate and evaluate on a manually labeled dataset selected primarily for clasped hands with 698 images of a single speaker with 1301 annotated left and right hands. Toward the goal of recognizing clasping hands, our method performs better than baseline on recall (0.66 vs. 0.53) without sacrificing precision (0.65 for both). We also evaluate its tracking efficacy through its ability to affect performance of a naive hand labeling heuristic, resulting in an improvement over the baseline (F-score of 0.59 vs. 0.48 baseline).
John R. Zhang, John R. Kender
ICIP2
2013 Learning by focusing: A new framework for concept recognition and feature selection
abstract
In this paper, we develop a new method for feature selection and category learning. We first introduce two observations from our experiments: (1) It is easier to distinguish two concepts than to learn an isolated concept. (2) To distinguish different concept pairs we can find different selections of optimal features. These two observations may partly explain the success of human vision learning, especially why an infant can simultaneously capture distinguished visual features when learning new concepts. Based on these two observations, we developed a new learning-by-focusing method which first constructs focalized concept discriminators for pairs of concepts, and then builds nonlinear classifiers using the discrimination scores. We build datasets for four concept structure: vehicle, human affliction, sports, and animals, and experiments on all the four datasets verify the success of our new approach.
Liangliang Cao, Leiguang Gong, John R. Kender, Noel Codella, John R. Smith
ICME3
2013 Tracking Large-Scale Video Remix in Real-World Events
abstract
Content sharing networks, such as YouTube, contain traces of both explicit online interactions (such as likes, comments, or subscriptions), as well as latent interactions (such as quoting, or remixing, parts of a video). We propose visual memes, or frequently re-posted short video segments, for detecting and monitoring such latent video interactions at scale. Visual memes are extracted by scalable detection algorithms that we develop, with high accuracy. We further augment visual memes with text, via a statistical model of latent topics. We model content interactions on YouTube with visual memes, defining several measures of influence and building predictive models for meme popularity. Experiments are carried out with over 2 million video shots from more than 40,000 videos on two prominent news events in 2009: the election in Iran and the swine flu epidemic. In these two events, a high percentage of videos contain remixed content, and it is apparent that traditional news media and citizen journalists have different roles in disseminating remixed content. We perform two quantitative evaluations for annotating visual memes and predicting their popularity. The proposed joint statistical model of visual memes and words outperforms an alternative concurrence model, with an average error of 2% for predicting meme volume and 17% for predicting meme lifespan.
Lexing Xie, Apostol Natsev, Xuming He 0001, John R. Kender, Matthew L. Hill, John R. Smith
IEEE Trans. Multim.4
2012 Fast Near-Duplicate Video Retrieval via Motion Time Series Matching
abstract
This paper introduces a method for the efficient comparison and retrieval of near duplicates of a query video from a video database. The method generates video signatures from histograms of orientations of optical flow of feature points computed from uniformly sampled video frames concatenated over time to produce time series, which are then aligned and matched. Major incline matching, a data reduction and peak alignment method for time series, is adapted for faster performance. The resultant method is compact and robust against a number of common transformations including: flipping, cropping, picture-in-picture, photometric, addition of noise and other artifacts. We evaluate on the MUSCLE VCD 2007 dataset and a dataset derived from TRECVID 2009. Good precision (average 88.8%) at significantly higher speeds (average durations: 45 seconds for signature generation plus 92 seconds for a linear search of 81-second query video in a 300 hour dataset) than results reported in the literature are shown.
John R. Zhang, Jennifer Ren, Fangzhe Chang, Thomas L. Wood, John R. Kender
ICME5
2011 Identifying salient poses in lecture videos
abstract
The communicative importance of gestures in teaching environments have been widely studied. Two classes of gestures - point and spread gestures - have been identified to indicate pedagogical importance in teaching discourse [1]. In this work, we propose a system for the identification of the poses of point and spread gestures as a preliminary step toward their identification in low-quality unstructured videos. We use a joint-angle descriptor derived from an automatic pose estimation framework to train an SVM in order to classify extracted video frames of an instructor giving a lecture. Ground-truth is collected in the form of 2500 manually annotated frames covering approximately 20 minutes of a video lecture. Cross validation on the ground-truth data showed initial classifier F-scores of 0.54 and 0.39 for point and spread poses.
John R. Zhang, John R. Kender
ICIP2
2011 Tracking Visual Memes in Rich-Media Social Communities
Lexing Xie, Apostol Natsev, John R. Kender, Matthew L. Hill, John R. Smith
ICWSM3
2011 Selecting the best faces to index presentation videos
abstract
We propose a system to select the most representative faces in unstructured presentation videos with respect to two criteria: to optimize matching accuracy between pairs of face tracks, and to select humanly preferred face icons for indexing purposes. We first extract face tracks using state-of-the-art face detection and tracking. A small subset of images are then selected per track in order to maximize matching accuracy between tracks. Finally, representative images are extracted for each speaker in order to build a face index of the video. We tested our approach on 3 unstructured presentation videos of approximately 45 minutes each, for a total of a quarter million frames. Compared to the standard min-min approach, our method achieves higher track matching accuracy (94.22%), while using 6% of the running time. Using an optimal combination of 3 user preference measures, we were able to build face indexes containing 54 speakers (out of the 58 present in the videos) indexing into 795 detected tracks.
Michele Merler, John R. Kender
ACM Multimedia2
2011 VastMM-Tag: a semantic tagging browser for unstructured videos
abstract
Quickly accessing the contents of a video is challenging for users, particularly for unstructured video, which contains no intentional shot boundaries, no chapters, and no apparent edited format. We approach this problem in the domain of lecture videos using machine learning and semantic display techniques. We extend an existing video browser, through a display of these machine-learned semantic labelings to provide the user with a multi-timeline semantic view. Each timeline corresponds to one semantic label and indicates the label's probable presence or absence in the associated frames. We also provide a full Boolean algebra over these labels, in order to accommodate more complex queries, such as 'text or code, but no instructor'. Finally, we quantify the effectiveness of our features and our browser through user studies on various tasks. We find that users follow a nearly fixed pattern of access, alternating between the use of these tags and keyframes, and also between the use of 'word bubbles' and the player. We show that the tag algebra is integral to the time efficient use of tag timelines, saving up to 27% of the time for various retrieval tasks.
Mitchell J. Morris, John R. Kender
ACM Multimedia2
2011 Visual memes in social media: tracking real-world news in YouTube videos
abstract
We propose visual memes, or frequently reposted short video segments, for tracking large-scale video remix in social media. Visual memes are extracted by novel and highly scalable detection algorithms that we develop, with over 96% precision and 80% recall. We monitor real-world events on YouTube, and we model interactions using a graph model over memes, with people and content as nodes, and meme postings as links. This allows us to define several measures of influence. These abstractions, using more than two million video shots from several large-scale event datasets, enable us to quantify and efficiently extract several important observations: over half of the videos contain re-mixed content, which appears rapidly; video view counts, particularly high ones, are poorly correlated with the virality of content; the influence of traditional news media versus citizen journalists varies from event to event; iconic single images of an event are easily extracted; and content that will have long lifespan can be predicted within a day after it first appears. Visual memes can be applied to a number of social media scenarios: brand monitoring, social buzz tracking, ranking content and users, among others.
Lexing Xie, Apostol Natsev, John R. Kender, Matthew L. Hill, John R. Smith
ACM Multimedia3
2010 Video genetics: a case study from YouTube
abstract
We explore in a single but large case study how videos within YouTube, competing for view counts, are like organisms within an ecology, competing for survival. We develop this analogy, whose core idea shows that short video clips, best detected across videos as near-duplicate keyframes, behave similarly to genes. We report work in progress, on a dataset of 5.4K videos with 210K keyframes on a single topic, which traces sequences, not bags, of "near-dups" over time, both within videos and across them. We demonstrate their utility to: cleanse responses to queries contaminated by over-eager YouTube query expansion; separate videos temporally according to their responses to external events; track the evolution and lifespan of continuing video "stories"; automatically locate video summaries already present within a video ecology; quickly verify video copying via a direct application of the Smith-Waterman algorithm used in genetics - which also provides useful feedback for tuning the near-dup detection and clustering process; and quickly classify videos via a kind of Lempel-Ziv encoding into the categories of news, monologue, dialogue, and slideshow. We demonstrate a number of novel visualizations of this large dataset, including a direct use of the Matlab black-body "hot" false-color map, together with the GraphViz package, to display the gene-like inheritance of viral properties of keyframes. We further speculate that, as with genes, there are "functional roles" for semantic categories of clips, and, as with species, there are differing rates of "genetic drift" for each video genre.
John R. Kender, Matthew L. Hill, Apostol Natsev, John R. Smith, Lexing Xie
ACM Multimedia1
2010 Hierarchical document clustering using local patterns
Hassan H. Malik, John R. Kender, Dmitriy Fradkin, Fabian Mörchen
Data Min. Knowl. Discov.2
2009 Semantic keyword extraction via adaptive text binarization of unstructured unsourced video
abstract
We propose a fully automatic method for summarizing and indexing unstructured presentation videos based on text extracted from the projected slides. We use changes of text in the slides as a means to segment the video into semantic shots. Unlike precedent approaches, our method does not depend on availability of the electronic source of the slides, but rather extracts and recognizes the text directly from the video. Once text regions are detected within keyframes, a novel binarization algorithm, Local Adaptive Otsu (LOA), is employed to deal with the low quality of video scene text, before feeding the regions to the open source Tesseract1OCR engine for recognition. We tested our system on a corpus of 8 presentation videos for a total of 1 hour and 45 minutes, achieving 0.5343 Precision and 0.7446 Recall Character recognition rates, and 0.4947 Precision and 0.6651 Recall Word recognition rates. Besides being used for multimedia documents, topic indexing, and cross referencing, our system can be integrated into summarization and presentation tools such as the VAST MultiMedia browser.
Michele Merler, John R. Kender
ICIP2
2009 Sort-Merge feature selection and fusion methods for classification of unstructured video
abstract
We explore the problem of rapid automatic semantic tagging of video frames of unstructured (unedited) videos. We apply the sort-merge algorithm for feature selection on a large (>1000) heterogeneous feature set for videos showing lectures, to quickly locate low-level image features most predictive for concepts such as "key frame with text" or "key frame with computer source code". For evaluation, we introduce a "keeper" heuristic for feature retention, which provides a baseline comparison. We then compare early fusion and late fusion of diverse feature types; based on experiments on 12,395 frames, we find that in general late computation cost, compared to early fusion. However, mergers of redundant feature types do not necessarily improve performance over single feature types; exploration of both merged and unmerged performance is necessary.
Mitchell J. Morris, John R. Kender
ICME2
2008 Classifying High-Dimensional Text and Web Data Using Very Short Patterns
abstract
In this paper, we propose the "democratic classifier", a simple pattern-based classification algorithm that uses very short patterns for classification, and does not rely on the minimum support threshold. Borrowing ideas from democracy, our training phase allows each training instance to vote for an equal number of candidate size-2 patterns. The training instances select patterns by effectively balancing between local, class, and global significance of patterns. The selected patterns are simultaneously added to the model for all applicable classes and a novel power law based weighing scheme adjusts their weights with respect of each class. Results of experiments performed on 121 common text and Web datasets show that our algorithm almost always outperforms state of the art classification algorithms, without any parameter tuning. On 100 real-life Web datasets, the average absolute classification accuracy improvement was as great as 9.4% over SVM, Harmony, C4.5 and KNN. Also, our algorithm ran about 3.5 times faster than the fastest existing pattern-based classification algorithm.
Hassan H. Malik, John R. Kender
ICDM2
2008 Accommodating sample size effect on similarity measures in speaker clustering
abstract
We investigate the symmetric Kullback-Leibler (KL2) distance in speaker clustering and its unreported effects for differently-sized feature matrices. Speaker data is represented as Mel frequency cepstral coefficient (MFCC) vectors, and features are compared using the KL2 metric to form clusters of speech segments for each speaker. We make two observations with respect to clustering based on KL2: 1.) The accuracy of clustering is strongly dependent on the absolute lengths of the speech segments and their extracted feature vectors. 2.) The accuracy of the similarity measure strongly degrades with the length of the shorter of the two speech segments. These effects of length can be attributed to the measure of covariance used in KL2. We demonstrate an empirical correction of this sample-size effect that increases clustering accuracy. We draw parallels to two vector quantization-based (VQ) similarity measures, one which exhibits an equivalent effect of sample size, and the second being less influenced by it.
Alexander Haubold, John R. Kender
ICME2
2008 Separability and refinement of hierarchical semantic video labels and their ground truth
abstract
We investigate several problems in the annotation of video shots by semantic labels which are implicitly embedded in a semantic hierarchy, leading to analyses and novel methods for refining video ontologies and their ground truth. First, in the large 449 LSCOM semantic concept data set, we show that within the implicit ldquouse ontologyrdquo, many concepts tags are ambiguous as to purposeful activity, visual scope, or social agency, or are absent altogether, but that better ldquouse sensesrdquo can be refined algorithmically. Second, we find that both traditional hard and fuzzy k-medoid clustering techniques are inadequate for hierarchical concepts, but a novel ldquofirm k-medoidrdquo clustering method both separates clusters and distributes superconcepts equitably. Third, we show how the scores of SVM semantic filters can be more reliably and quickly converted to probabilities by using a closed-form approximation to SVM behavior between its margins. Fourth, we show that the quality of SVM semantic filters for hierarchical concepts can be analyzed by their ability to separate their positive ground truth examples from those of any other concept in the hierarchy; the most discriminating are those with ground truth showing distinctive physical backgrounds.
John R. Kender
ICME1
2008 Evaluation of video browser features and user interaction with VAST MM
abstract
In this paper, we present extensive user studies on browsing and information retrieval in the domain of unstructured videos using the VAST MM video library browser. Our studies were performed over a 3-year period with more than 1,000 participants in the university setting. The majority of students use the video library for retrieval of student presentations in a large engineering design course. Through iterative analysis of context-specific audio, visual, and textual cues, we are able to measure significant improvements on typical retrieval tasks, such as searching for unfamiliar content in a large database with over 300 hours of video. We also present user studies conducted in two videotaped core computer science courses to measure the usefulness of the VAST MM (Video Audio Structure Text MultiMedia) resource for final exam preparation. We find that students who use the lecture video library experience significant improvement in final exam scores.
Alexander Haubold, Promiti Dutta, John R. Kender
ACM Multimedia3
2008 Robust Dominant Motion Estimation Using MPEG Information in Sport Sequences
abstract
In this paper, we introduce a new method to estimate a parametric description of the dominant motion existing in a video sequence, a key task needed to face more complex video analysis problems. In order to do so, we use motion data provided by the MPEG streams. We propose a method based on imaginary straight line tracking to retrieve the projective transformations that describe the dominant motion of a sequence by estimating 2-D homographies. Our method takes advantage not only of the MPEG motion data, but also of its structure. In order to overcome the noise introduced by the MPEG compression scheme, we also employ several robust estimators to attain reliability. We demonstrate its performance by displaying several image mosaics derived from the motion in real time.
Ramon Lluis Felip, Lluis Barceló, Xavier Binefa, John R. Kender
IEEE Trans. Circuits Syst. Video Technol.4
2007 Optimizing Frequency Queries for Data Mining Applications
abstract
Data mining algorithms use various Trie and bitmap-based representations to optimize the support (i.e., frequency) counting performance. In this paper, we compare the memory requirements and support counting performance of FP Tree, and Compressed Patricia Trie against several novel variants of vertical bit vectors. First, borrowing ideas from the VLDB domain, we compress vertical bit vectors using WAH encoding. Second, we evaluate the Gray code rank- based transaction reordering scheme, and show that in practice, simple lexicographic ordering, obtained by applying LSB Radix sort, outperforms this scheme. Led by these results, we propose HDO, a novel Hamming-distance-based greedy transaction reordering scheme, and aHDO, a linear-time approximation to HDO. We present results of experiments performed on 15 common datasets with varying degrees of sparseness, and show that HDO- reordered, WAH encoded bit vectors can take as little as 5% of the uncompressed space, while aHDO achieves similar compression on sparse datasets. Finally, with results from over a billion database and data mining style frequency query executions, we show that bitmap-based approaches result in up to hundreds of times faster support counting, and HDO-WAH encoded bitmaps offer the best space-time tradeoff.
Hassan H. Malik, John R. Kender
ICDM2
2007 Alignment of Speech to Highly Imperfect Text Transcriptions
abstract
We introduce a novel and inexpensive approach for the temporal alignment of speech to highly imperfect transcripts from automatic speech recognition (ASR). Transcripts are generated for extended lecture and presentation videos, which in some cases feature more than 30 speakers with different accents, resulting in highly varying transcription qualities. In our approach we detect a subset of phonemes in the speech track, and align them to the sequence of phonemes extracted from the transcript. We report on the results for 4 speech-transcript sets ranging from 22 to 108 minutes. The alignment performance is promising, showing a correct matching of phonemes within 10, 20, 30 second error margins for more than 60 %, 75 %, 90 % of text, respectively, on average. For perfect manually generated transcripts, more than 75 % of text is correctly aligned within 5 seconds.
Alexander Haubold, John R. Kender
ICME2
2007 Analysis, User Interface, and their Evaluation for Student Presentation Videos
abstract
In the domain of candidly-captured student presentation videos, we examine and evaluate approaches for multimodal analysis and indexing of audio and video. We apply visual segmentation techniques on unedited video to determine likely changes of topics. Speaker segmentation methods are employed to determine individual student appearances, which are linked to extracted headshots to create a visual speaker index. Videos are augmented with time-aligned filtered keywords and phrases from highly inaccurate speech transcripts. An experimental user interface (UI) combines streaming videos, visual, and textual indices for browsing and searching. We evaluate the UI and methods in a large engineering design course. We report on observations and statistics collected over 4 semesters and 598 student participants. Results suggest that our video indexing and retrieval approach is effective, and that our continuous improvements are reflected in an increase in accuracy and completion rates of user study tasks.
Alexander Haubold, John R. Kender
ICME2
2007 A Large Scale Concept Ontology for News Stories: Empirical Methods, Analysis, and Improvements
abstract
We analyze the completeness, accuracy, and utility of the largest known annotation ground truth database for video news stories, comprising nearly 680K individual tags on 62 K shots using a vocabulary of 449 semantic concepts. We find the vocabulary is not yet mature: it does not follow Zipf's law, although concepts derived from vocabulary intersection do so more closely. We find that because many concepts are sparse, the best method for exposing the implicit semantic space is to use the distance measure G2, complete link clustering, and a heuristic distance cutoff based on shot cluster evolution history; this yields 12 well-defined major shot categories. Because the database is errorful, we derive a model for annotator error, and using it, we extract a natural concept subsumption ontology from the database, including some counter-intuitive relationships. Again using intersection, we demonstrate a method for identifying "missing" subconcepts from superconcepts. Lastly, we note that without superconcepts, shot clustering fails. These methods are unbiased by any prior semantic assumptions, and only depend on a statistically sufficient body of ground truth. They are therefore applicable to any other specific video retrieval domain.
John R. Kender
ICME1
2007 Computational approaches to temporal sampling of video sequences
abstract
Video key frame extraction is one of the most important research problems for video summarization, indexing, and retrieval. For a variety of applications such as ubiquitous media access and video streaming, the temporal boundaries between video key frames are required for synchronizing visual content with audio. In this article, we define temporal video sampling as a unified process of extracting video key frames and computing their temporal boundaries, and formulate it as an optimization problem. We first provide an optimal approach that minimizes temporal video sampling error using a dynamic programming process. The optimal approach retrieves a key frame hierarchy and all temporal boundaries in O ( n 4 ) time and O ( n 2 ) space. To further reduce computational complexity, we also provide a suboptimal greedy algorithm that exploits the data structure of a binary heap and uses a novel “look-ahead” computational technique, enabling all levels of key frames to be extracted with an average-case computational time of O ( n log n ) and memory usage of O ( n ). Both the optimal and the greedy methods are free of parameters, thus avoiding the threshold-selection problem that exists in other approaches. We empirically compare the proposed optimal and greedy methods with several existing methods in terms of video sampling error, computational cost, and subjective quality. An evaluation of eight videos of different genres shows that the greedy approach achieves performance very close to that of the optimal approach while drastically reducing computational cost, making it suitable for processing long video sequences in large video databases.
Tiecheng Liu, John R. Kender
ACM Trans. Multim. Comput. Commun. Appl.2
2006 High Quality, Efficient Hierarchical Document Clustering Using Closed Interesting Itemsets
abstract
High dimensionality remains a significant challenge for document clustering. Recent approaches used frequent itemsets and closed frequent itemsets to reduce dimensionality, and to improve the efficiency of hierarchical document clustering. In this paper, we introduce the notion of "closed interesting" itemsets (i.e. closed itemsets with high interestingness). We provide heuristics such as "super item" to efficiently mine these itemsets and show that they provide significant dimensionality reduction over closed frequent itemsets. Using "closed interesting" itemsets, we propose a new, sub-linearly scalable, hierarchical document clustering method that outperforms state of the art agglomerative, partitioning and frequent-itemset based methods both in terms of clustering quality and runtime performance, without requiring dataset specific parameter tuning. We evaluate twenty interestingness measures and show that when used to generate "closed interesting" itemsets, and to select parent nodes, mutual information, added value, Yule's Q and Chi- Square offer best clustering performance.
Hassan H. Malik, John R. Kender
ICDM2
2006 Video News Shot Labeling Refinement via Shot Rhythm Models
abstract
We present a three-step post-processing method for increasing the precision of video shot labels in the domain of television news. First, we demonstrate that news shot sequences can be characterized by rhythms of alternation (due to dialogue), repetition (due to persistent background settings), or both. Thus a temporal model is necessarily third-order Markov. Second, we demonstrate that the output of feature detectors derived from machine learning methods (in particular, from SVMs) can be converted into probabilities in a more effective way than two suggested existing methods. This is particularly true when detectors are errorful due to sparse training sets, as is common in this domain. Third, we demonstrate that a straightforward application of the Viterbi algorithm on a third-order FSM, constructed from observed transition probabilities and converted feature detector outputs, can refine feature label precision at little cost. We show that on a test corpus of TRECVID 2005 news videos annotated with 39 LSCOM-lite features, the mean increase in the measure of average precision (AP) was 4%, with some of the rarer and more difficult features having relative increases in AP of as much as 67%
John R. Kender, Milind R. Naphade
ICME1
2006 Clustering web images using association rules, interestingness measures, and hypergraph partitions
abstract
This paper presents a new approach to cluster web images. Images are first processed to extract signal features such as color in HSV format and quantized orientation. Web pages referring to these images are processed to extract textual features (keywords) and feature reduction techniques such as stemming, stop word elimination, and Zipf's law are applied. All visual and textual features are used to generate association rules. Hypergraphs are generated from these rules, with features used as vertices and discovered associations as hyperedges. Twenty-two objective interestingness measures are evaluated on their ability to prune non-interesting rules and to assign weights to hyperedges. Then a hypergraph partitioning algorithm is used to generate clusters of features, and a simple scoring function is used to assign images to clusters. A tree-distance-based evaluation measure is used to evaluate the quality of image clustering with respect to manually generated ground truth. Our experiments indicate that combining textual and content-based features results in better clustering as compared to signal-only or text-only approaches. Online steps are done in real-time, which makes this approach practical for web images. Furthermore, we demonstrate that statistical interestingness measures such as Correlation Coefficient, Laplace, Kappa and J-Measure result in better clustering compared to traditional association rule interestingness measures such as Support and Confidence.
Hassan H. Malik, John R. Kender
ICWE2
2006 Designing an intelligent user interface for instructional video indexing and browsing
abstract
Instructional videos are used intensively in universities for remote education and e-learning, and a typical university course consists of videos of more than two thousand minutes in total length. This paper presents a novel graphics user interface for indexing and browsing such extensive but thematically related content. We present how the interface automatically extracts semantic indices from the visual content, and then presents both high- and low-level cues from five different conceptual viewpoints. We detail each of these novel UI units, and show how they are integrated into a user-adjustable main framework, and interconnected and navigated through user mouse events.
Lijun Tang, John R. Kender
IUI2
2005 Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment
abstract
In the summer of 2003, using an interactive intelligent tool, over 100 researchers in video understanding annotated from the NIST TRECVID database over 62 hours of news video spanning six months of 1998. These 47K shots with 43 3 K labels from over 1000 visual concept categories comprise the largest publicly available ground truth for this domain. Our analysis of this data, combining the tools of statistical natural language processing, machine learning, and computer vision, finds significant novel statistical patterns that can be exploited for the accurate tracking of the episodes of a given news story over time, by using semantic labels that are solely visual. We find that the ground "truth" is very muddy, but by using the feature selection tool of information gain, we extract 14 reliable visual concepts with mid-frequency use; all but one are visual concepts that refer to settings, rather than actors, objects, or events. We discover that the probability of another episode of a named story to recur after a gap of d days is proportional to 1/(d + 1). We define a novel similarity measure incorporating both semantic and temporal properties between episodes i and j as: Dice(i, j)/(1 + gap(i, j)). We exploit a low-level computer vision technique, normalized cut (Laplacian eigenmaps), for clustering these episodes into stories, and in the process document a weakness of this popular technique. We use these empirical results to make specific recommendations on how better visual semantic ontologies for news stories, and how better video annotation tools, should be designed.
John R. Kender, Milind R. Naphade
CVPR (1)1
2005 Educational Video Understanding: Mapping Handwritten Text to Textbook Chapters
abstract
Handwritten text frames appear frequently in educational videos and can be used as an important cue for semantic analysis of educational videos. We detect text frames using a motion pattern analyzing algorithm. Then, we extract binary handwritten word images from the text frames in various visual formats: handwritten slides, electronic slides, handwriting on chalkboard, etc. We propose a handwritten word recognition method, using combined dynamic programming stroke-based character segmentation with optimal statistical handwritten character recognition. In parallel, we construct a small vocabulary from topic words taken from table-of-contents of course materials such as the course textbook. We use the handwritten word recognition results to query this table-of-contents structure, implemented as latent semantic analysis matrix operations. We are able to spot the most likely discussed chapters and topic words for each frame. We evaluate the overall approach on 12 videos of two courses, and the results are encouraging.
Lijun Tang, John R. Kender
ICDAR2
2005 A unified text extraction method for instructional videos
abstract
Videotext can be an efficient semantic index and summary for instructional videos. However, videotext usually appears in different visual formats: handwritten slides, electronic slides, book pages, web pages, handwriting on chalkboard, etc. We propose a unified approach to handle all these kinds of videotext in three steps. First, we detect still video segments by analyzing motion energy patterns in instructional videos, and construct a quality-enhanced candidate text frame for each still video segment. Then, we use a trained SVM classifier to verify the candidate text frames, as well as to segment the text region and individual text blocks from the verified frames. Finally, we filter redundant text frames with similar text content by a Hausdorff distance-based image comparison algorithm. The resulting text frames are automatically organized into HTML and PDF documents to serve as an imagery summarization of the instructional videos. We show the application of our method to 75 instructional videos of five different courses, and discuss its applications.
Lijun Tang, John R. Kender
ICIP (3)2
2005 User Study for Generating Personalized Summary Profiles
abstract
The need for personalized summaries of media content has been driven by the recent and anticipated explosive growth in the media world. In this paper, we present a methodology and a supporting user study for generating user profiles and content features that can be used to automatically create personalized summaries of broadcast television content. We determined a mapping, from users' personality traits measured by commonly available personality tests, to computable video features that such personality traits appear to prefer. Three common personality profiles (Myers-Briggs, Merrill Reed, and Brain.exe) were elicited from 59 subjects, together with their preferred summary of news, music, and talk show videos. A factor analysis between the personality traits and the features in preferred summaries indicated that only some traits (e.g., gender, extroversion, control orientation, intuitiveness, etc.) and only some features (e.g., faces, reportage, text, chorus, host, etc.) had predictive value. The mapping of personality to feature also differed by genre. However, in general, extroverted users tended to prefer directly experienced content, while introverted users preferred content mediated through analysis. A validation user study is in progress
Lalitha Agnihotri, John R. Kender, Nevenka Dimitrova, John Zimmerman
ICME2
2005 Ontology Design for Video Semantic Threads
abstract
We propose that, at the highest level of video understanding, the human needs for meaning and the methodologies to extract it are both universal and generic. One must develop an ontology, then develop analyzers that learn the statistical correlates of that ontology, and finally use the analyzers to tie together common occurrences across individual videos. The first step towards adapting the ontology to the genre is the design of automated tools to assist in the annotation of the ground truth; these tools in turn provide feedback on the appropriateness of the filters and the ontology. We support this hypothesis by presenting and discussing some experiments conducted on the NIST TRECVID 2003 video corpus. We also validate this hypothesis by showing the connection between story tracking in our multimedia news and topic detection and tracking in the NIST TDT natural language effort. At the highest level, we find that our annotation tool shows that semantic concepts tend to cluster reliably into a few significant semantic dimensions. For news videos specifically, two of these clusters measure "presidentiality" and "outdoor-ness"
John R. Kender, Milind R. Naphade
ICME1
2005 Semantic Indexing for Instructional Video Via Combination of Handwriting Recognition and Information Retrieval
abstract
Efficient indexing and retrieval of digital videos are important needs within instructional video databases. Semantic indexing for instructional videos can be achieved by combining the analysis of the instructor's handwriting in the video with domain knowledge taken from course support materials such as the course textbook, syllabus, or slides. We propose such a semantic indexing method, by combining handwritten word recognition with information retrieval techniques. We first present a novel handwritten word segmentation and recognition approach for instructional videos. Then we construct a table-of-contents (TOC) structure from course materials. We use word recognition results to query the TOC, implemented as matrix operations, and spot the most likely discussed chapters and topic words for each video. We evaluate the overall approach on 12 videos of two courses, and the results are encouraging.
Lijun Tang, John R. Kender
ICME2
2005 Augmented segmentation and visualization for presentation videos
abstract
We investigate methods of segmenting, visualizing, and indexing presentation videos by both audio and visual data. The audio track is segmented by speaker, and augmented with key phrases which are extracted using an Automatic Speech Recognizer (ASR). The video track is segmented by visual dissimilarities and changes in speaker gesturing, and augmented by representative key frames. An interactive user interface combines a visual representation of audio, video, text, key frames, and allows the user to navigate presentation videos. User studies with 176 students of varying knowledge were conducted on 7.5 hours of student presentation video (32 presentations). Tasks included searching for various portions of presentations, both known and unknown to students, and summarizing presentations given the annotations. The results are favorable towards the video summaries and the interface, suggesting faster responses by a factor of 20% compared to having access to the actual video. Accuracy of responses remained the same on average. Follow-up surveys present a number of suggestions towards improving the interface, such as the incorporation of automatic speaker clustering and identification, and the display of an abstract topological view of the presentation. Surveys also show alternative contexts in which students would like to use the tool in the classroom environment.
Alexander Haubold, John R. Kender
ACM Multimedia2
2004 Design and evaluation of a music video summarization system
abstract
We present a system that summarizes the textual, audio, and video information of music videos in a format tuned to the preferences of a focus group of 20 users. First, we analyzed user-needs for the content and the layout of the music summaries. Then, we designed algorithms that segment individual song videos from full music video programs by noting changes in color palette, transcript, and audio classification. We summarize each song with automatically selected high level information such as title, artist, duration, title frame, and text as well as audio and visual segments of the chorus. Our system automatically determines with high recall and precision chorus locations, from the placement of repeated words and phrases in the text of the song's lyrics. Our Bayesian belief network then selects other significant video and audio content from the multiple media. Overall, we are able to compress content by a factor of 10. Our second user study has identified the principal variations between users in their choices of content desired in the summary, and in their choices of the platforms that should support their viewing.
Lalitha Agnihotri, Nevenka Dimitrova, John R. Kender
ICME3
2004 Video feature selection using fast-converging sort-merge tree
abstract
High time complexity is a bottle-neck in video segmentation, classification, analysis, and retrieval. In This work we use a heuristic method called fast-converging sort-merge tree (FSMT) to construct automatically a hierarchy of small subsets of features that are progressively more useful for video data exploration. The method combines the virtues of a wrapper model approach for high accuracy, with those of a filter method approach for deriving the appropriate features quickly. FSMT speeds up a more fundamental method, the basic sort-merge tree (BSMT) approach, while retaining its performance. We demonstrate FSMT's high accuracy: it has a 0.001 error rate in a frame classification task on 75 minutes of instructional video, and a 0.98 precision and 0.89 recall in a segment retrieval task on 30 minutes of sports video. Additionally, FSMT is more than 80% faster than its predecessor, BSMT.
Yan Liu 0004, John R. Kender
ICME2
2004 A method and user interface for instructional video indexing via recognition of handwritten table-of-contents words
abstract
Efficient indexing and retrieval of digital videos are important needs within instructional video databases. We construct a word-based video index by spotting words taken from the table of contents (TOC) of the textbooks for the course given in the video. We present a segment-free method of word spotting based on a two step analysis. First, for each candidate word in the vocabulary and each unknown handwritten word image, we search for the location, orientation, and scale of every character of the candidate word; we use a modified Hausdorff distance (MHD) to robustly locate the character's skeleton in the skeletons of the handwritten word image. Second, we verify words by creating a template image according to the result of character analysis, and by matching this template image to the skeleton of the handwritten word using the MHD. Results for recognition of handwritten words from real videos are shown (greater than 70% accuracy with 100 words). Applying these methods to three contiguous videos (of 25) of a course, we find that they recognize nearly disjoint subsets of highly descriptive indexing keywords.
Lijun Tang, John R. Kender
ICME2
2004 Video summaries and cross-referencing through mosaic-based representation
Aya Aner-Wolf, John R. Kender
Comput. Vis. Image Underst.2
2003 Semantic mosaic for indexing and compressing instructional videos
abstract
A new approach for content analysis and semantic compression of instructional videos is presented. Cameras often only capture a portion of a hand-drawn slide or blackboard panel that dominate this genre. We therefore have designed a novel semantic technique that retrieves the teaching content (text lines, figures, emphasis marks, etc.) visible in a frame, de-skews them and enhances their contrast, and stitches them back together into "virtual" slides that maintain the relative spatial relations of the captured fragments. Since this semantic mosaicing does not need precise pixel-wise matching, it has low computational complexity. These virtual slides maintain the useful teaching content of their real-world counterparts, and we further stitch them together to form a content summary slide to summarizes an extended video segment. By detecting content and recording its temporal development, we can then semantically compress a video into several such higher quality mosaics, and reconstruct the original instructional video by displaying the mosaics instead at the appropriate times. Our design runs in one pass and in real time. We apply the system to a 17-minute instructional video, and show the results of its semantic compression into virtual slide mosaics.
Tiecheng Liu, John R. Kender
ICIP (1)2
2003 Study on requirement specifications for personalized multimedia summarization
abstract
The ability to summarize and abstract information is an essential part of intelligent behavior in consumer devices. However, even for the emerging area of video content analysis, user preferences on summarization have not been explored. This paper reports on a panel which asked: who, why and when summarization is needed; what information should be summarized; and what forms should summaries take. In particular, we investigated the requirements with sensitivity to user needs, user context, media content, device capabilities, and the methods by which user and environment profiles can be assembled and exploited. We organize our findings as a wish list containing four major user questions, and we report the current and near-term states of the art and the major technical challenges in satisfying them. Our study suggests that user preferences should be derived from explicit user statements and from implicit trends inferred from viewing histories and summary usage.
Lalitha Agnihotri, Nevenka Dimitrova, John R. Kender, John Zimmerman
ICME3
2003 Analysis and interface for instructional video
abstract
We present a new method for segmenting, and a new user interface for indexing and visualizing, the semantic content of extended instructional videos. Using various visual filters, key frames are first assigned a media type (board, class, computer, illustration, podium, and sheet). Key frames of media type board and sheet are then clustered based on contents via an algorithm with near-linear cost. A novel user interface, the result of two user studies, displays related topics using icons linked topologically, allowing users to quickly locate semantically related portions of the video. We analyze the accuracy of the segmentation tool on 17 instructional videos, each of which is from 75 to 150 minutes in duration (a total of 40 hours); it exceeds 96%.
Alexander Haubold, John R. Kender
ICME2
2003 Fast scene segmentation using multi-level feature selection
abstract
High time cost is the bottle-neck of video scene segmentation. In this paper we use a heuristic method called sort-merge feature selection to construct automatically a hierarchy of small subsets of features that are progressively more useful for segmentation. A novel combination of fastmap for dimensionality reduction and Mahalanobis distance for likelihood determination is used as induction algorithm. Because these induced feature sets from a hierarchy with increasing classification accuracy, video segments can be segmented and categorized simultaneously in a coarse-fine manner that efficiently and progressively detects and refines their temporal boundaries. We analyze the performance of these methods, and demonstrate them in the domain of long (75 minute) instructional video.
Yan Liu 0004, John R. Kender
ICME2
2003 Music videos miner
abstract
No abstract available.
Lalitha Agnihotri, Nevenka Dimitrova, John R. Kender, John Zimmerman
ACM Multimedia3
2003 Sort-Merge Feature Selection for Video Data
abstract
Applying existing feature selection algorithms to video classification is impractical. A novel algorithm called Basic Sort-Merge Tree (BSMT) is proposed to choose a very small subset of features for video classification in linear time in the number of features. We reduce the cardinality of the input data by sorting the individual features by their effectiveness in categorization, and then merging pairwise these features into feature sets of cardinality two. Repeating this Sort-Merge process several times results in the learning of a small-cardinality, efficient, but highly accurate feature set. As the wrapper model, this paper exploits a novel combination of Fastmap for dimensionality reduction and Mahalanobis distance for likelihood determination. The time complexity of this induction part is linear in the number of training data. We provide theoretical proof of time cost and empirical validation of the accuracy.
Yan Liu 0004, John R. Kender
SDM2
2003 Fast video segment retrieval by Sort-Merge feature selection, boundary refinement, and lazy evaluation
Yan Liu 0004, John R. Kender
Comput. Vis. Image Underst.2
2002 Video Summaries through Mosaic-Based Shot and Scene Clustering
Aya Aner, John R. Kender
ECCV (4)2
2002 Optimization Algorithms for the Selection of Key Frame Sequences of Variable Length
Tiecheng Liu, John R. Kender
ECCV (4)2
2002 Rule-based semantic summarization of instructional videos
abstract
We present a new content-based approach to summarize instructional videos. We first redefine "scene" in instructional videos. Focusing on one dominant scene type, that of handwritten lecture notes, we define semantic content as "ink pixels", and present a low-level retrieval technique to extract this content from each frame with consideration of various occlusion and illumination effects. "Key frames" in this video genre are redefined as a set of frames that cover the semantic content, and the fluctuating amount of visible ink is used to drive a heuristic real-time key frame extraction method. A rule-based method is also provided to synchronize key frames with audio. We extend our method to the extraction of key frame hierarchies. We show its application to a 17-minute (30K frames) instructional video sequence, resulting in seven key frames. These techniques create tunable instructional summaries over a wide and dynamically varying range of compression factors.
Tiecheng Liu, John R. Kender
ICIP (1)2
2002 A method and browser for cross-referenced video summaries
abstract
We present an automatic tool for compact representation and cross-referencing of long video sequences, which is based on a novel visual abstraction of semantic content. Our highly compact hierarchical representation results from the non-temporal clustering of scene segments into a new conceptual form grounded in the recognition of real-world backgrounds. We represent shots and scenes using mosaics and employ a novel method for the comparison of scenes based on these representative mosaics. We then cluster scenes together into a higher level of abstraction-the physical setting. We demonstrate our work using situation comedies (sitcoms), where each half-hour episode is well structured by rules governing background use. Consequently, browsing, indexing and comparison across videos by physical setting is very fast. Further, we show that physical settings lead to a higher-level contextual identification of the main plots in each video. We demonstrate these contributions with a browsing tool whose top-level single page displays the settings of several episodes. This page expands to display windows for each episode, and each episode menu summary is further expanded into scenes and shots, all by mouse-clicking on appropriate plots and settings according to user interests.
Aya Aner, Lijun Tang, John R. Kender
ICME (2)3
2002 Analysis and enhancement of videos of electronic slide presentations
abstract
This paper presents a new approach to indexing videos of presentations which use electronic slides. By determining the images of slides in the video frames, and then matching the video sequences to the original electronic slides, the video can be indexed and searched, and the visual appearance of the segments can be improved. We first detect the "content area" in video frames using a color similarity weighted least square method. By monitoring the frame-to-frame "content difference", we then temporally segment the video into sequences which display the same slide. Since this differencing removes low-level imaging effects, we can then relate the video segments to slides by matching the content differences of adjacent video segments to the content differences of all possible slide pairs. By defining the transition probabilities of video segments, we are able to solve this matching efficiently in two steps, first by finding high likelihood matches, and then by using dynamic programming for the unmatched remainder. Once matched, the correspondence to the original slides can be used to enhance the presentation video quality in the content area, and can also be used for indexing and summarization. Experiments show high performance of segmentation and matching on several presentation videos which vary considerably in color and background style.
Tiecheng Liu, Rune Hjelsvold, John R. Kender
ICME (1)3
2002 An efficient error-minimizing algorithm for variable-rate temporal video sampling
abstract
We provide a novel algorithm for selecting key frames from a video at all possible temporal sampling densities. We define a measure for video reconstruction error (VRE), and use VRE to evaluate the representational quality of the selected key frames. The algorithm uses a heap-based greedy algorithm to build a hierarchy of increasingly sparsely sampled temporal sequences, each consisting of key frames with minimal VRE. By exploiting the heap and a novel computation technique of "forward computing", the complete hierarchy is constructed in O(nlog(n)) time and O(n) space. The algorithm also ranks all frames according to their importance in recovering the original video, potentially useful for applications in temporally scalable video coding and video streaming. Experiments show that our algorithm outperforms other existing methods in VRE, as compared via peak signal-noise ratio, in computation time and in guaranteed convergence.
Tiecheng Liu, John R. Kender
ICME (1)2
2002 Multiple feature temporal models for object detection in video
abstract
We present a general object detection system for video sequences based on the characterization of the temporal behavior of multiple image features, as well as any dependency relationships that may exist between them. This characterization is based on the coupling of multiple Markov chains and its generalization to coupled Markov random fields. The model automatically learns the temporal behavior of the features that better characterize an object from a training sequence. We show experiments on generic news videos with cluttered backgrounds and we propose a multi-scale process with automatic selection of the optimal scale of the object in the scene.
Juan María Sánchez, Xavier Binefa, John R. Kender
ICME (1)3
2001 Time-constrained Dynamic Semantic Compression for Video Indexing and Interactive Searching
abstract
We define and present a new model for semantic video compression and searching, and apply this model to several video genres, with special emphasis on instructional videos. In this model, a semantic distance based on the video genre is computed between adjacent frames in a dynamic video buffer of predetermined size, and redundant "unkey" frames leak from the buffer interior into a data structure "leakage history directed acyclic graph" (LHDAG) which records the relative importance of these frames. What exits the buffer forms a highly compressed video stream consisting of only the most semantically significant video frames, whereas the LHDAG permits efficient semantic exploration of the video interior. This novel hierarchical but context-sensitive data structure LHDAG permits the searching of the video at display rates that are proportional to visual significance, at levels of semantic density selectable by the user. The data structure is simple to create and query, and appears to be more psychologically plausible than more straightforward fixed sampling indexing schemes. We empirically display and mathematically analyze the relationships of the video buffer's leaking rate, buffer size, and exit delay, and demonstrate its performance on several extended videos. The flexibility of the method is indicated by its demonstration on two very different definitions of semantic significance: color similarity, and content ("ink pixel") similarity.
Tiecheng Liu, John R. Kender
CVPR (2)2
2001 A unified memory based approach to cut, dissolve, key frame and scene analysis
abstract
We review a memory-based buffer model of visual perception, that combines the lower and middle stages in the analysis of video. This model was originally developed for the detection of breaks between physical scene changes ("story units"). We show how this method can also be applied for shot detection. Moreover, it gives rise to a unified approach for detecting cuts and gradual changes. Additionally, as a straightforward corollary to its definition, it leads to a more natural definition of key frames. We derive the theoretic performance of the model, given arbitrary measures of frame-to-frame dissimilarity. We show several examples of its response, using standard color histogram differencing (with the L/sub 1/ norm as measure). We evaluate the model's performance in detecting cuts and dissolves against a complete hand-segmented situation comedy.
Aya Aner, John R. Kender
ICIP (3)2
2001 Visual Interfaces to Computers: A Systems-Oriented First Course in Reliable Control via Imagery ("Visual Interfaces")
abstract
We present the rationale, description, and critique of a first course in image computing that is not a traditional computer vision principles-and-tools course. "Visual Interfaces to Computers" is instead complementary to standard Computer Vision, User Interface, and Graphics courses; in fact, VI:CV::UI:G. It is organized by case studies of working visual systems that use camera input for data or control information in service of higher user goals, such as GUI control, user identification, or automobile steering. Many CV scientific principles and engineering tools are therefore taught, as well as those of psychophysics, AI, and EE, but taught selectively and always within the context of total system design. Course content is derived from conference and journal articles and Ph.D. theses, augmented with video tapes and real-time web site demos. Students do two homework assignments, one to design a "visual combination lock", and one to parse an image into English. They also do a final paper or project of their own choosing, often in teams of two, and often with surprisingly deep results. The course is assisted by a custom C-based tool kit, "XILite", a user-friendly (and comparatively bug-free) modification of Sun's X-windows Image Library for our lab's camera-equipped Sun workstations. The course has been offered twice to a wide audience with good reviews.
John R. Kender
Int. J. Pattern Recognit. Artif. Intell.1
2001 Sectored Snakes: Evaluating Learned-Energy Segmentations
abstract
We describe how to teach deformable models to maximize image segmentation correctness based on user-specified criteria, and present a method for evaluating which criteria work best. We show how to evaluate the efficacy of any resulting deformable model, given a sampling of ground truth, a model of the range of shapes tried during optimization, and a measure of shape closeness. In the domain of abdominal CT images, we demonstrate such evaluation on a simple "sectoring" of a snake in which intensity and perpendicular gradient are observed over equal-length segments. This specific set of qualities shows a measured improvement over an objective function that is uniform around the shape, and it follows naturally from examination of the latter's failures due to image variations around the organ boundary.
Samuel D. Fenster, John R. Kender
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 A Comparative Technique and Performance Results on Novel Learned Snakes in Two Dissimilar Medical Domains
abstract
We review our work on how to teach deformable models to maximize image segmentation correctness based on user-specified criteria. We then present new variants and applications of learned snakes, modeled by four different probability density functions (PDFs), at three scales, and in the two medical domains of abdominal CT slices and echocardiograms. We review and extend our method for evaluating which criteria work best. Success depends on the relation of objective function (the PDF) output to shape correctness. This relationship for all the above learned snake variants and domains, is evaluated on perturbed ground truth shapes in three ways: by the incidence of "false positives" of randomized shapes; by the monotonicity of the objective function versus shape closeness to ground truth, as given by a correlation coefficient; and by the distance of this relationship to the nearest monotonically increasing function, a new performance measure which we introduce. We demonstrate such evaluations on traditional snakes, and on snakes for which image intensity and perpendicular gradient are learned separately, and with their covariances, and with separate learning over equal-length "sectors". Optimal blur appears to depend on domain. Both sectoring and the use of covariance markedly improve results in abdominal CT images, where nearby image landmarks (i.e. organs) stabilize learning. Results on echocardiograms, however, are less striking, although the use of covariance does show improvements; this appears to be due to the non-Gaussian distribution of image features in this domain.
Samuel D. Fenster, John R. Kender
CVPR2
1998 Video Scene Segmentation via Continuous Video Coherence
abstract
In extended video sequences, individual frames are grouped into shots which are defined as a sequence taken by a single camera, and related shots are grouped into scenes which are defined as a single dramatic event taken by a small number of related cameras. This hierarchical structure is deliberately constructed, dictated by the limitations and preferences of the human visual and memory systems. We present three novel high-level segmentation results derived from these considerations, some of which are analogous to those involved in the perception of the structure of music. First and primarily, we derive and demonstrate a method for measuring probable scene boundaries, by calculating a short term memory-based model of shot-to-shot "coherence". The detection of local minima in this continuous measure permits robust and flexible segmentation of the video into scenes, without the necessity for first aggregating shots into clusters. Second, and independently of the first, we then derive and demonstrate a one-pass on-the-fly shot clustering algorithm. Third, we demonstrate partially successful results on the application of these two new methods to the next higher, "theme", level of video structure.
John R. Kender, Boon-Lock Yeo
CVPR1
1998 Sectored Snakes: Evaluating Learned-Energy Segmentations
abstract
We describe how to teach deformable models to maximize image segmentation correctness based on user-specified criteria, and we present a method for evaluating which criteria work best. We present sectored snakes, a formulation that demonstrably improves upon regular snakes. A traditional deformable model ("snake" in 2D) fails to find an object's boundary when the strongest nearby image edges are not the ones sought. But models can be trained to respond to other image features instead, by learning their probability distributions. The implementor must then decide on which of many image qualities to teach the model. To this end, we show how to evaluate the efficacy of any resulting deformable model, given a sampling of ground truth, a model of the range of shapes tried during optimization, and a measure of shape closeness. In the domain of abdominal CT images, we demonstrate such evaluation on a simple "sectoring" of a snake, in which intensity and perpendicular gradient are observed over equal-length segments. This specific set of qualities shows a measured improvement over an objective function that is uniform around the shape, and it follows naturally from examination of the latter's failures due to images variations around the organ boundary.
Samuel D. Fenster, John R. Kender
ICCV2
1997 Automatic Summarization of Radiographic Imagery
Alicia Abella, John R. Kender
COSIT2
1997 Interaction with On-Screen Objects using Visual Gesture Recognition
abstract
This paper will review the design of a working system that visually recognizes hand gestures for the control of a window based user interface. After an overview of the system, it will explore one aspect of gestural interaction in depth, hand tracking, and what is needed for the user to be able to interact comfortably with on-screen objects. We describe how the location of the hand is mapped to a location on the screen, and how it is both necessary and possible to smooth the camera input using a non-linear physical model of the cursor. The performance of the system is examined, especially with respect to abject selection. We show how a standard HCI model of object selection (Fitts' Law) can be extended to model the selection performance of free-hand pointing.
Rick Kjeldsen, John R. Kender
CVPR2
1996 Toward the use of gesture in traditional user interfaces
abstract
This work describes the design of a functioning user interface based on visual recognition of hand gestures, and details its performance. In the interface, gesture replaces the mouse for many actions including selecting, moving and resizing windows. A camera below the screen observes the user. The hand is segmented from the background using color. Features of the hand's motion are extracted from the sequence of segmented images, and when needed the hand's pose is classified using a neural net. This information is parsed by a task specific grammar. The system runs in real time on standard PC hardware. It has demonstrated its abilities with various users in several different office environments. Having experimented with a functioning gestural interface, the authors discuss the practicality and best applications of this technology.
Rick Kjeldsen, John R. Kender
FG2
1996 Finding skin in color images
abstract
This paper describes the techniques used to separate the hand from a cluttered background in a gesture recognition system. Target colors are identified using a histogram-like structure called a Color Predicate, which is trained in real-time using a novel algorithm. Running on standard PC hardware, the segmentation is of sufficient speed and quality to support an interactive user interface. The method has shown its flexibility in a range of different office environments, segmenting users with many different skin-tones. Variations have been applied to other problems including finding face candidates in video sequences.
Rick Kjeldsen, John R. Kender
FG2
1995 Topological Direction-Giving and Visual Navigation in Large Environments
abstract
In this paper, we propose and investigate a new model for robot navigation in large unstructured environments. Current models, which depend on metric information, have to deal with inherent mechanical and sensory errors. Instead we supply the navigator with qualitative information. Our model consists of two parts, a map-maker and a navigator. Given a source and a goal, the mapmaker derives a navigational path based on the topological relationships between landmarks. A navigational path is generated as a combination of “parkway” and “trajectory” paths, both of which are abstractions of the real world into topological data structures. Traversing within a parkway enables the navigator to follow landmarks that are continuously visible. Traversing on a trajectory enables the navigator to move reliably into featureless space, based on local headings formed by visible landmarks that are robust to positional and orientational errors. Reliability measures of parkway and trajectory traversals are defined by appropriate error models that account for the sensory errors of the navigator, the population of neighboring objects, and the rotational and translational errors of the navigator. The optimal path is further abstracted into a “custom map”, which consists of a list of symbolic directional instructions, the vocabulary of which is defined by our environmental description language. Based on the custom map generated by the map-maker, the navigating robot looks for events that are characterized by spatial properties of the environment. The map-maker and the navigator are implemented using two cameras, an IBM 7575 robot arm, and a PIPE (Pipelined Image Processing Engine.)
Il-Pyung Park, John R. Kender
Artif. Intell.2
1995 On Seeing Spaghetti: Self-Adjusting Piecewise Toroidal Recognition of Flexible Extruded Objects
abstract
We present a model for flexible extruded objects, such as wires, tubes, or grommets, and demonstrate a novel, self-adjusting, seven-dimensional Hough transform that derives their diameter and three-space curved axes from position and surface normal information. The method is purely local and is inexpensive to compute. The model considers such objects as piecewise toroidal, and decomposes the seven parameters of a torus into three nested subspaces, the structures of which counteract the errors implicit in the analysis of objects of great size and/or small curvature. We believe it is the first example of a parameter space structure designed to cluster ill-conditioned hypotheses together so that they can be easily detected and ignored. This work complements existing shape-from-contour approaches for analyzing tori: it uses no edge information, and it does not require the solution of high-degree nonlinear equations by iterative techniques.>
John R. Kender, Rick Kjeldsen
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Qualitatively Describing Objects Using Spatial Prepositions
Alicia Abella, John R. Kender
AAAI2
1993 Using isolated landmarks and trajectories in robot navigation
abstract
A purely topological method is explored for navigation in a large unstructured environment that contains featureless objects, using qualitative nonmetric information such as isolated landmarks and trajectories, which are defined. Given a source and a goal, the map-maker constructs a series of verbal directions, called custom map. Based on the custom map, the navigating robot is told to look for events that are characterized by spatial properties of the environment, such as isolated landmarks or distinguishing juxtapositions of landmarks. The result is efficient and reliable navigation that does not depend on detailed numeric information.>
Il-Pyung Park, John R. Kender
CVPR2
1991 On Seeing Spaghetti: A Novel Self-Adjusting Seven Parameter Hough Space for Analyzing Flexible Extruded Objects
John R. Kender, Rick Kjeldsen
IJCAI1
1989 Why direction-giving is hard: the complexity of using landmarks in one-dimensional navigation
abstract
The authors outline a formal model of topological navigation in one-dimensional spaces such as single roads. They formally define the concepts of direction giving and of custom map, as well as some specifications for feature detector models, including the idea of sensor synchronicity. The representations necessary to model and make use of differences between the world itself, the world as perceived by the map maker, and the world as experienced by the navigator are discussed. The difficulty of giving precise meaning to what is meant by a 'good' map is shown, and an operational definition of what is meant by a 'landmark' is given: regardless of starting position, any custom map can attain a landmark with but a single direction. It is shown that even in the simplest case, NP-complete problems arise in the efficient selection and sequencing of sensor modalities, even while attempting to navigate from one single object to another. A heuristic is provided that appears reasonable for map creation, and examples are given of several very different maps, each of which is optimal under eight reasonable criteria.>
John R. Kender, Avraham Leff
IEEE Trans. Syst. Man Cybern.1
1988 Solving the depth interpolation problem on a parallel architecture with a multigrid approach
abstract
The authors discuss solving the depth interpolation problem on a parallel architecture, a fine-grained SIMD (single-instruction, multiple data stream) machine with local and global communication networks. Many constraint propagation problems in early vision, including depth interpolation, can be cast as solving a large system of linear equations where the resulting matrix is symmetric and positive definite (SPD). Usually, the resulting SPD matrix is sparse. The authors show how the adaptive Chebyshev acceleration and the conjugate gradient methods accelerated further with a multigrid approach can be run on this parallel architecture for sparse SPD matrices. They give numerical results for fairly large synthetic images, and compare them with the results from the Gauss-Seidel method accelerated also with a multigrid approach.>
Dong J. Choi, John R. Kender
CVPR2
1988 An optimal algorithm for the derivation of shape from shadows
abstract
The authors study the problem of recovering a surface slice from the shadows it casts on itself when lighted by the sun at various times of the day. The problem is formulated and solved in a Hilbert space setting. The spine algorithm interpolating the data that result from the shadows is constructed. This algorithm is optimal in terms of the approximation error and has low cost. The authors implement the optimal error algorithm and show a series of test runs. Another modified version of the algorithm that improves the cost considerably is also shown. This version is suited for parallel computation with further reductions in the cost of the solution.>
Michael Hatzitheodorou, John R. Kender
CVPR2
1987 An Integrated System that Unifies Multiple Shape from Texture Algorithms
Mark L. Moerdler, John R. Kender
AAAI2
1987 What is a 'Degenerate' View?
John R. Kender, David G. Freudenstein
IJCAI1
1987 Low-Level Image Analysis Tasks on Fine-Grained Tree-Structured SIMD Machines
abstract
Abstract This paper examines the applicability of fine-grained tree-structured SIMD machines, which are amenable to highly efficient VLSI implementation, to several low-level image understanding tasks. Algorithms are presented for histogramming, thresholding, image correlation, connected component labeling, and computing Euler number. A particular massively parallel machine called NON-VON is used for purposes of explication and performance evaluation. Only NON-VON tree-structured communication capabilities and its SIMD mode of execution are considered in this paper. Novel algorithmic techniques are described, such as vertical pipelining, subproblem partitioning, associative matching, and data duplication, that effectively exploit the massive parallelism available in fine-grained SIMD tree machines while avoiding communication bottlenecks. Simulation results are presented and compared with results obtained or forecast for other highly parallel machines. The relative advantages and limitations of the class of machines under consideration are outlined; except for some types of image correlation, the fine-grained SIMD tree is exceptionally fast.
Hussein Ibrahim, John R. Kender, David E. Shaw
J. Parallel Distributed Comput.2
1986 SIMD Tree Algorithms for Image Correlation
Hussein Ibrahim, John R. Kender, David E. Shaw
AAAI2
1986 Shape from Darkness: Deriving Surface Information from Dynamic Shadows
John R. Kender, Earl Smith
AAAI1
1986 On the application of massively parallel SIMD tree machines to certain intermediate-level vision tasks
abstract
In this paper, we examine the implementation of two middle-level image understanding tasks on fine-grained tree-structured SIMD machines, which have highly efficient VLSI implementations. We first present one such massively parallel machine called NON-VON, and summarize the cost/performance trade-offs of such machines for vision taks. We follow with a more detailed description of the NON-VON architecture (a prototype of which has been operational since January 1985), and of the high-level parallel language in which our algorithms have been written and simulated. The heart of the paper consists of the description and analysis of algorithms for a representative Hough transform, and of an algorithm for the interpretation of moving light displays. Novel algorithmic techniques are motivated and described, and simulation timings are presented and discussed. We conclude that it is possible to exploit the available massive parallelism while avoiding many of the communication bottlenecks common at this level of image understanding, by carefully and inexpensively duplicating data and/or control information, and by delaying or avoiding the reporting of intermediate results.
Hussein A. H. Ibrahim, John R. Kender, David E. Shaw
Comput. Vis. Graph. Image Process.2
1986 Vision expert systems demand challenging expert interactions
abstract
The paper explores the usefulness and applicability of knowledge-based image interpretation. By limiting the analysis to ‘restricted’ scenes, a bottom-up strategy has been developed to improve a primal image segmentation. Two case studies are discussed: the first deals with medical X-rays, the second with satellite images (SPOT). In both projects, generic geometrical knowledge is encoded in the format of production rules. The results obtained so far are encouraging and are already of practical use. Ways to extend the knowledge bases by more specific domain knowledge are mentioned.
John R. Kender
Comput. Vis. Graph. Image Process.1
1983 Surface Constraints From Linear Extents
John R. Kender
AAAI1
1983 Environmental Labelings in Low-Level Image Understanding
John R. Kender
IJCAI1
1983 Gradient space under orthography and perspective
Steven A. Shafer, Takeo Kanade, John R. Kender
Comput. Vis. Graph. Image Process.3
1982 Why Perspective Is Difficult: How Two Algorithms Fail
John R. Kender
AAAI1
1980 Mapping Image Properties into Shape Constraints: Skewed Symmetry and Affine-Tramsfornable Patterns, and the Shape-from-Texture Paradigm
John R. Kender, Takeo Kanade
AAAI1
1979 Shape from Texture: An Aggregation Transform that Maps a Class of Textures into Surface Orientation
John R. Kender
IJCAI1