Alexander C. Loui

dblp:l/AlexanderCLoui · DBLP profile ↗
← Back
61ranked-venue papers
10as first author
4since 2021 · last 2025
0000-0002-7427-1503ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 5Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Segmentation and scene understanding · 46% Deep learning architectures and training · 19% Image recognition and object detection · 16%
Computer graphics and multimedia
10 papers
Multimedia analysis and retrieval · 52% Image and video coding · 35% Visual content generation and editing · 8%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › regularization
consistency training
0.912025
Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation · IEEE Trans. Image Process. 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.912025
Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation · IEEE Trans. Image Process. 2025
Computer vision › Segmentation and scene understanding › medical image segmentation
semi-supervised segmentation
0.912025
Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation · IEEE Trans. Image Process. 2025
Computer vision › Image recognition and object detection › object detection
feature pyramid network
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Machine learning › Graph learning
graph kernel
0.712023
Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence Learning · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding
instance segmentation
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Computer vision › Image recognition and object detection
object detection
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Tripartite Feature Enhanced Pyramid Network for Dense Prediction · IEEE Trans. Image Process. 2023
Image and video coding › image quality assessment
image aesthetics assessment
0.312018
Leveraging Expert Feature Knowledge for Predicting Image Aesthetics · IEEE Trans. Image Process. 2018
Multimedia analysis and retrieval › multimedia analysis › multimedia content description › multimedia semantics
semantic concept detection
0.222009
Short-term audio-visual atoms for generic video concept classification · ACM Multimedia 2009
Kernel Sharing With Joint Boosting For Multi-Class Concept Detection · CVPR 2007
Web and social media mining
online community analysis
0.112010
Kodak moments and Flickr diamonds: how users shape large-scale media · ACM Multimedia 2010
Web and social media mining › social media analysis
social media content analysis
0.112010
Kodak moments and Flickr diamonds: how users shape large-scale media · ACM Multimedia 2010
Visual content generation and editing
image cropping
0.112010
Towards aesthetics: a photo quality assessment and photo selection system · ACM Multimedia 2010
Image and video coding › image quality assessment
photo quality assessment
0.112010
Towards aesthetics: a photo quality assessment and photo selection system · ACM Multimedia 2010
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis › image collection analysis
photo selection
0.112010
Towards aesthetics: a photo quality assessment and photo selection system · ACM Multimedia 2010
Machine learning › Representation and self-supervised learning › visual representation › image representation
convolutional neural network features
0.112018
Leveraging Expert Feature Knowledge for Predicting Image Aesthetics · IEEE Trans. Image Process. 2018
Multimedia analysis and retrieval
audio-visual analysis
0.112009
Short-term audio-visual atoms for generic video concept classification · ACM Multimedia 2009
Multimedia analysis and retrieval › event detection
video event detection
0.112009
Short-term audio-visual atoms for generic video concept classification · ACM Multimedia 2009
Natural language and speech › Information extraction and text analysis › text classification
semantic classification
0.112008
Semantic Concept Classification by Joint Semi-supervised Learning of Feature Subspaces and Support Vector Machines · ECCV (4) 2008
Machine learning › Learning paradigms
semi-supervised learning
0.112008
Semantic Concept Classification by Joint Semi-supervised Learning of Feature Subspaces and Support Vector Machines · ECCV (4) 2008
Visualization and visual analytics › clustering
event clustering
0.012003
Automated event clustering and quality screening of consumer pictures for digital albuming · IEEE Trans. Multim. 2003
Image and video coding
image quality assessment
0.012003
Automated event clustering and quality screening of consumer pictures for digital albuming · IEEE Trans. Multim. 2003
Multimedia analysis and retrieval › audio-visual learning
audio-visual representation
0.012011
Audio-visual grouplet: temporal audio-visual interactions for general video concept classification · ACM Multimedia 2011
Multimedia analysis and retrieval
near-duplicate detection
0.012002
Duplicate detection in consumer photography and news video · ACM Multimedia 2002
Data mining › temporal data mining
event segmentation
0.012001
Using Event Segmentation to Improve Indexing of Consumer Photographs · SIGIR 2001
Multimedia analysis and retrieval › image analysis
image grouping
0.011999
A software system for automatic albuming of consumer pictures · ACM Multimedia (2) 1999
Multimedia analysis and retrieval › multimedia archives
photo collection management
0.012002
Duplicate detection in consumer photography and news video · ACM Multimedia 2002
Multimedia analysis and retrieval › multimedia analysis
visual content analysis
0.012001
Using Event Segmentation to Improve Indexing of Consumer Photographs · SIGIR 2001

Methods — techniques the papers use, named apart from their topics

weak-to-strong consistency · 0.9uncertainty estimation · 0.9copy-paste augmentation · 0.9self-supervised learning · 0.7hand-crafted features · 0.7feature reference module · 0.7feature fusion · 0.7feature feedback · 0.7feature calibration · 0.7contrastive learning · 0.7tree-based feature elimination · 0.3temporal correlation mining · 0.1dictionary learning · 0.1audio-visual grouplet · 0.1joint probabilistic topic model · 0.1cropping-based editing · 0.1aesthetic quality assessment · 0.1multiple instance learning · 0.1
YearPublicationVenuePosition
2025 Rethinking Copy-Paste for Consistency Learning in Medical Image Segmentation
abstract
Semi-supervised learning based on consistency learning offers significant promise for enhancing medical image segmentation. Current approaches use copy-paste as an effective data perturbation technique to facilitate weak-to-strong consistency learning. However, these techniques often lead to a decrease in the accuracy of synthetic labels corresponding to the synthetic data and introduce excessive perturbations to the distribution of the training data. Such over-perturbation causes the data distribution to stray from its true distribution, thereby impairing the model's generalization capabilities as it learns the decision boundaries. We propose a weak-to-strong consistency learning framework that integrally addresses these issues with two primary designs: 1) it emphasizes the use of highly reliable data to enhance the quality of labels in synthetic datasets through cross-copy-pasting between labeled and unlabeled datasets; 2) it employs uncertainty estimation and foreground region constraints to meticulously filter the regions for copy-pasting, thus the copy-paste technique implemented introduces a beneficial perturbation to the training data distribution. Our framework expands the copy-paste method by addressing its inherent limitations, and amplifying the potential of data perturbations for consistency learning. We extensively validated our model using six publicly available medical image segmentation datasets across different diagnostic tasks, including the segmentation of cardiac structures, prostate structures, brain structures, skin lesions, and gastrointestinal polyps. The results demonstrate that our method significantly outperforms state-of-the-art models. For instance, on the PROMISE12 dataset for the prostate structure segmentation task, using only 10% labeled data, our method achieves a 15.31% higher Dice score compared to the baseline models. Our experimental code will be made publicly available at https://github.com/slhuang24/RCP4CL.
Senlong Huang, Yongxin Ge, Dongfang Liu, Mingjian Hong, Junhan Zhao, Alexander C. Loui
IEEE Trans. Image Process.6
2024 Guest Editorial Introduction to the Special Issue on Label-Efficient Learning on Video Data
abstract
Currently, the success of image processing relies heavily on large well-annotated datasets. However, collecting and labeling video data are significantly more labor-intensive, posing major challenges for training video algorithms and limiting their practical applications. While label-efficient techniques for image data have advanced, solutions for video data are still emerging. Unlabeled video data, with their inherent structured nature, offer valuable assets for label-efficient learning. Unlike image data, video data naturally captures realistic transformations, providing rich samples for learning. Moreover, from a border perspective, video tasks hold great potential for applications like autonomous driving and video surveillance but present unique challenges due to the need to understand both spatial and temporal aspects. Leveraging label-efficient learning is essential for comprehensively understanding visual content and enabling a wide range of real-world video applications. This Special Issue on “Label-Efficient Learning for Video Data” seeks to advance research in this area, offering new insights and solutions to benefit both researchers and practitioners.
Wenguan Wang, Tianfei Zhou, Dongfang Liu, Zheng Thomas Tang, Alexander C. Loui
IEEE Trans. Circuits Syst. Video Technol.5
2023 Tripartite Feature Enhanced Pyramid Network for Dense Prediction
abstract
Learning pyramidal feature representations is important for many dense prediction tasks (e.g., object detection, semantic segmentation) that demand multi-scale visual understanding. Feature Pyramid Network (FPN) is a well-known architecture for multi-scale feature learning, however, intrinsic weaknesses in feature extraction and fusion impede the production of informative features. This work addresses the weaknesses of FPN through a novel tripartite feature enhanced pyramid network (TFPN), with three distinct and effective designs. First, we develop a feature reference module with lateral connections to adaptively extract bottom-up features with richer details for feature pyramid construction. Second, we design a feature calibration module between adjacent layers that calibrates the upsampled features to be spatially aligned, allowing for feature fusion with accurate correspondences. Third, we introduce a feature feedback module in FPN, which creates a communication channel from the feature pyramid back to the bottom-up backbone and doubles the encoding capacity, enabling the entire architecture to generate incrementally more powerful representations. The TFPN is extensively evaluated over four popular dense prediction tasks, i.e., object detection, instance segmentation, panoptic segmentation, and semantic segmentation. The results demonstrate that TFPN consistently and significantly outperforms the vanilla FPN. Our code is available at https://github.com/jamesliang819.
Dongfang Liu, James Liang, Tong Geng, Alexander C. Loui, Tianfei Zhou
IEEE Trans. Image Process.4
2023 Reformulating Graph Kernels for Self-Supervised Space-Time Correspondence Learning
abstract
Self-supervised space-time correspondence learning utilizing unlabeled videos holds great potential in computer vision. Most existing methods rely on contrastive learning with mining negative samples or adapting reconstruction from the image domain, which requires dense affinity across multiple frames or optical flow constraints. Moreover, video correspondence prediction models need to uncover more inherent properties of the video, such as structural information. In this work, we propose HiGraph+, a sophisticated space-time correspondence framework based on learnable graph kernels. By treating videos as a spatial-temporal graph, the learning objective of HiGraph+ is issued in a self-supervised manner, predicting the unobserved hidden graph via graph kernel methods. First, we learn the structural consistency of sub-graphs in graph-level correspondence learning. Furthermore, we introduce a spatio-temporal hidden graph loss through contrastive learning that facilitates learning temporal coherence across frames of sub-graphs and spatial diversity within the same frame. Therefore, we can predict long-term correspondences and drive the hidden graph to acquire distinct local structural representations. Then, we learn a refined representation across frames on the node-level via a dense graph kernel. The structural and temporal consistency of the graph forms the self-supervision of model training. HiGraph+ achieves excellent performance and demonstrates robustness in benchmark tests involving object, semantic part, keypoint, and instance labeling propagation tasks. Our algorithm implementations have been made publicly available at https://github.com/zyqin19/HiGraph.
Zheyun Qin, Xiankai Lu, Dongfang Liu, Xiushan Nie, Yilong Yin, Jianbing Shen, Alexander C. Loui
IEEE Trans. Image Process.7
2018 Leveraging Expert Feature Knowledge for Predicting Image Aesthetics
abstract
The ability to rank images based on their appearance finds many real-world applications such as image retrieval or image album creation. Despite the recent dominance of deep learning methods in computer vision which often result in superior performance, they are not always the methods of choice because they lack interpretability. In this work, we investigate the possibility of improving image aesthetic inference of convolutional neural networks with hand-designed features that rely on domain expertise in various fields. We perform a comparison of hand-crafted feature sets in their ability to predict fine-grained aesthetics scores on two image aesthetics datasets. We observe that even feature sets published earlier are able to compete with more recently published algorithms and, by combining the algorithms together, one can obtain a significant improvement in predicting image aesthetics. By using a tree-based learner, we perform feature elimination to understand the best performing features overall and across different image categories. Only roughly 15 % and 8 % of the features are needed to achieve full performance in predicting a fine-grained aesthetic score and binary classification respectively. By combining hand-crafted features with meta-features that predict the quality of an image based on CNN features, the model performs better than a baseline VGG16 model. One can, however, achieve more significant improvement in both aesthetics score prediction and binary classification by fusing the hand-crafted features and the penultimate layer activations. Our experiments indicate an improvement up to 2.2 % achieving current state-of-the-art binary classification accuracy on the AVA dataset when the hand-designed features are fused with activation from VGG16 and ResNet50 networks.
Michal Kucer, Alexander C. Loui, David W. Messinger
IEEE Trans. Image Process.2
2016 Event-enabled intelligent asset selection and grouping for photobook creation
Mark D. Wood, Madirakshi Das, Peter O. Stubler, Alexander C. Loui
Image Vis. Comput.4
2015 A Graph-Based Framework for Video Object Segmentation and Extraction in Feature Space
abstract
Video segmentation is the task of grouping pixels in successive video frames into perceptually coherent regions. It is a preliminary step to solve higher level problems such as automated surveillance, object tracking, video summarization, video indexing and retrieval. For consumer videos, video segmentation is a useful tool for extracting relevant and interesting content from such video sequence for further analysis or re-purposing of the visual content. Given an unannotated video sequence captured by either a static or hand-held camera, our graph-based approach first effectively models the data in a high dimensional feature space, which emphasizes the correlation between similar pixels while reducing the inter-class connectivity between different objects. The graph model fuses appearance, spatial, and temporal information to break a volumetric video sequence into semantic spatiotemporal key-segments. By further grouping the key-segments, a binary segmentation is able to extract a moving object of interest from a video sequence based on its unique and distinguishable regional properties. Experiment results show the robustness of our approach, which has achieved comparable or better performance when compared to several unsupervised methods.
Alexander C. Loui
ISM2
2013 Kinship classification by modeling facial feature heredity
abstract
We propose a new, challenging, problem in kinship classification: recognizing the family that a query person belongs to from a set of families. We propose a novel framework for recognizing kinship by modeling this problem as that of reconstructing the query face from a mixture of parts from a set of families. To accomplish this, we reconstruct the query face from a sparse set of samples among the candidate families. Our sparse group reconstruction roughly models the biological process of inheritance: a child inherits genetic material from two parents, and therefore may not appear completely similar to either parent, but is instead a composite of the parents. The family classification is determined based on the reconstruction error for each family. On our newly collected “Family101” dataset, we discover links between familial traits among family members and achieve state-of-the-art family classification performance.
Ruogu Fang, Andrew C. Gallagher, Tsuhan Chen, Alexander C. Loui
ICIP4
2012 Grouplet-Based Distance Metric Learning for Video Concept Detection
abstract
We investigate general concept detection in unconstrained videos. A distance metric learning algorithm is developed to use the information of the group let structure for improved detection. A group let is defined as a set of audio and/or visual code words that are grouped together according to their strong correlations in videos. By using the entire group lets as building elements, concepts can be more robustly detected than using discrete audio or visual code words. Compared with the traditional method of generating aggregated group let-based features for classification, our group let-based distance metric learning approach directly learns distances between data points, which better preserves the group let structure. Specifically, our algorithm uses an iterative quadratic programming formulation where the optimal distance metric can be effectively learned based on the large-margin nearest-neighbor setting. The framework is quite flexible, where various types of distances can be computed using individual group lets, and through the same distance metric learning algorithm the distances computed over individual group lets can be combined for final classification. We extensively evaluate our method over the large-scale Columbia Consumer Video set. Experiments demonstrate that our approach can achieve consistent and significant performance improvements.
Wei Jiang 0001, Alexander C. Loui
ICME2
2012 Tag Cloud++ - Scalable Tag Clouds for Arbitrary Layouts
abstract
Tag-clouds are becoming extremely popular in multimedia community as media of exploration and expression. In this work, we take tag-cloud construction to a new level by allowing a tag-cloud to take any arbitrary shape while preserving some order of tags (here alphabetical). Our method guarantees non-overlap among words and ensures compact representation within specified shape. The experiments on a variety of input set of tags and shapes of the tag-clouds show that the proposed method is promising and has real-time performance. Finally, we show the applicability of our method with an application wherein the tag-clouds specific to places, people, and keywords are constructed and used for digital media selection within a social network domain.
Dhiraj Joshi, Alexander C. Loui
ISM3
2012 Societally connected multimedia across cultures
abstract
The advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity.
Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu
J. Zhejiang Univ. Sci. C9
2012 Introduction to the ICME 2011 Special Issue
abstract
The 14 papers in this special issue are extended versions of papers presented at ICME 2011, held in Barcelona, Spain, on 11-15 July 2011.
Dinei A. F. Florêncio, Sethuraman Panchanathan, Philippe Salembier, Mohamed Hefeeda, Alexander C. Loui, Mrinal Mandal 0001
IEEE Trans. Multim.6
2011 Soundtrack classification by transient events
abstract
We present a method for video classification based on information in the soundtrack. Unlike previous approaches which describe the audio via statistics of mel-frequency cepstral coefficient (MFCC) features calculated on uniformly-spaced frames, we investigate an approach to focusing our representation on audio transients corresponding to sound-track events. These event-related features can reflect the "foreground" of the soundtrack and capture its short-term temporal structure better than conventional frame-based statistics. We evaluate our method on a test set of 1873 YouTube videos labeled with 25 semantic concepts. Retrieval results based on transient features alone are comparable to an MFCC-based system, and fusing the two representations achieves a relative improvement of 7.5% in mean average precision (MAP).
Courtenay V. Cotton, Daniel P. W. Ellis, Alexander C. Loui
ICASSP3
2011 A new video similarity measure model based on video time density function and dynamic programming
abstract
In this paper, we propose a novel video similarity measure model using video time density function (VTDF) and dynamic programming. First, we employ VTDF to describe the density of video activities in time domain by calculating the inter-frame mutual information. Second, a temporal partition solution is applied to divide each video sequence into equi-sized temporal segments. Third, a new VTDF based similarity measure using correlation is calculated to measure the similarity between two temporal segments. Fourth, dynamic programming is then developed to find the optimal non-linear mapping between two video sequences. A new normalized similarity measure function combing both visual characteristics and temporal information together is to evaluate the semantic similarity of two video sequences. Experimental results show that the proposed measurement model is effective to explore the semantic similarity of video sequences.
Xiao-Ping Zhang 0002, Alexander C. Loui
ICASSP3
2011 Key frame extraction from consumer videos using sparse representation
abstract
Key frame extraction algorithms select a subset of the most informative frames from videos. Key frame extraction finds applications in several broad areas of video processing research such as video summarization, creating “chapter titles” in DVDs, video indexing, and prints from video. In this paper, a sparse representation based method to extract key frames from unstructured consumer videos is presented. In the proposed approach, video frames are projected to a low dimensional random feature space and theory from sparse signal representation is used to analyze the spatio-temporal information of the video data and generate key frames. The proposed approach is computationally efficient and does not require shot(s) detection, segmentation, or semantic understanding. A comparison of the results obtained by this method with the ground truth agreed by multiple judges and another approach based on camera operator's intent clearly indicates the feasibility of the proposed approach.
Mrityunjay Kumar, Alexander C. Loui
ICIP2
2011 Automatic consumer video summarization by audio and visual analysis
abstract
Video summarization provides a condensed version of a video stream by analyzing the video content. Automatic summarization of consumer videos is an important tool that facilitates efficient browsing, searching, and album creation in large consumer video collections. This paper studies automatic video summarization in the consumer domain where most previous methods cannot be easily applied due to the challenging issues for content analysis, i.e., consumer videos are captured with uncontrolled conditions such as uneven illumination, clutter, and large camera motion, and with poor-quality soundtrack as a mix of multiple sound sources under severe noise. To pursue reliable summarization, a case study with actual consumer users is conducted, from which a set of consumer-oriented guidelines is obtained. The guidelines reflect the high-level semantic rules, in both visual and audio aspects, which are recognized by consumers as important to produce good video summaries. Following these guidelines, an automatic video summarization algorithm is developed where both visual and audio information are used to generate improved summaries. To the best of our knowledge, this is a first systematic study on automatic summarization of consumer-quality videos. Experimental evaluations from consumer subjects show the effectiveness of our approach.
Wei Jiang 0001, Courtenay V. Cotton, Alexander C. Loui
ICME3
2011 A content-based video fast-forward playback method using video time density function and rate distortion theory
abstract
In this paper, we propose a new video summary method using video time density function (VTDF) and rate distortion theory. The whole system has two main modules, processing and playing. In the processing part, we apply VTDF to describe the temporal dynamics of video data first. A VTDF-based temporal quantization method is then developed to find the best quanta and partition in time domain. The optimal quanta are used to extract the representative video frames. A temporal mean square error (TMSE) is introduced by using rate-distortion theory to evaluate the quantization performance. In the playing module, we develop a video player to only play all sampled frames in its intelligent fast-forward mode. The built video player can allow users to do fast-forward playback based on the semantic video content, which demonstrates the feasibility of proposed method in practice.
Xiao-Ping Zhang 0002, Alexander C. Loui
ICME3
2011 Saliency Detection Using Region-Based Incremental Center-Surround Distance
abstract
A new method to detect salient region(s) in images is proposed in this paper. The proposed approach, which is inspired by object-based visual attention theory, segments the input image into coherent regions and measures region-based center-surround distance (RBCSD), which is a distance between region attributes such as color histograms found in each region and its surrounding region. Furthermore, segmented regions are merged such that the RBCSD of the merged region is greater than the individual RBCSD of the component regions through region-based incremental center surround distance (RBCSD+I) process. Due to this RBCSD+I process, merged regions may contain incoherent color regions, which improves the robustness of the proposed approach. The key advantages of the proposed algorithm are: (1) it provides a salient region with plausible object boundaries, (2) it is robust to color incoherency present in the salient region, and (3) it is computationally efficient. Extensive qualitative and quantitative evaluation of the proposed algorithm on widely used data sets and comparison with the existing saliency detection approaches clearly indicates the feasibility and efficiency of the proposed approach.
Mrityunjay Kumar, Alexander C. Loui
ISM3
2011 Consumer video understanding: a benchmark database and an evaluation of human and machine performance
abstract
Recognizing visual content in unconstrained videos has become a very important problem for many applications. Existing corpora for video analysis lack scale and/or content diversity, and thus limited the needed progress in this critical area. In this paper, we describe and release a new database called CCV, containing 9,317 web videos over 20 semantic categories, including events like "baseball" and "parade", scenes like "beach", and objects like "cat". The database was collected with extra care to ensure relevance to consumer interest and originality of video content without post-editing. Such videos typically have very little textual annotation and thus can benefit from the development of automatic content analysis techniques.
Yu-Gang Jiang 0001, Guangnan Ye, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ICMR5
2011 Audio-visual grouplet: temporal audio-visual interactions for general video concept classification
abstract
We investigate general concept classification in unconstrained videos by joint audio-visual analysis. A novel representation, the Audio-Visual Grouplet (AVG), is extracted by studying the statistical temporal audio-visual interactions. An AVG is defined as a set of audio and visual codewords that are grouped together according to their strong temporal correlations in videos. The AVGs carry unique audio-visual cues to represent the video content, based on which an audio-visual dictionary can be constructed for concept classification. By using the entire AVGs as building elements, the audio-visual dictionary is much more robust than traditional vocabularies that use discrete audio or visual codewords. Specifically, we conduct coarse-level foreground/background separation in both audio and visual channels, and discover four types of AVGs by exploring mixed-and-matched temporal audio-visual correlations among the following factors: visual foreground, visual background, audio foreground, and audio background. All of these types of AVGs provide discriminative audio-visual patterns for classifying various semantic concepts. We extensively evaluate our method over the large-scale Columbia Consumer Video set. Experiments demonstrate that the AVG-based dictionaries can achieve consistent and significant performance improvements compared with other state-of-the-art approaches.
Wei Jiang 0001, Alexander C. Loui
ACM Multimedia2
2010 Searching consumer image collections using web-based concept expansion
abstract
As consumers accumulate more and more personal imagery, searching for specific images has become increasingly difficult. Consumers typically provide little or no annotations, and automated classifiers and concept tagging tools are limited in their scope and vocabulary. This work addresses this sparsity of semantic information by leveraging domain-specific information provided by online photo-sharing communities. Such information enables improved search by allowing user-provided search terms to be expanded into a set of semantically related concepts, using relevant semantic relationships provided by millions of users. Our system first extracts metadata using a modest number of image and event-based semantic classifiers, as well as any meaningful file or folder names. When users pose text-based queries, our system retrieves images from their personal image collections by leveraging Flickr's tag dataset for concept expansion. This approach enables users to search their collections without having to manually annotate their pictures. We compare the retrieval performance of using a Flickr-based concept expander with the performance obtained without concept expansion and with using a WordNet-based concept expander. The results demonstrate that common sense knowledge gleaned from online photo sharing communities can enable meaningful image search on consumer image collections, searches that would be impossible using only the available image metadata.
Mark D. Wood, Alexander C. Loui, Stacie Hibino
CIKM2
2010 Detecting local semantic concepts in environmental sounds using Markov model based clustering
abstract
Detecting the time of occurrence of an acoustic event (for instance, a cheer) embedded in a longer soundtrack is useful and important for applications such as search and retrieval in consumer video archives. We present a Markov-model based clustering algorithm able to identify and segment consistent sets of temporal frames into regions associated with different ground-truth labels, and simultaneously to exclude a set of uninformative frames shared in common from all clips. The labels are provided at the clip level, so this refinement of the time axis represents a variant of Multiple-Instance Learning (MIL). Evaluation shows that local concepts are effectively detected by this clustering technique based on coarse-scale labels, and that detection performance is significantly better than existing algorithms for classifying real-world consumer recordings.
Keansub Lee, Daniel P. W. Ellis, Alexander C. Loui
ICASSP3
2010 Aesthetic quality assessment of consumer photos with faces
abstract
Automatically assessing the subjective quality of a photo is a challenging area in visual computing. Previous works study the aesthetic quality assessment on a general set of photos regardless of the photo's content and mainly use features extracted from the entire image. In this work, we focus on a specific genre of photos: consumer photos with faces. This group of photos constitutes an important part of consumer photo collections. We first conduct an online study on Mechanical Turk to collect ground-truth and subjective opinions for a database of consumer photos with faces. We then extract technical features, perceptual features, and social relationship features to represent the aesthetic quality of a photo, by focusing on face-related regions. Experiments show that our features perform well for categorizing or predicting the aesthetic quality.
Andrew C. Gallagher, Alexander C. Loui, Tsuhan Chen
ICIP3
2010 A novel framework for fast scene matching in consumer image collections
abstract
The widespread utilization of digital visual media has motivated many research efforts towards efficient search and retrieval from large photo collections. Traditionally, SIFT feature-based methods have been widely used for matching photos taken at particular locations or places of interest. These methods are very time-consuming due to the complexity of the features and the large number of images typically contained in the image database being searched. In this paper, we propose a fast approach to matching images captured at particular locations or places of interest by selecting representative images from an image collection that have the best chance of being successfully matched by using SIFT, and relying on only these representative images for efficient scene matching. We present a unified framework incorporating a set of discriminative features that can effectively select the images containing signature elements of particular locations from a large number of images. The proposed approach produces an order of magnitude improvement in computational time for matching similar scenes in an image collection using SIFT features. The experimental results demonstrate the efficiency of our approach compared to the traditional SIFT, PCA-SIFT, and SURF-based approaches.
Madirakshi Das, Alexander C. Loui
ICME3
2010 Automatic aesthetic value assessment in photographic images
abstract
The automatic assessment of aesthetic values in consumer photographic images is an important issue for content management, organizing and retrieving images, and building digital image albums. This paper explores automatic aesthetic estimation in two different tasks: (1) to estimate fine-granularity aesthetic scores ranging from 0 to 100, a novel regression method, namely Diff-RankBoost, is proposed based on RankBoost and support vector techniques; and (2) to predict coarse-granularity aesthetic categories (e.g., visually “very pleasing” or “not pleasing”), multi-category classifiers are developed. A set of visual features describing various characteristics related to image quality and aesthetic values are used to generate multidimensional feature spaces for aesthetic estimation. Experiments over a consumer photographic image collection with user ground-truth indicate that the proposed algorithms provide promising results for automatic image aesthetic assessment.
Wei Jiang 0001, Alexander C. Loui, Cathleen Daniels Cerosaletti
ICME2
2010 Towards aesthetics: a photo quality assessment and photo selection system
abstract
Automatic photo quality assessment and selection systems are helpful for managing the large mount of consumer photos. In this paper, we present such a system based on evaluating the aesthetic quality of consumer photos. The proposed system focuses on photos with faces, which constitute an important part of consumer photo albums. The system has three contributions: 1) We propose an aesthetics-based photo assessment algorithm, by considering different aesthetics-related factors, including the technical characteristics of the photo and the specific features related to faces; 2) Based on the aesthetic measurement, we propose a cropping-based photo editing algorithm, which differs from prior works by eliminating unimportant faces before optimizing photo composition; 3) We also incorporate the aesthetic evaluation with other metrics to select quintessential photos for a large collection of photos. The entire system is delivered by a web interface, which allows users to submit images or albums, and returns promising results for photo evaluation, editing recommendation, and photo selection.
Alexander C. Loui, Tsuhan Chen
ACM Multimedia2
2010 Kodak moments and Flickr diamonds: how users shape large-scale media
abstract
In today's age of digital multimedia deluge, a clear understanding of the dynamics of online communities is capital. Users have abandoned their role of passive consumers and are now the driving force behind large-scale media repositories, whose dynamics and shaping factors are not yet fully understood. In this paper we present a novel human-centered analysis of two major photo sharing websites, Flickr and Kodak Gallery. On a combined dataset of over 5 million tagged photos, we investigate fundamental differences and similarities at the level of tag usage and propose a joint probabilistic topic model to provide further insight into semantic differences between the two communities. Our results show that the effects of the users' motivations and needs can be strongly observed in this large-scale data, in the form of what we call Kodak Moments and Flickr Diamonds. They are an indication that system designers should carefully take into account the target audience and its needs.
Radu Andrei Negoescu, Alexander C. Loui, Daniel Gatica-Perez
ACM Multimedia2
2010 Audio-visual atoms for generic video concept classification
abstract
We investigate the challenging issue of joint audio-visual analysis of generic videos targeting at concept detection. We extract a novel local representation, Audio-Visual Atom (AVA), which is defined as a region track associated with regional visual features and audio onset features. We develop a hierarchical algorithm to extract visual atoms from generic videos, and locate energy onsets from the corresponding soundtrack by time-frequency analysis. Audio atoms are extracted around energy onsets. Visual and audio atoms form AVAs, based on which discriminative audio-visual codebooks are constructed for concept detection. Experiments over Kodak's consumer benchmark videos confirm the effectiveness of our approach.
Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ACM Trans. Multim. Comput. Commun. Appl.5
2009 Event classification in personal image collections
abstract
In this paper, we investigate event classification that is specifically developed for use in consumer family photo collections. This domain is very different from news video collections that have been the focus of research in the area of scene content classification. We determine a set of broad event classes that are relevant to personal collections. We investigate the use of a variety of high-level visual and temporal features, and determine a set of features that show good correlation with the event class. We propose a Bayesian belief network for event classification that computes the a posteriori probability of the event class given the input features. The Bayes net is trained on a large set of manually annotated consumer collections. We obtain a classification accuracy of over 70% in this challenging domain.
Madirakshi Das, Alexander C. Loui
ICME2
2009 Short-term audio-visual atoms for generic video concept classification
abstract
We investigate the challenging issue of joint audio-visual analysis of generic videos targeting at semantic concept detection. We propose to extract a novel representation, the Short-term Audio-Visual Atom (S-AVA), for improved concept detection. An S-AVA is defined as a short-term region track associated with regional visual features and background audio features. An effective algorithm, named Short-Term Region tracking with joint Point Tracking and Region Segmentation (STR-PTRS), is developed to extract S-AVAs from generic videos under challenging conditions such as uneven lighting, clutter, occlusions, and complicated motions of both objects and camera. Discriminative audio-visual codebooks are constructed on top of S-AVAs using Multiple Instance Learning. Codebook-based features are generated for semantic concept detection. We extensively evaluate our algorithm over Kodak's consumer benchmark video set from real users. Experimental results confirm significant performance improvements - over 120% MAP gain compared to alternative approaches using static region segmentation without temporal tracking. The joint audio-visual features also outperform visual features alone by an average of 8.5% (in terms of AP) over 21 concepts, with many concepts achieving more than 20%.
Wei Jiang 0001, Courtenay V. Cotton, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ACM Multimedia5
2008 Semantic Concept Classification by Joint Semi-supervised Learning of Feature Subspaces and Support Vector Machines
Wei Jiang 0001, Shih-Fu Chang, Tony Jebara, Alexander C. Loui
ECCV (4)4
2008 Cross-domain learning methods for high-level visual concept classification
abstract
Exploding amounts of multimedia data increasingly require automatic indexing and classification, e.g. training classifiers to produce high-level features, or semantic concepts, chosen to represent image content, like car, person, etc. When changing the applied domain (i.e. from news domain to consumer home videos), the classifiers trained in one domain often perform poorly in the other domain due to changes in feature distributions. Additionally, classifiers trained on the new domain alone may suffer from too few positive training samples. Appropriately adapting data/models from an old domain to help classify data in a new domain is an important issue. In this work, we develop a new cross-domain SVM (CDSVM) algorithm for adapting previously learned support vectors from one domain to help classification in another domain. Better precision is obtained with almost no additional computational cost. Also, we give a comprehensive summary and comparative study of the state-of-the-art SVM-based cross-domain learning methods. Evaluation over the latest large-scale TRECVID benchmark data set shows that our CDSVM method can improve mean average precision over 36 concepts by 7.5%. For further performance gain, we also propose an intuitive selection criterion to determine which cross-domain learning method to use for each concept.
Wei Jiang 0001, Eric Zavesky, Shih-Fu Chang, Alexander C. Loui
ICIP4
2008 Multidimensional image value assessment and rating for automated albuming and retrieval
abstract
The ability to automatically assess image characteristics is an important function for content management, building digital image albums, storytelling with images, and retrieval of specific visual content. This capability is needed to organize and sort large numbers of image and video assets. This paper proposes a novel approach to assess and rate images based on multidimensional characteristics including image quality, social relationships, aesthetic quality, important events, and usage. This new approach provides additional flexibility for end user applications that utilize different aspects of image characteristics. Specifically, we describe a method for assessing image quality based upon technical characteristics of the image, and for predicting the significance of an image based upon the people portrayed in the image. Experimental results indicate that the proposed multidimensional approach provides a promising framework for automated image value assessment and rating.
Alexander C. Loui, Mark D. Wood, Anthony Scalise, John Birkelund
ICIP1
2008 Semantic event detection for consumer photo and video collections
abstract
The automatic detection of semantic events in userspsila image and video collections is an important technique for content management and retrieval. In this paper we propose a novel semantic event detection approach by considering an event-level bag-of-features (BOF) representation to model typical events. Based on this BOF representation, semantic events are detected in a concept space instead of the original low-level visual feature space. There are two advantages of our approach: we can avoid the sensitivity problem by decreasing the influence of difficult or erroneous images or videos in measuring the event-level similarity; also we can utilize the power of higher-level concept scores in describing semantic events. Experiments over a large real consumer database confirm the effectiveness of our approach.
Wei Jiang 0001, Alexander C. Loui
ICME2
2008 Semantics Meets UX: Mediating Intelligent Indexing of Consumers' Multimedia Collections for Multifaceted Visualization and Media Creation
abstract
Unorganized media collections hinder consumers from fully experiencing and enjoying their visual media. User interfaces can mediate the results of automated indexing by presenting data and interactions that leverage the strengths of individual and combined algorithm results, supporting multifaceted browsing, and enabling user correction in a way that is not disruptive to the userspsila activities. We describe the semantic system demonstration framework (SSDF), a flexible and extensible framework for combining multiple semantic indexing algorithms for consumer photo and video clip collections into one integrated system. We also describe key features of Koi, an SSDF desktop client application with a user interface designed to mediate and leverage the intelligent indexing incorporated in the SSDF server. Together, Koi and SSDF empower users to experience their personal multimedia in novel and sophisticated ways.
Stacie Hibino, Alexander C. Loui, Mark D. Wood, Samuel Fryer, Cathleen Daniels Cerosaletti
ISM2
2007 Kernel Sharing With Joint Boosting For Multi-Class Concept Detection
abstract
Object/scene detection by discriminative kernel-based classification has gained great interest due to its promising performance and flexibility. In this paper, unlike traditional approaches that independently build binary classifiers to detect individual concepts, we proposed a new framework for multi-class concept detection based on kernel sharing and joint learning. By sharing "good" kernels among concepts, accuracy of individual weak detectors can be greatly improved; by joint learning of common detectors among classes, the required kernels and the computational complexity for detecting each individual concept can be reduced. We demonstrated our approach by developing an extended JointBoost framework, which was used to choose the optimal kernel and subset of sharing classes in an iterative boosting process. In addition, we constructed multi-resolution visual vocabularies by hierarchical clustering and computed kernels based on spatial matching. We tested our method in detecting 12 concepts (objects, scenes, etc) over 80+ hours of broadcast news videos from the challenging TRECVID 2005 corpus. Significant performance gains were achieved -10% in mean average precision (MAP) and up to 34% average precision (AP) for some concepts like maps, building, and boat-ship. Extensive analysis of the results also revealed interesting and important underlying relations among concepts.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
CVPR3
2007 Context-Based Concept Fusion with Boosted Conditional Random Fields
abstract
The contextual relationships among different semantic concepts provide important information for automatic concept detection in images/videos. We propose a new context-based concept fusion (CBCF) method for semantic concept detection. Our work includes two folds. (1) We model the inter-conceptual relationships by a conditional random field (CRF) that improves detection results from independent detectors by taking into account the inter-correlation among concepts. CRF directly models the posterior probability of concept labels and is more accurate for the discriminative concept detection than previous statistical inferencing techniques. The boosted CRF framework is incorporated to further enhance performance by combining the power of boosting with CRF. (2) We develop an effective criterion to predict which concepts may benefit from CBCF. As reported in previous works, CBCF has inconsistent performance gain on different concepts. With accurate prediction, computational and data resources can be allocated to enhance concepts that are promising to gain performance. Evaluation on TRECVID2005 development set demonstrates the effectiveness of our algorithm.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
ICASSP (1)3
2007 User-Assisted People Search in Consumer Image Collections
abstract
In this paper, we investigate the process of searching for images of specified people in the consumer family photo domain. This domain is very different from the controlled environment of secure-access applications that have been extensively studied and face recognition packages are available in the market. Instead of the typical frontal mug shot, consumer photos are more likely to show people with unconstrained pose and illumination. This domain is also unique in that there are a large number of instances of a limited number of unique individuals. We develop and test facial recognition that is specifically targeted to this domain, using facial features that are derived from active shape modeling of faces, followed by a combination of the features using AdaBoost. We also provide a workflow that is suitable for lay users, and which rewards user inputs with improved performance. Test results show good performance on a challenging data set of consumer images.
Andrew C. Gallagher, Madirakshi Das, Alexander C. Loui
ICME3
2006 Active Context-Based Concept Fusionwith Partial User Labels
abstract
In this paper we propose a new framework, called active context-based concept fusion, for effectively improving the accuracy of semantic concept detection in images and videos. Our approach solicits user annotations for a small number of concepts, which are used to refine the detection of the rest of concepts. In contrast with conventional methods, our approach is active, by using information theoretic criteria to automatically determine the optimal concepts for user annotation. Our experiments over TRECVID 2005 development set (about 80 hours) show significant performance gains. In addition, we have developed an effective method to predict concepts that may benefit from context-based fusion.
Wei Jiang 0001, Shih-Fu Chang, Alexander C. Loui
ICIP3
2005 RhythmPix: a multimedia composition and albuming system for consumer images
abstract
With the advent of new digital consumer electronic devices, the interest and demand for audiovisual media and content authoring has increased in recent years. The RhythmPix project is to explore, develop, and demonstrate technology for albuming and authoring of multimedia image content for playback on consumer electronic devices such as DVD players, as well as PCs. In this paper, we describe an end-to-end multimedia albuming system that we have developed as a platform for combining images and sound/music content. The three main functional modules-media composition, media encoding, custom recording, and related implementation issues, are described. A set of APIs has been developed for rapid system adaptation. Various user studies and focus groups have been conducted in the US and Asia to validate the concept and functionality of the system.
Alexander C. Loui, Bryan Kraus, Jon Riek
ICME1
2005 Classification-based multidimensional adaptation prediction for scalable video coding using subjective quality evaluation
abstract
Scalable video coding offers a flexible representation for video adaptation in multiple dimensions comprising spatial detail and temporal resolution, thus providing great benefits for universal media access (UMA) applications. However, currently most of the approaches address the multidimensional adaptation (MDA) problem in an ad hoc manner. One challenging issue affecting the systematic MDA solution is the difficulty in constructing analytical models in theoretical optimization that capture the relations between video utility and MDA operations. In this paper, we propose a general classification-based prediction framework for selecting the preferred MDA operations based on subjective quality evaluation. For this purpose, we first apply domain-specific knowledge or general unsupervised clustering to construct distinct categories within which the videos share similar preferred MDA operations. Thereafter, a machine learning based method is applied where the low level content features extracted from the compressed video streams are employed to train a framework for the problem of joint signal-to-noise ratio (SNR)-temporal adaptation selection based on the motion compensated three-dimensional subband coding (MC-3DSBC) system. We conduct extensive subjective tests involving 31 subjects, 128 video clips, and formal subjective quality metrics. Statistical analysis of the experimental results confirms the excellent accuracy in using domain knowledge and content features to predict the MDA operation.
Mihaela van der Schaar, Shih-Fu Chang, Alexander C. Loui
IEEE Trans. Circuits Syst. Video Technol.4
2004 Subjective preference of spatio-temporal rate in video adaptation using multi-dimensional scalable coding
abstract
Video adaptation allows for direct manipulation of existing encoded video streams to meet new resource constraints without having to encode the video from scratch. Multi-dimensional scalable coding, such as motion-compensated subband coding (MCSBC), offers an effective and flexible representation for video adaptation. In order to develop robust criteria for selecting optimal spatio-temporal rates used in adaptation, knowledge about subjective preference of spatio-temporal rates is needed. We study the optimal temporal frame rate over a wide range of bandwidth (50 kbps to 1 Mbps) using subjective quality evaluation with 128 clips and 31 subjects. We analyze the results using statistical testing methods and investigate the dependence of optimal frame rate on user, bandwidth, and video content characteristics. Our findings indicate agreement among most users and the existence of switching bandwidths at which preferred frame rates change. Dependence of the preference on video content types is also revealed.
Shih-Fu Chang, Alexander C. Loui
ICME3
2003 High resolution multimedia slide show composition for video CD and DVD rendering
abstract
With the advance of digital and information technology, there is an urgent need to make image taking and sharing easy, fun, and cost and time efficient. In this paper, we present a multimedia system to compose visual and audio information into high resolution multimedia slide show, a video compact disc (VCD) compliant bitstream readable by electronic devices such as video CD/DVD players, for TV-centered multimedia rendering, sharing and entertainment. By using the technologies of multimedia composition, multimedia coding, and the sophisticated bit allocation scheme, we can achieve the same spatial resolution of DVD (without full motion). Special efforts are devoted to the multimedia experience, image quality improvement, and automation in the system.
Zhaohui Sun, Jon Riek, Alexander C. Loui
ICME3
2003 Automatic face-based image grouping for albuming
abstract
In this paper, a system for automatic albuming of consumer photographs is described, that uses face-based information extracted from images. The target image sets for this work are snapshots from family photo collections. The aim is to automatically provide the user with the option of selecting image groups based on the people present in them. The core components of age/gender classification and clustering based on facial similarity for this domain are discussed. Age and gender classification based on individual facial feature measurements combined to produce a single classifier using the AdaBoost algorithm. Face-based information is used both in locating suitable images and layout/design phases of the albuming process. Face-based information is combined with earlier work on event segmentation and layout design to provide a more effective system. Performance is tested on two family photo databases covering a 5 year time-span.
Madirakshi Das, Alexander C. Loui
SMC2
2003 An observation-constrained generative approach for probabilistic classification of image regions
Sanjiv Kumar, Alexander C. Loui, Martial Hebert
Image Vis. Comput.2
2003 Finding structure in home videos by probabilistic hierarchical clustering
abstract
Accessing, organizing, and manipulating home videos present technical challenges due to their unrestricted content and lack of storyline. We present a methodology to discover cluster structure in home videos, which uses video shots as the unit of organization, and is based on two concepts: (1) the development of statistical models of visual similarity, duration, and temporal adjacency of consumer video segments and (2) the reformulation of hierarchical clustering as a sequential binary Bayesian classification process. A Bayesian formulation allows for the incorporation of prior knowledge of the structure of home video and offers the advantages of a principled methodology. Gaussian mixture models are used to represent the class-conditional distributions of intra- and inter-segment visual and temporal features. The models are then used in the probabilistic clustering algorithm, where the merging order is a variation of highest confidence first, and the merging criterion is maximum a posteriori. The algorithm does not need any ad-hoc parameter determination. We present extensive results on a 10-h home-video database with ground truth which thoroughly validate the performance of our methodology with respect to cluster detection, individual shot-cluster labeling, and the effect of prior selection.
Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
2003 Automated event clustering and quality screening of consumer pictures for digital albuming
abstract
In this paper, algorithms for automatic albuming of consumer photographs are described. Specifically, two core algorithms namely event clustering and screening of low-quality images, are introduced and their performance is evaluated. Event clustering and image quality screening have many applications including albuming services, image management and organization, and digital photofinishing. These are difficult tasks because there is, in general, none (or very limited) contextual information about picture content, and the final interpretation could be subjective. A novel event-clustering algorithm is created to automatically segment pictures into events and subevents for albuming, based on date/time metadata information, as well as color content of the pictures. A block-based color histogram correlation technique is developed for image content comparison of general consumer pictures. A new quality-screening algorithm is developed based on object quality measures, to detect problematic images caused by underexposure, low contrast, and camera defocus or movement.
Alexander C. Loui, Andreas E. Savakis
IEEE Trans. Multim.1
2002 Probabilistic home video structuring: feature selection and performance evaluation
abstract
We previously proposed a method to find the cluster structure in home videos based on statistical models of visual and temporal features of video segments and sequential binary Bayesian classification. In this paper, we present analysis and improved results on two key issues: feature selection and performance evaluation, using a ten-hour database (30 video clips, 1,075,000 frames). From multiple features and similarity measures, visual features are selected in order to minimize the empirical probability of misclassification. Temporal features are chosen to reflect the patterns existing in both shot and cluster duration and adjacency. Finally, we describe a detailed performance evaluation procedure that includes cluster detection, individual shot-cluster labeling, and prior selection.
Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun
ICIP (1)2
2002 Event clustering of consumer pictures using foreground/background segmentation
abstract
This paper describes a new algorithm to classify consumer photographs into different events when date and time information is not available. Without any information about the context of the pictures, we have to rely on the image content. Our approach involves using an efficient segmentation scheme and extraction of low-level features to detect event boundaries. Specifically, we have developed a foreground/background segmentation algorithm based on block-based clustering. This block segmentation provides less precision, but still gives good results with low computation cost. A third-party ground truth database has been created with the help of the Human Factors Laboratory at Kodak, to benchmark our approaches. Based on these results, we concluded that a simple block-based segmentation scheme performed better than the original block-based event clustering algorithm without segmentation. We believe that many improvements, especially on segmentation and feature extraction, should lead to better results in the future.
Alexander C. Loui, Matthieu Jeanson
ICME (1)1
2002 Duplicate detection in consumer photography and news video
abstract
Consumers often make more than one photograph of the same scene, creating non-identical duplicates and near duplicates. In Kodak's consumer photography database, on average, 19% of the images, per roll, fall into this category. Automatic detection of duplicates, therefore, is extremely useful in applications that help users organize their image collections. We introduce the challenging problem of non-identical duplicate image detection in consumer photography, describe STELLA (a novel interactive personal image collection organization system), and give an overview of our novel framework for detecting duplicate and near duplicate consumer photographs and news videos.
Alejandro Jaimes, Shih-Fu Chang, Alexander C. Loui
ACM Multimedia3
2001 Consumer Video Structuring by Probabilistic Merging of Video Segments
abstract
Accessing, organizing, and manipulating home videos constitutes a technical challenge due to their unrestricted content and the lack of storyline. In this paper, we present a methodology for structuring consumer video, based on the development of statistical models of similarity and adjacency between video segments in a probabilistic formulation. Learned Gaussian mixture models of inter-segment visual similarity, temporal adjacency, and segment duration are used to represent the classconditional densities of observed features. Such models are then used in a sequential merging algorithm consisting of a binary Bayes classifier, where the merging order is determined by a variation of Highest Confidence First (HCF), and the merging criterion is Maximum a Posteriori (MAP). The merging algorithm can be efficiently implemented and does not need any empirical parameter determination. Finally, the representation of the merging sequence by a tree provides for hierarchical, nonlinear access to the video content. Results on an eight-hour home video database illustrate the validity of our approach.
Daniel Gatica-Perez, Ming-Ting Sun, Alexander C. Loui
ICME3
2001 Using Event Segmentation to Improve Indexing of Consumer Photographs
abstract
Automatic albuming --- the automatic organization of photographs, either as an end in itself or for use in other applications -- is an application that promises to be of great assistance to photographers. Relatively sophisticated image content analysis techniques have been used for image indexing, organization and retrieval. In this paper, we describe a method of organizing photographs into events using spoken photograph captions. The results of this process can be used to improve image indexing and retrieval.
Amanda Stent, Alexander C. Loui
SIGIR2
2000 Discovering Recurrent Visual Semantics in Consumer Photographs
abstract
We present techniques to semi-automatically discover recurrent visual semantics (RVS)-the repetitive appearance of visually similar elements such as objects and scenes-in consumer photographs. First, we introduce the detection of "bracketing" (very similar photographs) using an edge-correlation metric, which outperforms the color histogram. Then, we use color and novel composition features (based on automatic region segmentation) to perform scene-level clustering of images. We use a novel sequence-weighted technique, which uses the structure of standard film (only image sequence information), to perform hierarchical clustering. We show performance results of bracketing, explore clustering evaluation, and discuss STELLA, an interactive albuming and story telling application that uses these techniques to assist users in building digital albums. The STELLA system uses a new approach to album creation: instead of automatically creating albums, it provides an interactive environment that assists users in digital album creation.
Alejandro Jaimes, Ana B. Benitez, Shih-Fu Chang, Alexander C. Loui
ICIP4
1999 A software system for automatic albuming of consumer pictures
abstract
The wide-spread use of image capturing and scanning devices such as digital cameras and low-cost scanners is rapidly resulting in the digital equivalent of the overstuffed shoebox full of snapshots.This paper describes a software system for automatic albuming of pictures to facilitate the creation of consumer albums and album pages.This software system provides consumers an efficient way of organizing and albuming their pictures digitized born different sources.The system includes advanced features such as image event clustering, dud detection, and duplicate detection necessary for automated albuming.The software architecture is based on a simple database structure that can efficiently handle metadata information.The system includes a substantial user interface component that enables the user to preview and edit the generated album.
Alexander C. Loui, Mark D. Wood
ACM Multimedia (2)1
1998 An Image Database for Benchmarking of Automatic Face Detection and Recognition Algorithms
abstract
This paper describes a face image database which has been created and developed at Kodak as a common database for direct benchmarking of automatic face detection and recognition algorithms. This consumer application-oriented face image database is composed of two main sub-databases, one for face detection, and one for face recognition. The database is intended to be distributed to researchers both inside and outside of Kodak working in face detection and recognition research. The database contains pictures taken using consumer digital cameras, scanned in from a photo scanner, as well as pictures from Kodak Image Magic picture disks. The details of the face database such as the development process and the image file formats, are described, together with a discussion on some application scenarios, as well as current benchmarking activities.
Alexander C. Loui, Charles N. Judice
ICIP (1)1
1997 A coded-domain video combiner for multipoint continuous presence video conferencing
abstract
This paper presents a coded-domain combining technique which can accommodate up to six users for continuous presence multipoint video conferencing. This technique is based on a group-of-block (GOB) video unit in the H.261 video standard. Three important technical issues including frame rate synchronization, combiner delay accumulation, and potential quality degradation at the GOB boundary, are addressed. A frame rate synchronization scheme is proposed for the coded-domain combiner. It is observed that delay accumulation will not occur at the combiner with the proposed multiplexing strategy. Further, simulation results indicate that the output perceptual quality is very comparable even without special motion search handling at the GOB boundary. This implies that the proposed scheme can be applied to standard terminals. Finally, we demonstrate the viability of this software-based coded-domain combiner for continuous-presence multipoint video conferencing with a fully functional experimental prototype built around a low-cost personal computer.
Ming-Ting Sun, Alexander C. Loui, Ting-Chung Chen
IEEE Trans. Circuits Syst. Video Technol.2
1993 Application of a multiresolution optical-flow-based method for motion estimation to video coding
Pierre Moulin, Alexander C. Loui
ISCAS2
1993 High-resolution still-image transmission based on CCITT H.261 codec
abstract
An efficient scheme is given for transmitting high-resolution still images based on a lower-resolution CCITT H.261 codec. Two transmission schemes based on a subsampling approach are studied in terms of image quality, transmission delay, and compatibility with H.261. It is found that the sequential frame-repeat scheme provides a good compromise among the different criteria. In addition, a simple progressive-reconstruction procedure for the decoded subimages is described. Simulation results show that the visual effects of such a reconstruction procedure are quite pleasing.>
Alexander C. Loui, Ming Lei Liou
IEEE Trans. Circuits Syst. Video Technol.1
1992 Flexible architectures for morphological image processing and analysis
abstract
An architecture for the efficient and high-speed realization of morphological filters is presented. Since morphological filtering can be described in terms of erosion and dilation, two basic building units performing these functions are required for the realization of any morphological filter. Dual architectures for erosion and dilation are proposed and their operations are described. Their structure, similar to the systolic array architecture as used in the implementation of linear digital filters, is highly modular and suitable for efficient very-large-scale integration (VLSI) implementation. A decomposition scheme is proposed to facilitate the implementation of two-dimensional morphological filters based on one-dimensional structuring elements constructed using the dual architectures. The proposed architectures, which also allow the processing of gray-scale images, are appropriate for applications where speed, size, and cost are of critical significance.>
Alexander C. Loui, Anastasios N. Venetsanopoulos, Kenneth C. Smith
IEEE Trans. Circuits Syst. Video Technol.1
1992 Morphological autocorrelation transform: A new representation and classification scheme for two-dimensional images
abstract
A methodology based on mathematical morphology is proposed for efficient recognition of two-dimensional (2D) objects or shapes. It is based on the introduction a shape descriptor called the morphological autocorrelation transform (MAT). The MAT of an image is composed of a family of geometrical correlation functions (GCFs) which define its morphological covariance in a specific direction. The MAT is translation-, scale-, and rotation-invariant. It is shown that in most situations, a small subset of the MAT suffices for image representation. The characteristics and performance of a shape recognition system based on the MAT are investigated and analyzed. Computational complexity of the proposed morphological-based recognition system is examined. It is shown that shape properties, such as area, perimeter, and orientation, are readily derived from the MAT representation, and that the proposed system is well suited for shape representation and classification.
Alexander C. Loui, Anastasios N. Venetsanopoulos, Kenneth C. Smith
IEEE Trans. Image Process.1
1990 Two-dimensional shape representation using morphological correlation functions
abstract
A new descriptor for representing two-dimensional continuous or discrete signals is introduced. The proposed shape descriptor, which is called the geometrical correlation function (GCF), is based on the principle of mathematical morphology. The properties of this shape descriptor are examined. It is shown that the family of GCFs associated with different orientations of a particular shape is translation, scale, and rotation invariant. Geometrical properties such as the area and perimeter of the shape can be derived from the GCF family. The utilization of the GCFs for shape recognition is considered. The GCF family can be computed using the associated morphological correlator which is composed of m parallel computation units and a feature-function selection unit where a small subset of the GCF family is selected for classification. It is shown that with a suitable criterion for selecting the feature function, promising results for successful classification are obtained.>
Alexander C. Loui, Anastasios N. Venetsanopoulos, Kenneth C. Smith
ICASSP1