EDBT 2026 Demo / reviewers in the wild / expert
James Z. Wang 0001
dblp:w/JamesZeWang · also James Ze Wang
· DBLP profile ↗
114ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0003-4379-4173ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 42 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7Computer networks · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval QueriesabstractVisual content memorability has intrigued the scientific community for decades, with applications spanning from understanding nuanced aspects of human memory to enhancing content design. A significant challenge in progressing the field lies in the high cost of collecting memorability annotations from humans, which constrains both the diversity and scalability of available datasets. Existing datasets typically provide only aggregate memorability scores for visual content, overlooking the nuanced signals embedded in natural, open-ended recall descriptions. In this work, we introduce the first large-scale, unsupervised dataset designed explicitly for modeling visual memorability signals, containing over 82,000 videos paired with descriptive recall data. We leverage tip-of-the-tongue (ToT) retrieval queries from online platforms such as Reddit. We demonstrate that this unsupervised dataset provides rich signals for two memorability-related tasks: recall generation and ToT retrieval. Large vision-language models fine-tuned on our dataset outperform state-of-the-art models such as GPT-4o in generating open-ended memorability descriptions for visual content. In addition, we employ a contrastive training strategy to create the first model capable of multimodal ToT retrieval. Our dataset and models present a new research direction and provide scalable tools for advancing work on visual content memorability.1 Sree Bhattacharyya, Yaman Singla, Sudhir Yarram, Somesh Singh 0003, Harini S. I, James Z. Wang 0001 |
WACV | 6 |
| 2026 | Computational Investigation of Abstraction in Claude Monet's Water Lilies Through Brushstroke AnalysisabstractClaude Monet's late paintings of Water Lilies exhibit stylistic transformations that are often characterized by art historians as increasingly abstract and gesturally expressive. However, it remains challenging to define and systematically identify this stylistic shift. Here, we introduce a machine learning framework for analyzing Monet's evolving brushwork using streamline curves: computational representations that capture the dynamic movement patterns inherent in brushstrokes. From 554 image patches sampled from 47 paintings spanning early (pre-1913) and later (post-1913) periods of Monet's output, we extract streamlines and compute geometric features for each, including smoothness of curvature and directional variability. Each image is represented as a set of streamline feature vectors, a data type referred to as distributional. A new deep neural network architecture named Composition to Attribute (C2A) is designed for classifying distributional data. We hypothesize that Monet's so-called 'abstract' style does not uniformly characterize all late-period Water Lilies, and that non-abstract flowers, regardless of period, share similar brushwork qualities. Under these assumptions, building on C2A, we propose a novel learning paradigm named Discover Embedded Group with Asymmetry (DEGA) which enforces a shared distribution of DNN-extracted features for non-abstract flower patches across both periods while distinguishing the abstract ones. DEGA reveals a meaningful two-dimensional feature space, where one dimension differentiates abstract from mimetic Water Lilies, while the other separates abstract flowers from close-up flowers of the early period. Our findings suggest that the so-called 'abstract' qualities of Monet's late style retain certain visual affinities with his earlier approach to depicting close-up floral motifs. When this brushwork is used in more expansive scenes, the depiction of flowers shifts away from realistic renderings of individual petals toward a looser, more allusive expression, conveying a sense of floral presence rather than botanical detail. This study highlights the value of computational analysis for a more accurate understanding of an artist's stylistic development. Jia Li 0001, Chaewan Chun, Kathryn Brown, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Refining pseudo-labels through iterative mix-up for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS) aims to provide accurate pixel-level annotation based on only weak guidance, primarily derived from image-level labels. Recent WSSS methods exploit pseudo-labels generated from improved class activation maps (CAMs) to train a fine-grained classification model for semantic segmentation. However, these pseudo-labels are unreliable because they tend to either miss parts of the objects or include irrelevant regions due to weak guidance from individual images. In this paper, we propose a simple yet effective iterative mix-up strategy, Pseudo-Label-based Mix (PL-Mix), that refines pseudo-labels iteratively, thereby further enhancing WSSS performance. During each iteration, we migrate object regions from pseudo-labels produced in previous steps and render them with new contexts in a mix-up fashion. Due to model consistency enforcement across varied backgrounds and new combinations of multiple objects from enriched image samples, these pseudo-labels progressively become more accurate and reliable. Further enhanced by a masking strategy and a CAM-based earth mover’s distance loss, we achieve state-of-the-art performance on the PASCAL VOC2012 and MS COCO2014 benchmark datasets. Yifan Wang 0008, Kunhao Yuan, Gerald Schaefer, Xiyao Liu 0001, Linglin Jing, Kehua Guo, James Z. Wang 0001, Hui Fang 0003 |
Pattern Recognit. | 7 |
| 2026 | The Body in Affective Robotics: A Survey and Conceptual Positioning Using the Performing Arts as a Scaffold for Understanding Bodily Expressed EmotionabstractAffective robotics centers on recognizing emotional states and generating artificial emotions through embodied robotic systems. This paper surveys the current state of the field, with a particular focus on bodily expressed emotion—both in recognizing affect through body movements and postures, and in generating movement that is parsed as affect by human observers. Framed through the lens of the performing arts, this examination provides insights into the expressive potential of robots, motivates key open questions, and highlights challenging problems, as presented through an art-inspired case study and foundational background material. A close engagement with the performing arts suggests intense malleability and diversity of bodily expression, challenging some of the field's prevailing goals—such as designing generally “happy” robotic movement—and emphasizing the importance of variables such as context and interactional intent. The paper concludes by proposing future directions for bodily expressed affective robotics that integrate advances from both robotics and the performing arts. Damith Chandana Herath, Amy LaViers, Sitao Zhang, Nipuni Hansika Wijesinghe, Sharni Konrad, Stelarc, Janie Busby Grant, James Z. Wang 0001 |
IEEE Trans. Affect. Comput. | 8 |
| 2025 | S2S2: Semantic Stacking for Robust Semantic Segmentation in Medical ImagingabstractRobustness and generalizability in medical image segmentation are often hindered by scarcity and limited diversity of training data, which stands in contrast to the variability encountered during inference. While conventional strategies-such as domain-specific augmentation, specialized architectures, and tailored training procedures-can alleviate these issues, they depend on the availability and reliability of domain knowledge. When such knowledge is unavailable, misleading, or improperly applied, performance may deteriorate. In response, we introduce a novel, domain-agnostic, add-on, and data-driven strategy inspired by image stacking in image denoising. Termed "semantic stacking," our method estimates a denoised semantic representation that complements the conventional segmentation loss during training. This method does not depend on domain-specific assumptions, making it broadly applicable across diverse image modalities, model architectures, and augmentation techniques. Through extensive experiments, we validate the superiority of our approach in improving segmentation performance under diverse conditions. Yimu Pan, Sitao Zhang, Alison D. Gernand, Jeffery A. Goldstein, James Z. Wang 0001 |
AAAI | 5 |
| 2025 | Enhancing AI-Assisted Stroke Emergency Triage with Adaptive Uncertainty Estimation
Tongan Cai, Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Sharon X. Huang, Stephen T. C. Wong |
MICCAI (14) | 8 |
| 2025 | Guest Editorial: Special Issue on Fuzzy Affective Computing Systems
Sicheng Zhao, Hongxun Yao, Xinde Li, James Z. Wang 0001, Björn W. Schuller |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | Deep Rank-Consistent Pyramid Model for Enhanced Crowd CountingabstractMost conventional crowd counting methods utilize a fully-supervised learning framework to establish a mapping between scene images and crowd density maps. They usually rely on a large quantity of costly and time-intensive pixel-level annotations for training supervision. One way to mitigate the intensive labeling effort and improve counting accuracy is to leverage large amounts of unlabeled images. This is attributed to the inherent self-structural information and rank consistency within a single image, offering additional qualitative relation supervision during training. Contrary to earlier methods that utilized the rank relations at the original image level, we explore such rank-consistency relation within the latent feature spaces. This approach enables the incorporation of numerous pyramid partial orders, strengthening the model representation capability. A notable advantage is that it can also increase the utilization ratio of unlabeled samples. Specifically, we propose a Deep Rank-consist Ent pyrAmid Model (DREAM), which makes full use of rank consistency across coarse-to-fine pyramid features in latent spaces for enhanced crowd counting with massive unlabeled images. In addition, we have collected a new unlabeled crowd counting dataset, FUDAN-UCC, comprising 4000 images for training purposes. Extensive experiments on four benchmark datasets, namely UCF-QNRF, ShanghaiTech PartA and PartB, and UCF-CC-50, show the effectiveness of our method compared with previous semi-supervised methods. The codes are available at https://github.com/bridgeqiqi/DREAM. Zhizhong Huang, Hongming Shan, James Z. Wang 0001, Fei-Yue Wang 0001, Junping Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | A Heterogeneous Multimodal Graph Learning Framework for Recognizing User Emotions in Social NetworksabstractThe rapid expansion of social media platforms has provided unprecedented access to massive amounts of multi-modal user-generated content. Comprehending user emotions can provide valuable insights for improving communication and understanding of human behaviors. Despite significant advancements in Affective Computing, the diverse factors influencing user emotions in social networks remain relatively understudied. Moreover, there is a notable lack of deep learning-based methods for predicting user emotions in social networks, which could be addressed by leveraging the extensive multimodal data available. This work presents a novel formulation of personalized emotion prediction in social networks based on heterogeneous graph learning. Building upon this formulation, we design HMG-Emo, a Heterogeneous Multimodal Graph Learning Framework that utilizes deep learning-based features for user emotion recognition. Additionally, we include a dynamic context fusion module in HMG-Emo that is capable of adaptively integrating the different modalities in social media data. Through extensive experiments, we demonstrate the effectiveness of HMG-Emo and verify the superiority of adopting a graph neural network-based approach, which outperforms existing baselines that use rich hand-crafted features. To the best of our knowledge, HMG-Emo is the first multimodal and deep-learning-based approach to predict personalized emotions within online social networks. Our work highlights the significance of exploiting advanced deep learning techniques for less-explored problems in Affective Computing. Sree Bhattacharyya, James Z. Wang 0001 |
ACII | 3 |
| 2024 | A Machine Learning Paradigm for Studying Pictorial Realism: How Accurate are Constable's Clouds?abstractThe British landscape painter John Constable is considered foundational for the Realist movement in 19th-century European painting. Constable's painted skies, in particular, were seen as remarkably accurate by his contemporaries, an impression shared by many viewers today. Yet, assessing the accuracy of realist paintings like Constable's is subjective or intuitive, even for professional art historians, making it difficult to say with certainty what set Constable's skies apart from those of his contemporaries. Our goal is to contribute to a more objective understanding of Constable's realism. We propose a new machine-learning-based paradigm for studying pictorial realism in an explainable way. Our framework assesses realism by measuring the similarity between clouds painted by artists noted for their skies, like Constable, and photographs of clouds. The experimental results of cloud classification show that Constable approximates more consistently than his contemporaries the formal features of actual clouds in his paintings. The study, as a novel interdisciplinary approach that combines computer vision and machine learning, meteorology, and art history, is a springboard for broader and deeper analyses of pictorial realism. Zhuomin Zhang, Elizabeth C. Mansfield, Jia Li 0001, George S. Young, Catherine Adams, Kevin A. Bowley, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | HICEM: A High-Coverage Emotion Model for Artificial Emotional IntelligenceabstractAs social robots and other intelligent machines enter the home, artificial emotional intelligence (AEI) is taking center stage to address users' desire for deeper, more meaningful human-machine interaction. To accomplish such efficacious interaction, the next-generation AEI needs comprehensive human emotion models for training. Unlike theory of emotion, which has been the historical focus in psychology, emotion models are a descriptive tool. In practice, the strongest models need robust coverage, which means defining the smallest core set of emotions from which all others can be derived. To achieve the desired coverage, we turn to word embeddings from natural language processing. Using unsupervised clustering techniques, our experiments show that with as few as 15 discrete emotion categories, we can provide maximum coverage across six major languages–Arabic, Chinese, English, French, Spanish, and Russian. In support of our findings, we also examine annotations from two large-scale emotion recognition datasets to assess the validity of existing emotion models compared to human perception at scale. Because robust, comprehensive emotion models are foundational for developing real-world affective computing applications, this work has broad implications in social robotics, human-machine interaction, mental healthcare, computational psychology, and entertainment. Benjamin Wortman, James Z. Wang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Learning Emotion Representations from Verbal and Nonverbal CommunicationabstractEmotion understanding is an essential but highly challenging component of artificial general intelligence. The absence of extensive annotated datasets has significantly impeded advancements in this field. We present Emotion-CLIp, the first pre-training paradigm to extract visual emotion representations from verbal and nonverbal communication using only uncurated data. Compared to numerical labels or descriptions used in previous methods, communication naturally contains emotion information. Furthermore, acquiring emotion representations from communication is more congruent with the human learning process. We guide EmotionCLIP to attend to nonverbal emotion cues through subject-aware context encoding and verbal emotion cues using sentiment-guided contrastive learning. Extensive experiments validates the effectiveness and transferability of Emotion Clip. Using merely linear-probe evaluation protocol, EmotionCLIP outperforms the state-of-the-art supervised visual emotion recognition methods and rivals many multimodal approaches across various benchmarks. We anticipate that the advent of Emotion Clip will address the prevailing issue of data scarcity in emotion understanding, thereby fostering progress in related domains. The code and pretrained models are available at https://github.com/Xeaver/EmotionCLIP. Sitao Zhang, Yimu Pan, James Z. Wang 0001 |
CVPR | 3 |
| 2023 | Enhancing Automatic Placenta Analysis Through Distributional Feature Recomposition in Vision-Language Contrastive Learning
Yimu Pan, Tongan Cai, Manas Mehta, Alison D. Gernand, Jeffery A. Goldstein, Leena B. Mithal, Delia Mwinyelle, Kelly Gallagher, James Z. Wang 0001 |
MICCAI (6) | 9 |
| 2023 | Forget less, count better: a domain-incremental self-distillation learning benchmark for lifelong crowd countingabstractCrowd counting has important applications in public safety and pandemic control. A robust and practical crowd counting system has to be capable of continuously learning with the newly incoming domain data in real-world scenarios instead of fitting one domain only. Off-the-shelf methods have some drawbacks when handling multiple domains: (1) the models will achieve limited performance (even drop dramatically) among old domains after training images from new domains due to the discrepancies in intrinsic data distributions from various domains, which is called catastrophic forgetting; (2) the well-trained model in a specific domain achieves imperfect performance among other unseen domains because of domain shift; (3) it leads to linearly increasing storage overhead, either mixing all the data for training or simply training dozens of separate models for different domains when new ones are available. To overcome these issues, we investigate a new crowd counting task in incremental domain training setting called lifelong crowd counting. Its goal is to alleviate catastrophic forgetting and improve the generalization ability using a single model updated by the incremental domains. Specifically, we propose a self-distillation learning framework as a benchmark (forget less, count better, or FLCB) for lifelong crowd counting, which helps the model leverage previous meaningful knowledge in a sustainable manner for better crowd counting to mitigate the forgetting when new data arrive. A new quantitative metric, normalized Backward Transfer (nBwT), is developed to evaluate the forgetting degree of the model in the lifelong learning process. Extensive experimental results demonstrate the superiority of our proposed benchmark in achieving a low catastrophic forgetting degree and strong generalization ability. Hongming Shan, Yanyun Qu, James Z. Wang 0001, Fei-Yue Wang 0001, Junping Zhang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding EmotionabstractThe emergence of artificial emotional intelligence technology is revolutionizing the fields of computers and robotics, allowing for a new level of communication and understanding of human behavior that was once thought impossible. While recent advancements in deep learning have transformed the field of computer vision, automated understanding of evoked or expressed emotions in visual media remains in its infancy. This foundering stems from the absence of a universally accepted definition of "emotion," coupled with the inherently subjective nature of emotions and their intricate nuances. In this article, we provide a comprehensive, multidisciplinary overview of the field of emotion analysis in visual media, drawing on insights from psychology, engineering, and the arts. We begin by exploring the psychological foundations of emotion and the computational principles that underpin the understanding of emotions from images and videos. We then review the latest research and systems within the field, accentuating the most promising approaches. We also discuss the current technological challenges and limitations of emotion analysis, underscoring the necessity for continued investigation and innovation. We contend that this represents a "Holy Grail" research problem in computing and delineate pivotal directions for future inquiry. Finally, we examine the ethical ramifications of emotion-understanding technologies and contemplate their potential societal impacts. Overall, this article endeavors to equip readers with a deeper understanding of the domain of emotion analysis in visual media and to inspire further research and development in this captivating and rapidly evolving field. James Z. Wang 0001, Sicheng Zhao, Chenyan Wu, Reginald B. Adams Jr., Michelle G. Newman, Tal Shafir, Rachelle Tsachor |
Proc. IEEE | 1 |
| 2022 | Asymmetry Disentanglement Network for Interpretable Acute Ischemic Stroke Infarct Segmentation in Non-contrast CT Scans
Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, Stephen T. C. Wong, James Z. Wang 0001, Sharon X. Huang |
MICCAI (8) | 6 |
| 2022 | Patcher: Patch Transformers with Mixture of Experts for Precise Medical Image Segmentation
Yanglan Ou, Sharon X. Huang, Stephen T. C. Wong, John Volpi, James Z. Wang 0001, Kelvin K. Wong |
MICCAI (5) | 6 |
| 2022 | Vision-Language Contrastive Learning Approach to Robust Automatic Placenta Analysis Using Photographic Images
Yimu Pan, Alison D. Gernand, Jeffery A. Goldstein, Leena B. Mithal, Delia Mwinyelle, James Z. Wang 0001 |
MICCAI (3) | 6 |
| 2022 | DeepStroke: An efficient stroke screening framework for emergency rooms with multimodal adversarial deep learning
Tongan Cai, Haomiao Ni, Mingli Yu, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
Medical Image Anal. | 7 |
| 2022 | SSAS: Spatiotemporal Scale Adaptive Selection for Improving Bias Correction on PrecipitationabstractBy utilizing physical models of the atmosphere collected from the current weather conditions, the numerical weather prediction model developed by the European Centre for Medium-range Weather Forecasts (ECMWF) can provide the indicators of severe weather such as heavy precipitation for an early-warning system. However, the performance of precipitation forecasts from ECMWF often suffers from considerable prediction biases due to the high complexity and uncertainty for the formation of precipitation. The bias correcting on precipitation (BCoP) was thus utilized for correcting these biases via forecasting variables, including the historical observations and variables of precipitation, and these variables, as predictors, from ECMWF are highly relevant to precipitation. The existing BCoP methods, such as model output statistics and ordinal boosting autoencoder, do not take advantage of both spatiotemporal (ST) dependencies of precipitation and scales of related predictors that can change with different precipitation. We propose an end-to-end deep-learning BCoP model, called the ST scale adaptive selection (SSAS) model, to automatically select the ST scales of the predictors via ST Scale-Selection Modules (S3M/TS2M) for acquiring the optimal high-level ST representations. Qualitative and quantitative experiments carried out on two benchmark datasets indicate that SSAS can achieve state-of-the-art performance, compared with 11 published BCoP methods, especially on heavy precipitation. Yiqun Liu 0009, Junping Zhang, Hai Chu, James Z. Wang 0001, Leiming Ma |
IEEE Trans. Cybern. | 5 |
| 2022 | CAPTAIN: Comprehensive Composition Assistance for Photo TakingabstractMany people are interested in taking astonishing photos and sharing them with others. Emerging high-tech hardware and software facilitate the ubiquitousness and functionality of digital photography. Because composition matters in photography, researchers have leveraged some common composition techniques, such as the rule of thirds and the perspective-related techniques, in providing photo-taking assistance. However, composition techniques developed by professionals are far more diverse than well-documented techniques can cover. We present a new approach to leverage the underexplored photography ideas, which are virtually unlimited, diverse, and correlated. We propose a comprehensive fork-join framework, named CAPTAIN ( C omposition A ssistance for P hoto Ta k in g), to guide a photographer with a variety of photography ideas. The framework consists of a few components: integrated object detection, photo genre classification, artistic pose clustering, and personalized aesthetics-aware image retrieval. CAPTAIN is backed by a large managed dataset crawled from a Website with ideas from photography enthusiasts and professionals. The work proposes steps to decompose a given amateurish shot into composition ingredients and compose them to bring the photographer a list of useful and related ideas. The work addresses personal preferences for composition by presenting a user-specified preference list of photography ideas. We have conducted many experiments on the newly proposed components and reported findings. A user study demonstrates that the work is useful to those taking photos. Farshid Farhat, Mohammad Mahdi Kamani, James Z. Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | LambdaUNet: 2.5D Stroke Lesion Segmentation of Diffusion-Weighted MR Images
Yanglan Ou, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
MICCAI (1) | 6 |
| 2020 | MEBOW: Monocular Estimation of Body Orientation in the WildabstractBody orientation estimation provides crucial visual cues in many applications, including robotics and autonomous driving. It is particularly desirable when 3-D pose estimation is difficult to infer due to poor image resolution, occlusion or indistinguishable body parts. We present COCO-MEBOW (Monocular Estimation of Body Orientation in the Wild), a new large-scale dataset for orientation estimation from a single in-the-wild image. The body-orientation labels for around 130K human bodies within 55K images from the COCO dataset have been collected using an efficient and high-precision annotation pipeline. We also validated the benefits of the dataset. First, we show that our dataset can substantially improve the performance and the robustness of a human body orientation estimation model, the development of which was previously limited by the scale and diversity of the available training data. Additionally, we present a novel triple-source solution for 3-D human pose estimation, where 3-D pose labels, 2-D pose labels, and our body-orientation labels are all used in joint training. Our model significantly outperforms state-of-the-art dual-source solutions for monocular 3-D human pose estimation, where training only uses 3-D pose labels and 2-D pose labels. This substantiates an important advantage of MEBOW for 3-D human pose estimation, which is particularly appealing because the per-instance labeling cost for body orientations is far less than that for 3-D poses. The work demonstrates high potential of MEBOW in addressing real-world challenges involving understanding human behaviors. Further information of this work is available at https://chenyanwu.github.io/MEBOW/. Chenyan Wu, Jiajia Luo, Che-Chun Su, Anuja Dawane, Bikramjot Hanzra, Bilan Liu, James Z. Wang 0001, Cheng-Hao Kuo |
CVPR | 9 |
| 2020 | Targeted Data-driven Regularization for Out-of-Distribution GeneralizationabstractDue to biases introduced by large real-world datasets, deviations of deep learning models from their expected behavior on out-of-distribution test data are worrisome. Especially when data come from imbalanced or heavy-tailed label distributions, or minority groups of a sensitive feature. Classical approaches to address these biases are mostly data- or application-dependent, hence are burdensome to tune. Some meta-learning approaches, on the other hand, aim to learn hyperparameters in the learning process using different objective functions on training and validation data. However, these methods suffer from high computational complexity and are not scalable to large datasets. In this paper, we propose a unified data-driven regularization approach to learn a generalizable model from biased data. The proposed framework, named as targeted data-driven regularization (TDR), is model- and dataset-agnostic, and employs a target dataset that resembles the desired nature of test data in order to guide the learning process in a coupled manner. We cast the problem as a bilevel optimization and propose an efficient stochastic gradient descent based method to solve it. The framework can be utilized to alleviate various types of biases in real-world applications. We empirically show, on both synthetic and real-world datasets, the superior performance of TDR for resolving issues stem from these biases. Mohammad Mahdi Kamani, Sadegh Farhang, Mehrdad Mahdavi, James Z. Wang 0001 |
KDD | 4 |
| 2020 | Toward Rapid Stroke Diagnosis with Multimodal Deep Learning
Mingli Yu, Tongan Cai, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
MICCAI (3) | 6 |
| 2020 | ARBEE: Towards Automated Recognition of Bodily Expression of Emotion in the Wild
Yu Luo 0009, Jianbo Ye, Reginald B. Adams Jr., Jia Li 0001, Michelle G. Newman, James Z. Wang 0001 |
Int. J. Comput. Vis. | 6 |
| 2020 | Multi-region saliency-aware learning for cross-domain placenta image segmentationabstractWe propose a multi-region saliency-aware learning (MSL) method for cross-domain placenta image segmentation. Unlike most existing image-level transfer learning methods that fail to preserve the semantics of paired regions, our MSL incorporates the attention mechanism and a saliency constraint into the adversarial translation process, which can realize multi-region mappings in the semantic level. Specifically, the built-in attention module serves to detect the most discriminative semantic regions that the generator should focus on. Then we use the attention consistency as another guidance for retaining semantics after translation. Furthermore, we exploit the specially designed saliency-consistent constraint to enforce the semantic consistency by requiring the saliency regions unchanged. We conduct experiments using two real-world placenta datasets we have collected. We examine the efficacy of this approach in (1) segmentation and (2) prediction of the placental diagnoses of fetal and maternal inflammatory response (FIR, MIR). Experimental results show the superiority of the proposed approach over the state of the art. Zhuomin Zhang, Dolzodmaa Davaasuren, Chenyan Wu, Jeffery A. Goldstein, Alison D. Gernand, James Z. Wang 0001 |
Pattern Recognit. Lett. | 6 |
| 2020 | PaDNet: Pan-Density Crowd CountingabstractCrowd counting is a highly challenging problem in computer vision and machine learning. Most previous methods have focused on consistent density crowds, i.e., either a sparse or a dense crowd, meaning they performed well in global estimation while neglecting local accuracy. To make crowd counting more useful in the real world, we propose a new perspective, named pan-density crowd counting, which aims to count people in varying density crowds. Specifically, we propose the Pan-Density Network (PaDNet) which is composed of the following critical components. First, the Density-Aware Network (DAN) contains multiple subnetworks pretrained on scenarios with different densities. This module is capable of capturing pandensity information. Second, the Feature Enhancement Layer (FEL) effectively captures the global and local contextual features and generates a weight for each density-specific feature. Third, the Feature Fusion Network (FFN) embeds spatial context and fuses these density-specific features. Further, the metrics Patch MAE (PMAE) and Patch RMSE (PRMSE) are proposed to better evaluate the performance on the global and local estimations. Extensive experiments on four crowd counting benchmark datasets, the ShanghaiTech, the UCF-CC-50, the UCSD, and the UCFQNRF, indicate that PaDNet achieves state-of-the-art recognition performance and high robustness in pan-density crowd counting. Yukun Tian, Junping Zhang, James Z. Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | PlacentaNet: Automatic Morphological Characterization of Placenta Photos with Deep Learning
Chenyan Wu, Zhuomin Zhang, Jeffery A. Goldstein, Alison D. Gernand, James Z. Wang 0001 |
MICCAI (1) | 6 |
| 2019 | A recommender system for component-based applications using machine learning techniques
Antonio Jesús Fernández-García, Luis Iribarne, Antonio Corral, Javier Criado, James Z. Wang 0001 |
Knowl. Based Syst. | 5 |
| 2019 | Probabilistic Multigraph Modeling for Improving the Quality of Crowdsourced Affective DataabstractWe proposed a probabilistic approach to joint modeling of participants' reliability and humans' regularity in crowdsourced affective studies. Reliability measures how likely a subject will respond to a question seriously; and regularity measures how often a human will agree with other seriously-entered responses coming from a targeted population. Crowdsourcing-based studies or experiments, which rely on human self-reported affect, pose additional challenges as compared with typical crowdsourcing studies that attempt to acquire concrete non-affective labels of objects. The reliability of participants has been massively pursued for typical non-affective crowdsourcing studies, whereas the regularity of humans in an affective experiment in its own right has not been thoroughly considered. It has been often observed that different individuals exhibit different feelings on the same test question, which does not have a sole correct response in the first place. High reliability of responses from one individual thus cannot conclusively result in high consensus across individuals. Instead, globally testing consensus of a population is of interest to investigators. Built upon the agreement multigraph among tasks and workers, our probabilistic model differentiates subject regularity from population reliability. We demonstrate the method's effectiveness for in-depth robust analysis of large-scale crowdsourced affective data, including emotion and aesthetic assessments collected by presenting visual stimuli to human subjects. Jianbo Ye, Jia Li 0001, Michelle G. Newman, Reginald B. Adams Jr., James Z. Wang 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2019 | Detecting Comma-Shaped Clouds for Severe Weather Forecasting Using Shape and MotionabstractMeteorologists use shapes and movements of clouds in satellite images as indicators of several major types of severe storms. Yet, because satellite image data are in increasingly higher resolution, both spatially and temporally, meteorologists cannot fully leverage the data in their forecasts. Automatic satellite image analysis methods that can find storm-related cloud patterns are thus in demand. We propose a machine learning and pattern recognition-based approach to detect “comma-shaped” clouds in satellite images, which are specific cloud distribution patterns strongly associated with cyclone formulation. In order to detect regions with the targeted movement patterns, we use manually annotated cloud examples represented by both shape and motion-sensitive features to train the computer to analyze satellite images. Sliding windows in different scales ensure the capture of dense clouds, and we implement effective selection rules to shrink the region of interest among these sliding windows. Finally, we evaluate the method on a hold-out annotated comma-shaped cloud data set and cross match the results with recorded storm events in the severe weather database. The validated utility and accuracy of our method suggest a high potential for assisting meteorologists in weather forecasting. Xinye Zheng, Jianbo Ye, Stephen Wistar, Jia Li 0001, Jose A. Piedra-Fernández, Michael A. Steinberg, James Z. Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2019 | Crowd Counting With Limited Labeling Through Submodular Frame SelectionabstractAutomated crowd counting is valuable for intelligent transportation systems, as it can help to improve the emergency planning and prevent congestion in transit hubs such as train stations and airports. Semi-supervised crowd counting aims to estimate the number of pedestrians in an ongoing scene using a combination of a small number of labeled frames and a large number of unlabeled ones. However, existing methods do not incorporate ways to effectively select informative frames as labeled training samples, resulting in low accuracy on unseen crowd scenes. We propose a submodular method to select the most informative frames from the image sequences of crowds. Specifically, the method selects the most representative images to guarantee the information coverage, by maximizing the similarities between the group of selected images and the image sequence. In addition, these frames are chosen to avoid redundancies and preserve diversity. Finally, our semi-supervised method incorporates graph Laplacian regularization and spatiotemporal constraints. Extensive experiments on three benchmark data sets demonstrate that our proposed approach achieves higher accuracy compared with the state-of-the-art regression methods and competitive performance with deep convolutional models, especially when the number of labeled data is exceptionally small. Junping Zhang, Lingfu Che, Hongming Shan, James Z. Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2018 | Rethinking the Smaller-Norm-Less-Informative Assumption in Channel Pruning of Convolution Layers
Jianbo Ye, Xin Lu 0006, Zhe Lin 0001, James Z. Wang 0001 |
ICLR (Poster) | 4 |
| 2018 | Discovering Triangles in Portraits for Supporting Photographic CreationabstractIncorporating the concept of triangles in photos is an effective composition technique used by professional photographers for making pictures more interesting or dynamic. Information on the locations of the embedded triangles is valuable for comparing the composition of portrait photos which can be further leveraged by a retrieval system or used by the photographers. This paper presents a system to automatically detect embedded triangles in portrait photos. The problem is challenging because the triangles used in portraits are often not clearly defined by straight lines. The system first extracts a set of filtered line segments as candidate triangle sides and then utilizes a modified random sample consensus algorithm to fit triangles onto the set of line segments. We propose two metrics Continuity Ratio and Total Ratio to evaluate the fitted triangles; those with high fitting scores are taken as detected triangles. Experimental results have demonstrated high accuracy in locating preeminent triangles in portraits without dependence on the camera or lens parameters. To demonstrate the benefits of our method to digital photography we have developed two novel applications that aim to help users compose high-quality photos. In the first application we develop a human position and pose recommendation system by retrieving and presenting compositionally similar photos taken by competent photographers. The second application is a novel sketch-based triangle retrieval system which searches for photos containing a specific triangular configuration. User studies have been conducted to validate the effectiveness of these approaches. Siqiong He, Zihan Zhou 0001, Farshid Farhat, James Z. Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | An investigation into three visual characteristics of complex scenes that evoke human emotionabstractPrior computational studies have examined hundreds of visual characteristics related to color, texture, and composition in an attempt to predict human emotional responses. Beyond those myriad features examined in computer science, roundness, angularity, and visual complexity have also been found to evoke emotions in human perceivers, as demonstrated in psychological studies of facial expressions, dance poses, and even simple synthetic visual patterns. Capturing these characteristics algorithmically to incorporate in computational studies, however, has proven difficult. Here we expand the scope of previous computer vision work by examining these three visual characteristics in computer analysis of complex scenes, and compare the results to the hundreds of visual qualities previously examined. A large collection of ecologically valid stimuli (i.e., photos that humans regularly encounter on the web), named the EmoSet and containing more than 40,000 images crawled from web albums, was generated using crowd-sourcing and subjected to human subject emotion ratings. We developed computational methods to the separate indices of roundness, angularity, and complexity, thereby establishing three new computational constructs. Critically, these three new physically interpretable visual constructs achieve comparable classification accuracy to the hundreds of shape, texture, composition, and facial feature characteristics previously examined. In addition, our experimental results show that color features related most strongly with the positivity of perceived emotions, the texture features related more to calmness or excitement, and roundness, angularity, and simplicity related similarly with both of these emotions dimensions. Xin Lu 0006, Reginald B. Adams Jr., Jia Li 0001, Michelle G. Newman, James Z. Wang 0001 |
ACII | 5 |
| 2017 | Determining Gains Acquired from Word Embedding Quantitatively Using Discrete Distribution ClusteringabstractWord embeddings have become widelyused in document analysis.While a large number of models for mapping words to vector spaces have been developed, it remains undetermined how much net gain can be achieved over traditional approaches based on bag-of-words.In this paper, we propose a new document clustering approach by combining any word embedding with a state-of-the-art algorithm for clustering empirical distributions.By using the Wasserstein distance between distributions, the word-to-word semantic relationship is taken into account in a principled way.The new clustering method is easy to use and consistently outperforms other methods on a variety of data sets.More importantly, the method provides an effective framework for determining when and how much word embeddings contribute to document analysis.Experimental results with multiple embedding models are reported. Jianbo Ye, Yanran Li, James Z. Wang 0001, Wenjie Li 0002, Jia Li 0001 |
ACL (1) | 4 |
| 2017 | A Simulated Annealing Based Inexact Oracle for Wasserstein Loss MinimizationabstractLearning under a Wasserstein loss, a.k.a. Wasserstein loss minimization (WLM), is an emerging research topic for gaining insights from a large set of structured objects. Despite being conceptually simple, WLM problems are computationally challenging because they involve minimizing over functions of quantities (i.e. Wasserstein distances) that themselves require numerical algorithms to compute. In this paper, we introduce a stochastic approach based on simulated annealing for solving WLMs. Particularly, we have developed a Gibbs sampler to approximate effectively and efficiently the partial gradients of a sequence of Wasserstein losses. Our new approach has the advantages of numerical stability and readiness for warm starts. These characteristics are valuable for WLM problems that often require multiple levels of iterations in which the oracle for computing the value and gradient of a loss function is embedded. We applied the method to optimal transport with Coulomb cost and the Wasserstein non-negative matrix factorization problem, and made comparisons with the existing method of entropy regularization. Jianbo Ye, James Z. Wang 0001, Jia Li 0001 |
ICML | 2 |
| 2017 | Microexpression Identification and Categorization Using a Facial Dynamics MapabstractUnlike conventional facial expressions, microexpressions are instantaneous and involuntary reflections of human emotion. Because microexpressions are fleeting, lasting only a few frames within a video sequence, they are difficult to perceive and interpret correctly, and they are highly challenging to identify and categorize automatically. Existing recognition methods are often ineffective at handling subtle face displacements, which can be prevalent in typical microexpression applications due to the constant movements of the individuals being observed. To address this problem, a novel method called the Facial Dynamics Map is proposed to characterize the movements of a microexpression in different granularity. Specifically, an algorithm based on optical flow estimation is used to perform pixel-level alignment for microexpression sequences. Each expression sequence is then divided into spatiotemporal cuboids in the chosen granularity. We also present an iterative optimal strategy to calculate the principal optical flow direction of each cuboid for better representation of the local facial dynamics. With these principal directions, the resulting Facial Dynamics Map can characterize a microexpression sequence. Finally, a classifier is developed to identify the presence of microexpressions and to categorize different types. Experimental results on four benchmark datasets demonstrate higher recognition performance and improved interpretability. Junping Zhang, James Z. Wang 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Severe Thunderstorm Detection by Visual Learning Using Satellite ImagesabstractComputers are widely utilized in today's weather forecasting as a powerful tool to leverage an enormous amount of data. Yet, despite the availability of such data, current techniques often fall short of producing reliable detailed storm forecasts. Each year severe thunderstorms cause significant damage and loss of life, some of which could be avoided if better forecasts were available. We propose a computer algorithm that analyzes satellite images from historical archives to locate visual signatures of severe thunderstorms for short-term predictions. While computers are involved in weather forecasts to solve numerical models based on sensory data, they are less competent in forecasting based on visual patterns from both current and past satellite images. In our system, we extract and summarize important visual storm evidence from satellite image sequences in the way that meteorologists interpret the images. In particular, the algorithm extracts and fits local cloud motion from image sequences to model the storm-related cloud patches. Image data from the year 2008 have been adopted to train the model, and historical severe thunderstorm reports in continental U.S. from 2000 to 2013 have been used as the ground truth and priors in the modeling process. Experiments demonstrate the usefulness and potential of the algorithm for producing more accurate severe thunderstorm forecasts. Yu Zhang 0013, Stephen Wistar, Jia Li 0001, Michael A. Steinberg, James Z. Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Detecting Dominant Vanishing Points in Natural Scenes with Application to Composition-Sensitive Image RetrievalabstractLinear perspective is widely used in landscape photography to create the impression of depth on a 2D photo. Automated understanding of linear perspective in landscape photography has several real-world applications, including aesthetics assessment, image retrieval, and on-site feedback for photo composition, yet adequate automated understanding has been elusive. We address this problem by detecting the dominant vanishing point and the associated line structures in a photo. However, natural landscape scenes pose great technical challenges because often the number of strong edges converging to the dominant vanishing point is inadequate. To overcome this difficulty, we propose a novel vanishing point detection method that exploits global structures in the scene via contour detection. We show that our method significantly outperforms state-of-the-art methods on a public ground truth landscape image dataset that we have created. Based on the detection results, we further demonstrate how our approach to linear perspective understanding provides on-site guidance to amateur photographers on their work through a novel viewpoint-specific image retrieval system. Zihan Zhou 0001, Farshid Farhat, James Z. Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Shape matching using skeleton context for automated bow echo detectionabstractSevere weather conditions cause enormous amount of damages around the globe. Bow echo patterns in radar images are associated with a number of these destructive thunderstorm conditions such as damaging winds, hail and tornadoes. They are detected manually by meteorologists. In this paper, we propose an automatic framework to detect these patterns with high accuracy by introducing novel skeletonization and shape matching approaches. In this framework, first we extract regions with high probability of occurring bow echo from radar images, and apply our skeletonization method to extract the skeleton of those regions. Next, we prune these skeletons using our innovative pruning scheme with fuzzy logic. Then, using our proposed shape descriptor, Skeleton Context, we can extract bow echo features from these skeletons in order to use them in shape matching algorithm and classification step. The output of classification indicates whether these regions include a bow echo with over 97% accuracy. Mohammad Mahdi Kamani, Farshid Farhat, Stephen Wistar, James Z. Wang 0001 |
IEEE BigData | 4 |
| 2016 | Joint Image and Text Representation for Aesthetics AnalysisabstractImage aesthetics assessment is essential to multimedia applications such as image retrieval, and personalized image search and recommendation. Primarily relying on visual information and manually-supplied ratings, previous studies in this area have not adequately utilized higher-level semantic information. We incorporate additional textual phrases from user comments to jointly represent image aesthetics utilizing multimodal Deep Boltzmann Machine. Given an image, without requiring any associated user comments, the proposed algorithm automatically infers the joint representation and predicts the aesthetics category of the image. We construct the AVA-Comments dataset to systematically evaluate the performance of the proposed algorithm. Experimental results indicate that the proposed joint representation improves the performance of aesthetics assessment on the benchmarking AVA dataset, comparing with only visual features. Xin Lu 0006, Junping Zhang, James Z. Wang 0001 |
ACM Multimedia | 4 |
| 2015 | Deep Multi-patch Aggregation Network for Image Style, Aesthetics, and Quality EstimationabstractThis paper investigates problems of image style, aesthetics, and quality estimation, which require fine-grained details from high-resolution images, utilizing deep neural network training approach. Existing deep convolutional neural networks mostly extracted one patch such as a down-sized crop from each image as a training example. However, one patch may not always well represent the entire image, which may cause ambiguity during training. We propose a deep multi-patch aggregation network training approach, which allows us to train models using multiple patches generated from one image. We achieve this by constructing multiple, shared columns in the neural network and feeding multiple patches to each of the columns. More importantly, we propose two novel network layers (statistics and sorting) to support aggregation of those patches. The proposed deep multi-patch aggregation network integrates shared feature learning and aggregation function learning into a unified framework. We demonstrate the effectiveness of the deep multi-patch aggregation network on the three problems, i.e., image style recognition, aesthetic quality categorization, and image quality estimation. Our models trained using the proposed networks significantly outperformed the state of the art in all three applications. Xin Lu 0006, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, James Z. Wang 0001 |
ICCV | 5 |
| 2015 | Modeling Perspective Effects in Photographic CompositionabstractAutomatic understanding of photo composition is a valuable technology in multiple areas including digital photography, multimedia advertising, entertainment, and image retrieval. In this paper, we propose a method to model geometrically the compositional effects of linear perspective. Comparing with existing methods which have focused on basic rules of design such as simplicity, visual balance, golden ratio, and the rule of thirds, our new quantitative model is more comprehensive whenever perspective is relevant. We first develop a new hierarchical segmentation algorithm that integrates classic photometric cues with a new geometric cue inspired by perspective geometry. We then show how these cues can be used directly to detect the dominant vanishing point in an image without extracting any line segments, a technique with implications for multimedia applications beyond this work. Finally, we demonstrate an interesting application of the proposed method for providing on-site composition feedback through an image retrieval system. Zihan Zhou 0001, Siqiong He, Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 4 |
| 2015 | Contextual and Hierarchical Classification of Satellite Images Based on Cellular AutomataabstractSatellite image classification is an important technique used in remote sensing for the computerized analysis and pattern recognition of satellite data, which facilitates the automated interpretation of a large amount of information. Today, there exist many types of classification algorithms, such as parallelepiped and minimum distance classifiers, but it is still necessary to improve their performance in terms of accuracy rate. On the other hand, over the last few decades, cellular automata have been used in remote sensing to implement processes related to simulations. Although there is little previous research of cellular automata related to satellite image classification, they offer many advantages that can improve the results of classical classification algorithms. This paper discusses the development of a new classification algorithm based on cellular automata which not only improves the classification accuracy rate in satellite images by using contextual techniques but also offers a hierarchical classification of pixels divided into levels of membership degree to each class and includes a spatial edge detection method of classes in the satellite image. Moisés Espínola, Jose A. Piedra-Fernández, Rosa Ayala, Luis Iribarne, James Z. Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Image-Specific Prior Adaptation for DenoisingabstractImage priors are essential to many image restoration applications, including denoising, deblurring, and inpainting. Existing methods use either priors from the given image (internal) or priors from a separate collection of images (external). We find through statistical analysis that unifying the internal and external patch priors may yield a better patch prior. We propose a novel prior learning algorithm that combines the strength of both internal and external priors. In particular, we first learn a generic Gaussian mixture model from a collection of training images and then adapt the model to the given image by simultaneously adding additional components and refining the component parameters. We apply this image-specific prior to image denoising. The experimental results show that our approach yields better or competitive denoising results in terms of both the peak signal-to-noise ratio and structural similarity. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | Rating Image Aesthetics Using Deep LearningabstractThis paper investigates unified feature learning and classifier training approaches for image aesthetics assessment . Existing methods built upon handcrafted or generic image features and developed machine learning and statistical modeling techniques utilizing training examples. We adopt a novel deep neural network approach to allow unified feature learning and classifier training to estimate image aesthetics. In particular, we develop a double-column deep convolutional neural network to support heterogeneous inputs, i.e., global and local views, in order to capture both global and local characteristics of images . In addition, we employ the style and semantic attributes of images to further boost the aesthetics categorization performance . Experimental results show that our approach produces significantly better results than the earlier reported results on the AVA dataset for both the generic image aesthetics and content -based image aesthetics. Moreover, we introduce a 1.5-million image dataset (IAD) for image aesthetics assessment and we further boost the performance on the AVA test set by training the proposed deep neural networks on the IAD dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | Parallel Massive Clustering of Discrete DistributionsabstractThe trend of analyzing big data in artificial intelligence demands highly-scalable machine learning algorithms, among which clustering is a fundamental and arguably the most widely applied method. To extend the applications of regular vector-based clustering algorithms, the Discrete Distribution (D2) clustering algorithm has been developed, aiming at clustering data represented by bags of weighted vectors which are well adopted data signatures in many emerging information retrieval and multimedia learning applications. However, the high computational complexity of D2-clustering limits its impact in solving massive learning problems. Here we present the parallel D2-clustering (PD2-clustering) algorithm with substantially improved scalability. We developed a hierarchical multipass algorithm structure for parallel computing in order to achieve a balance between the individual-node computation and the integration process of the algorithm. Experiments and extensive comparisons between PD2-clustering and other clustering algorithms are conducted on synthetic datasets. The results show that the proposed parallel algorithm achieves significant speed-up with minor accuracy loss. We apply PD2-clustering to image concept learning. In addition, by extending D2-clustering to symbolic data, we apply PD2-clustering to protein sequence clustering. For both applications, we demonstrate the high competitiveness of our new algorithm in comparison with other state-of-the-art methods. Yu Zhang 0013, James Z. Wang 0001, Jia Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Locating visual storm signatures from satellite imagesabstractWeather forecasting is a problem where an enormous amount of data must be processed. Severe storms cause a significant amount of damages and loss every year in part due to the insufficiency of the current techniques in producing reliable forecasts. We propose an algorithm that analyzes satellite images from the vast historical archives to predict severe storms. Conventional weather forecasting involves solving numerical models based on sensory data. It has been challenging for computers to make forecasts based on the visual patterns from satellite images. In our system we extract and summarize important visual storm evidence from satellite image sequences in a way similar to how meteorologists interpret these images. Particularly, the algorithm extracts and fits local cloud motions from image sequences to model the storm-related cloud patches. Image data of an entire year are adopted to train the model. The historical storm reports since the year 2000 are used as the ground-truth and statistical priors in the modeling process. Experiments demonstrate the usefulness and potential of the algorithm for producing improved storm forecasts. Yu Zhang 0013, Stephen Wistar, Jose A. Piedra-Fernández, Jia Li 0001, Michael A. Steinberg, James Z. Wang 0001 |
IEEE BigData | 6 |
| 2014 | RAPID: Rating Pictorial Aesthetics using Deep LearningabstractEffective visual features are essential for computational aesthetic quality rating systems. Existing methods used machine learning and statistical modeling techniques on handcrafted features or generic image descriptors. A recently-published large-scale dataset, the AVA dataset, has further empowered machine learning based approaches. We present the RAPID (RAting PIctorial aesthetics using Deep learning) system, which adopts a novel deep neural network approach to enable automatic feature learning. The central idea is to incorporate heterogeneous inputs generated from the image, which include a global view and a local view, and to unify the feature learning and classifier training using a double-column deep convolutional neural network. In addition, we utilize the style attributes of images to help improve the aesthetic quality categorization accuracy. Experimental results show that our approach significantly outperforms the state of the art on the AVA dataset. Xin Lu 0006, Zhe Lin 0001, Hailin Jin, Jianchao Yang, James Z. Wang 0001 |
ACM Multimedia | 5 |
| 2014 | Fuzzy Content-Based Image Retrieval for Oceanic Remote SensingabstractThe detection of mesoscale oceanic structures, such as upwellings or eddies, from satellite images has significance for marine environmental studies, coastal resource management, and ocean dynamics studies. Nevertheless, there is a lack of tools that allow us to retrieve automatically relevant mesoscale structures from large satellite image databases. This paper focuses on the development and validation of a content-based image retrieval system to classify and retrieve oceanic structures from satellite images. The images were obtained from the National Oceanic and Atmospheric Administration satellite's Advanced Very High Resolution Radiometer sensor. The study area is about W2°- 21°, N19°- 45°. This system conducts labeling and retrieval of the most relevant and typical mesoscale oceanic structures, such as upwellings, eddies, and island wakes located in the Canary Islands area and in the Mediterranean and Cantabrian seas. Our work is based on several soft computing technologies such as fuzzy logic and neurofuzzy systems. Jose A. Piedra-Fernández, Gloria Ortega, James Z. Wang 0001, Manuel Cantón-Garbín |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Enhancing Training Collections for Image Annotation: An Instance-Weighted Mixture Modeling ApproachabstractTagged Web images provide an abundance of labeled training examples for visual concept learning. However, the performance of automatic training data selection is susceptible to highly inaccurate tags and atypical images. Consequently, manually curated training data sets are still a preferred choice for many image annotation systems. This paper introduces ARTEMIS - a scheme to enhance automatic selection of training images using an instance-weighted mixture modeling framework. An optimization algorithm is derived to learn instance-weights in addition to mixture parameter estimation, essentially adapting to the noise associated with each example. The mechanism of hypothetical local mapping is evoked so that data in diverse mathematical forms or modalities can be cohesively treated as the system maintains tractability in optimization. Finally, training examples are selected from top-ranked images of a likelihood-based image ranking. Experiments indicate that ARTEMIS exhibits higher resilience to noise than several baselines for large training data collection. The performance of ARTEMIS-trained image annotation system is comparable with usage of manually curated data sets. Neela Sawant, James Z. Wang 0001, Jia Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | On shape and the computability of emotionsabstractWe investigated how shape features in natural images influence emotions aroused in human beings. Shapes and their characteristics such as roundness, angularity, simplicity, and complexity have been postulated to affect the emotional responses of human beings in the field of visual arts and psychology. However, no prior research has modeled the dimensionality of emotions aroused by roundness and angularity. Our contributions include an in-depth statistical analysis to understand the relationship between shapes and emotions. Through experimental results on the International Affective Picture System (IAPS) dataset we provide evidence for the significance of roundness-angularity and simplicity-complexity on predicting emotional content in images. We combine our shape features with other state-of-the-art features to show a gain in prediction and classification accuracy. We model emotions from a dimensional perspective in order to predict valence and arousal ratings which have advantages over modeling the traditional discrete emotional categories. Finally, we distinguish images with strong emotional content from emotionally neutral images with high accuracy. Xin Lu 0006, Poonam Suryanarayan, Reginald B. Adams Jr., Jia Li 0001, Michelle G. Newman, James Z. Wang 0001 |
ACM Multimedia | 6 |
| 2012 | OSCAR: On-Site Composition and Aesthetics Feedback Through Exemplars for Photographers
Poonam Suryanarayan, James Z. Wang 0001, Jia Li 0001 |
Int. J. Comput. Vis. | 4 |
| 2012 | Thin Cloud Detection of All-Sky Images Using Markov Random FieldsabstractThin cloud detection for all-sky images is a challenge in ground-based sky-imaging systems because of low contrast and vague boundaries between cloud and sky regions. We treat cloud detection as a labeling problem based on the Markov random field model. In this model, each pixel is represented by a combined-feature vector that aims at improving the disparity between thin cloud and sky. The distribution of each label in the feature space is defined as a Gaussian model. Spatial information is coded by a generalized Potts model. During the estimation, thin cloud is detected by minimizing the posterior energy with an iterative procedure. Both subjective and objective evaluation results demonstrate higher accuracy of the algorithm compared with some other algorithms. Qingyong Li, Weitao Lu, James Z. Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2012 | Rhythmic Brushstrokes Distinguish van Gogh from His Contemporaries: Findings via Automated Brushstroke ExtractionabstractArt historians have long observed the highly characteristic brushstroke styles of Vincent van Gogh and have relied on discerning these styles for authenticating and dating his works. In our work, we compared van Gogh with his contemporaries by statistically analyzing a massive set of automatically extracted brushstrokes. A novel extraction method is developed by exploiting an integration of edge detection and clustering-based segmentation. Evidence substantiates that van Gogh's brushstrokes are strongly rhythmic. That is, regularly shaped brushstrokes are tightly arranged, creating a repetitive and patterned impression. We also found that the traits that distinguish van Gogh's paintings in different time periods of his development are all different from those distinguishing van Gogh from his peers. This study confirms that the combined brushwork features identified as special to van Gogh are consistently held throughout his French periods of production (1886-1890). Jia Li 0001, Ella Hendriks, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Analysis of cypriot icon faces using ICA-enhanced active shape model representationabstractReligious iconography is an integral component of the cultural heritage of Cyprus, which was once a part of the great Byzantine empire. On one hand, icons exhibit strict adherence to conventional symbols, poses and apparel. On the other hand, there is a great variety in the style of depiction that can be attributed to different schools and periods. This paper proposes an active shape model (ASM) based technique for icon face representation that can be used for style comparison and attribution. For centuries-old icons suffering from loss of paint, cracks and added noise from digitization artifacts, we apply an independent component analysis (ICA) technique to enhance the paintings' original work. The experimental results show that our method can effectively characterize Cypriot icons. Guifang Duan, Neela Sawant, James Z. Wang 0001, Dean R. Snow, Danni Ai, Yen-Wei Chen 0001 |
ACM Multimedia | 3 |
| 2011 | SHIRAZ: an automated histology image annotation system for zebrafish phenomicsabstractHistological characterization is used in clinical and research contexts as a highly sensitive method for detecting the morphological features of disease and abnormal gene function. Histology has recently been accepted as a phenotyping method for the forthcoming Zebrafish Phenome Project, a large-scale community effort to characterize the morphological, physiological, and behavioral phenotypes resulting from the mutations in all known genes in the zebrafish genome. In support of this project, we present a novel content-based image retrieval system for the automated annotation of images containing histological abnormalities in the developing eye of the larval zebrafish. Brian A. Canada, Georgia K. Thomas, Keith C. Cheng, James Z. Wang 0001 |
Multim. Tools Appl. | 4 |
| 2011 | Automatic image semantic interpretation using social action and tagging data
Neela Sawant, Jia Li 0001, James Z. Wang 0001 |
Multim. Tools Appl. | 3 |
| 2010 | Training data collection system for a learning-based photographic aesthetic quality inference engineabstractWe present a novel data collection system deployed for the ACQUINE - Aesthetic Quality Inference Engine. The goal of the system is to collect online user opinions, both structured and unstructured, for training future generation learning-based aesthetic quality inference engines. The development of the system was based on an analysis of over 60,000 user comments of photographs. For photos processed and rated by our engine, all users are invited to provide manual ratings. The users can also choose up to three key photographic features that the user liked, from a list, or to add features not in the list. Within a few months that the system is available for public used more than 20,000 photos have received manual ratings and key features for over 1,800 photos have been identified. We expect the data generated over time will be critical in the study of computational inferencing of visual aesthetics in photographs. The system is demonstrated at http://acquine.alipr.com Razvan Orendovici, James Z. Wang 0001 |
ACM Multimedia | 2 |
| 2010 | Determining the sexual identities of prehistoric cave artists using digitized handprints: a machine learning approachabstractThe sexual identities of human handprints inform hypotheses regarding the roles of males and females in prehistoric contexts. Sexual identity has previously been manually determined by measuring the ratios of the lengths of the individual's fingers as well as by using other physical features. Most conventional studies measure the lengths manually and thus are often constrained by the lack of scaling information on published images. We have created a method that determines sex by applying modern machine-learning techniques to relative measures obtained from images of human hands. This is the known attempt at substituting automated methods for time-consuming manual measurement in the study of sexual identities of prehistoric cave artists. Our study provides quantitative evidence relevant to sexual dimorphism and the sexual division of labor in Upper Paleolithic societies. In addition to analyzing historical handprint records, this method has potential applications in criminal forensics and human-computer interaction. James Z. Wang 0001, Weina Ge, Dean R. Snow, Prasenjit Mitra 0001, C. Lee Giles |
ACM Multimedia | 1 |
| 2010 | Feature Selection in AVHRR Ocean Satellite Images by Means of Filter MethodsabstractAutomatic retrieval and interpretation of satellite images is critical for managing the enormous volume of environmental remote sensing data available today. It is particularly useful in oceanography and climate studies for examination of the spatio-temporal evolution of mesoscalar ocean structures appearing in the satellite images taken by visible, infrared, and radar sensors. This is because they change so quickly and several images of the same place can be acquired at different times within the same day. This paper describes the use of filter measures and the Bayesian networks to reduce the number of irrelevant features necessary for ocean structure recognition in satellite images, thereby improving the overall interpretation system performance and reducing the computational time. We present our results for the National Oceanographic and Atmospheric Administration satellite Advanced Very High Resolution Radiometer (AVHRR) images. We have automatically detected and located mesoscale ocean phenomena of interest in our study area (North-East Atlantic and the Mediterranean), such as upwellings, eddies, and island wakes, using an automatic selection methodology which reduces the features used for description by about 80%. Finally, Bayesian network classifiers are used to assess classification quality. Knowledge about these structures is represented with numeric and nonnumeric features. Jose A. Piedra-Fernández, Manuel Cantón-Garbín, James Z. Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2009 | Characterizing elegance of curves computationally for distinguishing Morrisseau paintings and the imitationsabstractComputerized analysis of paintings has recently gained interest. The rapid technological advancements and the expanding interdisciplinary collaboration present us a promising prospect of computer-assisted authentication. We focus on the characterization of curve elegance. Specifically, we propose measures of curve steadiness and neighborhood coherence from brushstrokes. The technique has been applied to the paintings of renowned aboriginal Canadian artist Norval Morrisseau. Through computerized analysis of his authentic works and the imitations, it is revealed that the curves in his authentic paintings exhibit his commanding painting skills. The smooth and steady flow of the curves show less hesitancy of the artist than the authors of counterfeit works. The tangent angles tend to be more consistent along curves in the authentic paintings than in the imitations. Jia Li 0001, James Z. Wang 0001 |
ICIP | 3 |
| 2009 | Automated analysis of images in documents for intelligent document search
Saurabh Kataria 0003, William Browuer, James Z. Wang 0001, Prasenjit Mitra 0001, C. Lee Giles |
Int. J. Document Anal. Recognit. | 4 |
| 2009 | Exploiting the human-machine gap in image recognition for designing CAPTCHAsabstractSecurity researchers have, for a long time, devised mechanisms to prevent adversaries from conducting automated network attacks, such as denial-of-service, which lead to significant wastage of resources. On the other hand, several attempts have been made to automatically recognize generic images, make them semantically searchable by content, annotate them, and associate them with linguistic indexes. In the course of these attempts, the limitations of state-of-the-art algorithms in mimicking human vision have become exposed. In this paper, we explore the exploitation of this limitation for potentially preventing automated network attacks. While undistorted natural images have been shown to be algorithmically recognizable and searchable by content to moderate levels, controlled distortions of specific types and strengths can potentially make machine recognition harder without affecting human recognition. This difference in recognizability makes it a promising candidate for automated Turing tests [completely automated public Turing test to tell computers and humans apart (CAPTCHAs)] which can differentiate humans from machines. We empirically study the application of controlled distortions of varying nature and strength, and their effect on human and machine recognizability. While human recognizability is measured on the basis of an extensive user study, machine recognizability is based on memory-based content-based image retrieval (CBIR) and matching algorithms. We give a detailed description of our experimental image CAPTCHA system, IMAGINATION, that uses systematic distortions at its core. A significant research topic within signal analysis, CBIR is actually conceived here as a tool for an adversary, so as to help us design more foolproof image CAPTCHAs. Ritendra Datta, Jia Li 0001, James Z. Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2008 | Automatic lattice detection in near-regular histology array imagesabstractNear-regular texture (NRT), denoting deviations from otherwise symmetric wallpaper patterns, is commonly observable in the real world. Existing lattice detection algorithms capture the underlying lattice of an NRT pattern and all of its individual texels, facilitating an automated analysis of NRT. Many real world images, as in those of zebrafish larval histology arrays, depart significantly from regularity and challenge the current state of the art wallpaper group theory-based lattice detection methods. We propose an alternative 2D lattice detection algorithm that exploits translation and reflection symmetries and specific imaging cues. By outperforming existing methods on histology array images, our algorithm leads us towards complete automation of high-throughput histological image processing while broadening the spectrum of NRT computation. Brian A. Canada, Georgia K. Thomas, Keith C. Cheng, James Z. Wang 0001, Yanxi Liu 0001 |
ICIP | 4 |
| 2008 | Algorithmic inferencing of aesthetics and emotion in natural images: An expositionabstractInitial studies have shown that automatic inference of high-level image quality or aesthetics is very challenging. The ability to do so, however, can prove beneficial in many applications. In this paper, we define the aesthetics gap and discuss key aspects of the problem of aesthetics and emotion inference in natural images. We introduce precise, relevant questions to be answered, the effect that the target audience has on the problem specification, broad technical solution approaches, and assessment criteria. We then report on our effort to build real-world datasets that provide viable approaches to test and compare algorithms for these problems, presenting statistical analysis of and insights into them. Ritendra Datta, Jia Li 0001, James Z. Wang 0001 |
ICIP | 3 |
| 2008 | Boosted cannabis image recognitionabstractWith the large number of Web sites promoting the use of illicit drugs, it has become important to screen these sites for the protection of children on the Internet. Conventional keyword-based approaches are not sufficient because these Web sites often have lots of images and little meaningful words than prices. We propose an AdaBoost-based algorithm for cannabis image recognition. This is the first known attempt at computerized detection of illicit drug Web contents using images. The main technical contributions of our work are two-fold. First, we introduce a novel weak classifier which considers the inherently structural property or ldquoself-similarityrdquo of the cannabis plants. The self-correlation structural characteristics of cannabis can be used as a discriminative property for the purpose of cannabis image recognition. Second, we propose a rapid weak classifier finder, which can efficiently select discriminative weak classifiers from the weak classifier space with little degradation to the classification accuracy. Experiments on real world images have demonstrated improved performance of our method over other methods. Nianhua Xie, Xi Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, James Z. Wang 0001 |
ICPR | 5 |
| 2008 | Real-Time Computerized Annotation of PicturesabstractDeveloping effective methods for automated annotation of digital pictures continues to challenge computer scientists. The capability of annotating pictures by computers can lead to breakthroughs in a wide range of applications, including Web image search, online picture-sharing communities, and scientific experiments. In this work, the authors developed new optimization and estimation techniques to address two fundamental problems in machine learning. These new techniques serve as the basis for the Automatic Linguistic Indexing of Pictures - Real Time (ALIPR) system of fully automatic and high speed annotation for online pictures. In particular, the D2-clustering method, in the same spirit as k-means for vectors, is developed to group objects represented by bags of weighted vectors. Moreover, a generalized mixture modeling technique (kernel smoothing as a special case) for non-vector data is developed using the novel concept of Hypothetical Local Mapping (HLM). ALIPR has been tested by thousands of pictures from an Internet photo-sharing site, unrelated to the source of those pictures used in the training process. Its performance has also been studied at an online demo site where arbitrary users provide pictures of their choices and indicate the correctness of each annotation word. The experimental results show that a single computer processor can suggest annotation terms in real-time and with good accuracy. Jia Li 0001, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Real-World Image Annotation and Retrieval: An Introduction to the Special SectionabstractIndexing and retrieving large quantities of image data is an extremely challenging and increasingly topical problem for both industry and academia. Massive volumes of image data are all around us-in personal and commercial collections and on public Websites accessible via the Internet. According to a recent study by the market researcher IDC, digital camera sales rose 15 percent in 2006 to 105.7 million units worldwide. A four-year old online photo sharing website, Flickr, has more than 40 million monthly visitors and 2 billion photos uploaded; in fact, in a single day, a few million photos are uploaded. These developments have spurred enormous interest in digital images and a corresponding demand, both from the public and from industry, for better ways of cataloging, annotating, and accessing these data. This in turn has motivated researchers in pattern analysis and machine intelligence to address these tasks. Indeed, in a recent survey of the field of image annotation and retrieval, Wang et al. noticed an exponential growth over the last 10 years in the number of publications arising from researchers in computer vision, database management, machine learning, mathematical statistics, and signal and image processing. James Z. Wang 0001, Donald Geman, Jiebo Luo 0001, Robert M. Gray |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Automatic Extraction of Data from 2-D Plots in DocumentsabstractTwo-dimensional (2-D) plots in digital documents contain important information. Often, the results of scientific experiments and performance of businesses are summarized using plots. Although 2-D plots are easily understood by human users, current search engines rarely utilize the information contained in the plots to enhance the results returned in response to queries posed by end- users. We propose an automated algorithm for extracting information from line curves in 2-D plots. The extracted information can be stored in a database and indexed to answer end-user queries and enhance search results. We have collected 2-D plot images from a variety of resources and tested our extraction algorithms. Experimental evaluation has demonstrated that our method can produce results suitable for real world use. James Z. Wang 0001, Prasenjit Mitra 0001, C. Lee Giles |
ICDAR | 2 |
| 2007 | Tagging over time: real-world image annotation by lightweight meta-learningabstractAutomatic image annotation has been a hot-pursuit among multimedia researchers of late. Modest performance guarantees and limited adaptability often restrict its applicability to real-world settings. We propose tagging over time (T/T) to push the technology toward real-world applicability. Of particular interest are online systems that receive user-provided images and feedback over time, with user focus possibly changing and evolving. The T/T framework consists of a principled probabilistic approach to meta-learning, which acts as a go-between for a 'black-box' annotation system and the users. Inspired by inductive transfer, the approach attempts to harness available information, including the black-box model's performance, the image representations, and the WordNet ontology. Being computationally 'lightweight', this meta-learner efficiently re-trains over time, to improve and/or adapt to changes. The black-box annotation model is not required to be re-trained, allowing computationally intensive algorithms to be used. We experiment with standard image datasets and real-world data streams, using two existing annotation systems as black-boxes. Both batch and online annotation settings are experimented with. It is observed that the addition of this meta-learning layer produces much improved results that outperform best-known results. For the online setting, the T/T approach produces progressively better annotation with time, significantly outperforming the black-box as well as the static form of the meta-learner, on real-world data. Ritendra Datta, Dhiraj Joshi, Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 4 |
| 2007 | Learning the consensus on visual quality for next-generation image managementabstractWhile personal and community-based image collections grow by the day, the demand for novel photo management capabilities grows with it. Recent research has shown that it is possible to learn the consensus on visual quality measures such as aesthetics with a moderate degree of success. Here, we seek to push this performance to more realistic levels and use it to (a) help select high-quality pictures from collections, and (b) eliminate low-quality ones, introducing appropriate performance metrics in each case. To achieve this, we propose a sequential arrangement of a weighted linear least squares regressor and a naive Bayes' classifier, applied to a set of visual features previously found useful for quality prediction. Experiments on real-world data for these tasks show promising performance, with significant improvements over a previously proposed SVM-based method. Ritendra Datta, Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 3 |
| 2007 | Deriving knowledge from figures for digital librariesabstractFigures in digital documents contain important information. Current digital libraries do not summarize and index information available within figures for document retrieval. We present our system on automatic categorization of figures and extraction of data from 2-D plots. A machine-learning based method is used to categorize figures into a set of predefined types based on image features. An automated algorithm is designed to extract data values from solid line curves in 2-D plots. The semantic type of figures and extracted data values from 2-D plots can be integrated with textual information within documents to provide more effective document retrieval services for digital library users. Experimental evaluation has demonstrated that our system can produce results suitable for real-world use. James Z. Wang 0001, Prasenjit Mitra 0001, C. Lee Giles |
WWW | 2 |
| 2006 | QCHARM: A Novel Computational and Scientific Visualization Framework for Facilitating Discovery and Improving Diagnostic Reliability in Medicine
Brian A. Canada, Keith C. Cheng, James Z. Wang 0001 |
AMIA | 3 |
| 2006 | Studying Aesthetics in Photographic Images Using a Computational Approach
Ritendra Datta, Dhiraj Joshi, Jia Li 0001, James Z. Wang 0001 |
ECCV (3) | 4 |
| 2006 | Toward bridging the annotation-retrieval gap in image search by a generative modeling approachabstractWhile automatic image annotation remains an actively pursued research topic, enhancement of image search through its use has not been extensively explored. We propose an annotation-driven image retrieval approach and argue that under a number of different scenarios, this is very effective for semantically meaningful image search. In particular, our system is demonstrated to effectively handle cases of partially tagged and completely untagged image databases, multiple keyword queries, and example based queries with or without tags, all in near-realtime. Because our approach utilizes extra knowledge from a training dataset, it outperforms state-of-the-art visual similarity based retrieval techniques. For this purpose, a novel structure-composition model constructed from Beta distributions is developed to capture the spatial relationship among segmented regions of images. This model combined with the Gaussian mixture model produces scalable categorization of generic images. The categorization results are found to surpass previously reported results in speed and accuracy. Our novel annotation framework utilizes the categorization results to select tags based on term frequency, term saliency, and a WordNet-based measure of congruity, to boost salient tags while penalizing potentially unrelated ones. A bag of words distance measure based on WordNet is used to compute semantic similarity. The effectiveness of our approach is shown through extensive experiments. Ritendra Datta, Weina Ge, Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 4 |
| 2006 | Real-time computerized annotation of picturesabstractAutomated annotation of digital pictures has been a highly challenging problem for computer scientists since the invention of computers. The capability of annotating pictures by computers can lead to breakthroughs in a wide range of applications including Web image search, online picture-sharing communities, and scientific experiments. In our work, by advancing statistical modeling and optimization techniques, we can train computers about hundreds of semantic concepts using example pictures from each concept. The ALIPR (Automatic Linguistic Indexing of Pictures -Real Time)system of fully automatic and high speed annotation for online pictures has been constructed. Thousands of pictures from an Internet photo-sharing site, unrelated to the source of those pictures used in the training process, have been tested. The experimental results show that a single computer processor can suggest annotation terms in real-time and with good accuracy. Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 2 |
| 2006 | PARAgrab: A Comprehensive Architecture for Web Image Management and Multimodal Querying
Dhiraj Joshi, Ritendra Datta, Ziming Zhuang, W. P. Weiss, Marc Friedenberg, Jia Li 0001, James Z. Wang 0001 |
VLDB | 7 |
| 2006 | MILES: Multiple-Instance Learning via Embedded Instance SelectionabstractMultiple-instance problems arise from the situations where training class labels are attached to sets of samples (named bags), instead of individual samples within each bag (called instances). Most previous multiple-instance learning (MIL) algorithms are developed based on the assumption that a bag is positive if and only if at least one of its instances is positive. Although the assumption works well in a drug activity prediction problem, it is rather restrictive for other applications, especially those in the computer vision area. We propose a learning method, MILES (Multiple-Instance Learning via Embedded instance Selection), which converts the multiple-instance learning problem to a standard supervised learning problem that does not impose the assumption relating instance labels to bag labels. MILES maps each bag into a feature space defined by the instances in the training bags via an instance similarity measure. This feature mapping often provides a large number of redundant or irrelevant features. Hence, 1-norm SVM is applied to select important features as well as construct classifiers simultaneously. We have performed extensive experiments. In comparison with other methods, MILES demonstrates competitive classification accuracy, high computation efficiency, and robustness to labeling uncertainty. Yixin Chen 0002, Jinbo Bi, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | A Computationally Efficient Approach to the Estimation of Two- and Three-Dimensional Hidden Markov ModelsabstractStatistical modeling methods are becoming indispensable in today's large-scale image analysis. In this paper, we explore a computationally efficient parameter estimation algorithm for two-dimensional (2-D) and three-dimensional (3-D) hidden Markov models (HMMs) and show applications to satellite image segmentation. The proposed parameter estimation algorithm is compared with the first proposed algorithm for 2-D HMMs based on variable state Viterbi. We also propose a 3-D HMM for volume image modeling and apply it to volume image segmentation using a large number of synthetic images with ground truth. Experiments have demonstrated the computational efficiency of the proposed parameter estimation technique for 2-D HMMs and a potential of 3-D HMM as a stochastic modeling tool for volume images. Dhiraj Joshi, Jia Li 0001, James Z. Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2006 | The Story Picturing Engine - a system for automatic text illustrationabstractWe present an unsupervised approach to automated story picturing. Semantic keywords are extracted from the story, an annotated image database is searched. Thereafter, a novel image ranking scheme automatically determines the importance of each image. Both lexical annotations and visual content play a role in determining the ranks. Annotations are processed using the Wordnet. A mutual reinforcement-based rank is calculated for each image. We have implemented the methods in our Story Picturing Engine (SPE) system. Experiments on large-scale image databases are reported. A user study has been performed and statistical analysis of the results has been presented. Dhiraj Joshi, James Z. Wang 0001, Jia Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2005 | A Sparse Support Vector Machine Approach to Region-Based Image CategorizationabstractAutomatic image categorization using low-level features is a challenging research topic in computer vision. In this paper, we formulate the image categorization problem as a multiple-instance learning (MIL) problem by viewing an image as a bag of instances, each corresponding to a region obtained from image segmentation. We propose a new solution to the resulting MIL problem. Unlike many existing MIL approaches that rely on the diverse density framework, our approach performs an effective feature mapping through a chosen metric distance function. Thus the MIL problem becomes solvable by a regular classification algorithm. Sparse SVM is adopted to dramatically reduce the regions that are needed to classify images. The selected regions by a sparse SVM approximate to the target concepts in the traditional diverse density framework. The proposed approach is a lot more efficient in computation and less sensitive to the class label uncertainty. Experimental results are included to demonstrate the effectiveness and robustness of the proposed method. Jinbo Bi, Yixin Chen 0002, James Z. Wang 0001 |
CVPR (1) | 3 |
| 2005 | ALIP: The Automatic Linguistic Indexing of Pictures SystemabstractWe present the Automatic Linguistic Indexing of Pictures (ALIP) system. The system annotates images with linguistic terms, chosen among hundreds of such terms. The system uses a wavelet-based approach for feature extraction, a statistical modeling process for training, and a statistical significance processor to annotate images. We implemented and tested our ALIP system on a photographic image database of 600 different concepts, each with about 40 training images. The ALIP system has been used to annotate about 60,000 photographic images. In this demonstration, we illustrate the algorithms in the system and show the annotation results. With distributed computation, the annotation of an image can be provided in real-time. Jia Li 0001, James Z. Wang 0001 |
CVPR (2) | 2 |
| 2005 | Parameter estimation of multi-dimensional hidden Markov models - a scalable approachabstractParameter estimation is a key computational issue in all statistical image modeling techniques. In this paper, we explore a computationally efficient parameter estimation algorithm for multi-dimensional hidden Markov models. 2-D HMM has been applied to supervised aerial image classification and comparisons have been made with the first proposed estimation algorithm. An extensive parametric study has been performed with 3-D HMM and the scalability of the estimation algorithm has been discussed. Results show the great applicability of the explored algorithm to multi-dimensional HMM based image modeling applications. Dhiraj Joshi, Jia Li 0001, James Z. Wang 0001 |
ICIP (3) | 3 |
| 2005 | IMAGINATION: a robust image-based CAPTCHA generation systemabstractWe propose IMAGINATION (IMAge Generation for INternet AuthenticaTION), a system for the generation of attack-resistant, user-friendly, image-based CAPTCHAs. In our system, we produce controlled distortions on randomly chosen images and present them to the user for annotation from a given list of words. The distortions are performed in a way that satisfies the incongruous requirements of low perceptual degradation and high resistance to attack by content-based image retrieval systems. Word choices are carefully generated to avoid ambiguity as well as to avoid attacks based on the choices themselves. Preliminary results demonstrate the attack-resistance and user-friendliness of our system compared to text-based CAPTCHAs. Ritendra Datta, Jia Li 0001, James Z. Wang 0001 |
ACM Multimedia | 3 |
| 2005 | Multimedia Systems and Content-Based Image Retrieval. By Sagarmay Deb, Idea Group Publishing, 2004, $79.95 ISBN 1-59140-156-9
Dhiraj Joshi, James Z. Wang 0001 |
Inf. Process. Manag. | 2 |
| 2005 | CLUE: cluster-based retrieval of images by unsupervised learningabstractIn a typical content-based image retrieval (CBIR) system, target images (images in the database) are sorted by feature similarities with respect to the query. Similarities among target images are usually ignored. This paper introduces a new technique, cluster-based retrieval of images by unsupervised learning (CLUE), for improving user interaction with image retrieval systems by fully exploiting the similarity information. CLUE retrieves image clusters by applying a graph-theoretic clustering algorithm to a collection of images in the vicinity of the query. Clustering in CLUE is dynamic. In particular, clusters formed depend on which images are retrieved in response to the query. CLUE can be combined with any real-valued symmetric similarity measure (metric or nonmetric). Thus, it may be embedded in many current CBIR systems, including relevance feedback systems. The performance of an experimental image retrieval system using CLUE is evaluated on a database of around 60,000 images from COREL. Empirical results demonstrate improved performance compared with a CBIR system using the same image similarity measure. In addition, results on images returned by Google's Image Search reveal the potential of applying CLUE to real-world image data and integrating CLUE as a part of the interface for keyword-based image retrieval systems. Yixin Chen 0002, James Z. Wang 0001, Robert Krovetz |
IEEE Trans. Image Process. | 2 |
| 2004 | Stochastic modeling of volume images with a 3-d hidden markov modelabstractOver the years, researchers in the image analysis community have successfully used various statistical modeling methods to segment, classify and annotate digital images. In this paper, we propose a 3-D hidden Markov model (HMM) for volume image modeling. A computationally efficient algorithm is developed to estimate the model. The 3-D HMM is applied to volume image segmentation and tested using synthetic images with ground truth. Experiments have demonstrated that 3-D HMM outperforms Gaussian mixture model based clustering by an order of magnitude in accuracy. Jia Li 0001, Dhiraj Joshi, James Z. Wang 0001 |
ICIP | 3 |
| 2004 | Image Categorization by Learning and Reasoning with Regions
Yixin Chen 0002, James Z. Wang 0001 |
J. Mach. Learn. Res. | 2 |
| 2004 | Studying digital imagery of ancient paintings by mixtures of stochastic modelsabstractThis paper addresses learning-based characterization of fine art painting styles. The research has the potential to provide a powerful tool to art historians for studying connections among artists or periods in the history of art. Depending on specific applications, paintings can be categorized in different ways. In this paper, we focus on comparing the painting styles of artists. To profile the style of an artist, a mixture of stochastic models is estimated using training images. The two-dimensional (2-D) multiresolution hidden Markov model (MHMM) is used in the experiment. These models form an artist's distinct digital signature. For certain types of paintings, only strokes provide reliable information to distinguish artists. Chinese ink paintings are a prime example of the above phenomenon; they do not have colors or even tones. The 2-D MHMM analyzes relatively large regions in an image, which in turn makes it more likely to capture properties of the painting strokes. The mixtures of 2-D MHMMs established for artists can be further used to classify paintings and compare paintings or artists. We implemented and tested the system using high-resolution digital photographs of some of China's most renowned artists. Experiments have demonstrated good potential of our approach in automatic analysis of paintings. Our work can be applied to other domains. Jia Li 0001, James Z. Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2003 | Kernel machines and additive fuzzy systems: classification and function approximationabstractThis paper investigates the connection between additive fuzzy systems and kernel machines. We prove that, under quite general conditions, these two seemingly quite distinct models are essentially equivalent. As a result, algorithms based upon support vector (SV) learning are proposed to build fuzzy systems for classification and function approximation. The performance of the proposed algorithm is illustrated using extensive experimental results. Yixin Chen 0002, James Z. Wang 0001 |
FUZZ-IEEE | 2 |
| 2003 | Looking beyond region boundaries: a robust image similarity measure using fuzzified region featuresabstractThe performance of most region-based image retrieval systems depend critically on the accuracy of object segmentation. We propose a region matching approach, unified feature matching (UFM), which greatly increases the robustness of the retrieval system against segmentation related uncertainties. In our retrieval system, an image is represented by a set of segmented regions each of which is characterized by a fuzzy feature reflecting color, texture, and shape properties. The resemblance between two images is then defined as the overall similarity between two families of fuzzy features, and quantified by the UFM measure. The system has been tested on a database of about 60,000 general-purpose images. Experimental results demonstrate improved accuracy and robustness. Yixin Chen 0002, James Z. Wang 0001 |
FUZZ-IEEE | 2 |
| 2003 | Evaluation strategies for automatic linguistic indexing of picturesabstractWith the rapid technological advances in machine learning and data mining, it is now possible to train computers with hundreds of semantic concepts for the purpose of annotating images automatically using keywords and textual descriptions. We have developed a system, the automatic linguistic indexing of pictures (ALIP) system, using a 2-D multiresolution hidden Markov model. The evaluation of such approaches opens up challenges and interesting research questions. The goals of linguistic indexing are often different from those of other fields including image retrieval, image classification, and computer vision. In many application domains, computer programs that can provide semantically relevant keyword annotations are desired, even if the predicted annotations are different from those of the gold standard. In this paper, we discuss evaluation strategies for automatic linguistic indexing of pictures. We provide both objective and subjective evaluation methods. Finally, we report experimental results using our ALIP system. James Z. Wang 0001, Jia Li 0001, Sui Ching Lin |
ICIP (3) | 1 |
| 2003 | Automatic Linguistic Indexing of Pictures by a Statistical Modeling ApproachabstractAutomatic linguistic indexing of pictures is an important but highly challenging problem for researchers in computer vision and content-based image retrieval. In this paper, we introduce a statistical modeling approach to this problem. Categorized images are used to train a dictionary of hundreds of statistical models each representing a concept. Images of any given concept are regarded as instances of a stochastic process that characterizes the concept. To measure the extent of association between an image and the textual description of a concept, the likelihood of the occurrence of the image based on the characterizing stochastic process is computed. A high likelihood indicates a strong association. In our experimental implementation, we focus on a particular group of stochastic processes, that is, the two-dimensional multiresolution hidden Markov models (2D MHMMs). We implemented and tested our ALIP (Automatic Linguistic Indexing of Pictures) system on a photographic image database of 600 different concepts, each with about 40 training images. The system is evaluated quantitatively using more than 4,600 images outside the training database and compared with a random annotation scheme. Experiments have demonstrated the good accuracy of the system and its high potential in linguistic indexing of photographic images. Jia Li 0001, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Support vector learning for fuzzy rule-based classification systemsabstractTo design a fuzzy rule-based classification system (fuzzy classifier) with good generalization ability in a high dimensional feature space has been an active research topic for a long time. As a powerful machine learning approach for pattern recognition problems, the support vector machine (SVM) is known to have good generalization ability. More importantly, an SVM can work very well on a high- (or even infinite) dimensional feature space. This paper investigates the connection between fuzzy classifiers and kernel machines, establishes a link between fuzzy rules and kernels, and proposes a learning algorithm for fuzzy classifiers. We first show that a fuzzy classifier implicitly defines a translation invariant kernel under the assumption that all membership functions associated with the same input variable are generated from location transformation of a reference function. Fuzzy inference on the IF-part of a fuzzy rule can be viewed as evaluating the kernel function. The kernel function is then proven to be a Mercer kernel if the reference functions meet a certain spectral requirement. The corresponding fuzzy classifier is named positive definite fuzzy classifier (PDFC). A PDFC can be built from the given training samples based on a support vector learning approach with the IF-part fuzzy rules given by the support vectors. Since the learning process minimizes an upper bound on the expected risk (expected prediction error) instead of the empirical risk (training error), the resulting PDFC usually has good generalization. Moreover, because of the sparsity properties of the SVMs, the number of fuzzy rules is irrelevant to the dimension of input space. In this sense, we avoid the "curse of dimensionality." Finally, PDFCs with different reference functions are constructed using the support vector learning approach. The performance of the PDFCs is illustrated by extensive experimental results. Comparisons with other methods are also provided. Yixin Chen 0002, James Z. Wang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2002 | Progressive display of very high resolution images using wavelets
Ya Zhang 0002, James Z. Wang 0001 |
AMIA | 2 |
| 2002 | Learning-based linguistic indexing of pictures with 2--d MHMMsabstractAutomatic linguistic indexing of pictures is an important but highly challenging problem for researchers in computer vision and content-based image retrieval. In this paper, we introduce a statistical modeling approach to this problem. Categorized images are used to train a dictionary of hundreds of concepts automatically based on statistical modeling. Images of any given concept category are regarded as instances of a stochastic process that characterizes the category. To measure the extent of association between an image and the textual description of a category of images, the likelihood of the occurrence of the image based on the stochastic process derived from the category is computed. A high likelihood indicates a strong association. In our experimental implementation, the ALIP (Automatic Linguistic Indexing of Pictures) system, we focus on a particular group of stochastic processes for describing images, that is, the two-dimensional multiresolution hidden Markov models (2-D MHMMs). We implemented and tested the system on a photographic image database of 600 different semantic categories, each with about 40 training images. Tested using 3,000 images outside the training database, the system has demonstrated good accuracy and high potential in linguistic indexing of these test images. James Z. Wang 0001, Jia Li 0001 |
ACM Multimedia | 1 |
| 2002 | SST: an algorithm for finding near-exact sequence matches in time proportional to the logarithm of the database sizeabstractMOTIVATION: Searches for near exact sequence matches are performed frequently in large-scale sequencing projects and in comparative genomics. The time and cost of performing these large-scale sequence-similarity searches is prohibitive using even the fastest of the extant algorithms. Faster algorithms are desired. RESULTS: We have developed an algorithm, called SST (Sequence Search Tree), that searches a database of DNA sequences for near-exact matches, in time proportional to the logarithm of the database size n. In SST, we partition each sequence into fragments of fixed length called 'windows' using multiple offsets. Each window is mapped into a vector of dimension 4(k) which contains the frequency of occurrence of its component k-tuples, with k a parameter typically in the range 4-6. Then we create a tree-structured index of the windows in vector space, with tree-structured vector quantization (TSVQ). We identify the nearest neighbors of a query sequence by partitioning the query into windows and searching the tree-structured index for nearest-neighbor windows in the database. When the tree is balanced this yields an O(logn) complexity for the search. This complexity was observed in our computations. SST is most effective for applications in which the target sequences show a high degree of similarity to the query sequence, such as assembling shotgun sequences or matching ESTs to genomic sequence. The algorithm is also an effective filtration method. Specifically, it can be used as a preprocessing step for other search methods to reduce the complexity of searching one large database against another. For the problem of identifying overlapping fragments in the assembly of 120 000 fragments from a 1.5 megabase genomic sequence, SST is 15 times faster than BLAST when we consider both building and searching the tree. For searching alone (i.e. after building the tree index), SST 27 times faster than BLAST. AVAILABILITY: Request from the authors. Eldar Giladi, Michael G. Walker 0001, James Z. Wang 0001, Wayne Volkmuth |
Bioinform. | 3 |
| 2002 | A Region-Based Fuzzy Feature Matching Approach to Content-Based Image RetrievalabstractThis paper proposes a fuzzy logic approach, UFM (unified feature matching), for region-based image retrieval. In our retrieval system, an image is represented by a set of segmented regions, each of which is characterized by a fuzzy feature (fuzzy set) reflecting color, texture, and shape properties. As a result, an image is associated with a family of fuzzy features corresponding to regions. Fuzzy features naturally characterize the gradual transition between regions (blurry boundaries) within an image and incorporate the segmentation-related uncertainties into the retrieval algorithm. The resemblance of two images is then defined as the overall similarity between two families of fuzzy features and quantified by a similarity measure, UFM measure, which integrates properties of all the regions in the images. Compared with similarity measures based on individual regions and on all regions with crisp-valued feature representations, the UFM measure greatly reduces the influence of inaccurate segmentation and provides a very intuitive quantification. The UFM has been implemented as a part of our experimental SIMPLIcity image retrieval system. The performance of the system is illustrated using examples from an image database of about 60,000 general-purpose images. Yixin Chen 0002, James Z. Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | A scalable integrated region-based image retrieval systemabstractWe present a scalable algorithm for indexing and retrieving images based on region segmentation. The method uses statistical clustering on region features and IRM (integrated region matching), a measure developed to evaluate overall similarity between images. It incorporates properties of all the regions in the images by a region-matching scheme. The algorithm has been implemented as a part of our experimental SIMPLIcity (Semantics-sensitive Integrated Matching for Picture LIbraries) image retrieval system and tested on large-scale image databases of both general-purpose images and pathology slides. Experiments have demonstrated that this technique maintains the accuracy of the original system while reducing the matching time significantly. Yanping Du, James Z. Wang 0001 |
ICIP (1) | 2 |
| 2001 | FIRM: fuzzily integrated region matching for content-based image retrievalabstractWe propose FIRM (Fuzzily Integrated Region Matching), an efficient and robust similarity measure for region-based image retrieval. Each image in our retrieval system is represented by a set of regions that are characterized by fuzzy sets. The FIRM measure, representing the overall similarity between two images, is defined as the similarity between two families of fuzzy sets. Compared with similarity measures based on individual regions and on all regions with crisp feature representations, our approach greatly reduces the influence of inaccurate segmentation. Experimental results based on a database of about 200,000 general-purpose images demonstrate improved accuracy, robustness, and high speed. Yixin Chen 0002, James Z. Wang 0001, Jia Li 0001 |
ACM Multimedia | 2 |
| 2001 | Wavelets and Imaging Informatics: A Review of the Literature
James Z. Wang 0001 |
J. Biomed. Informatics | 1 |
| 2001 | Unsupervised Multiresolution Segmentation for Images with Low Depth of FieldabstractUnsupervised segmentation of images with low depth of field (DOF) is highly useful in various applications. This paper describes a novel multiresolution image segmentation algorithm for low DOF images. The algorithm is designed to separate a sharply focused object-of-interest from other foreground or background objects. The algorithm is fully automatic in that all parameters are image independent. A multi-scale approach based on high frequency wavelet coefficients and their statistics is used to perform context-dependent classification of individual blocks of the image. Unlike other edge-based approaches, our algorithm does not rely on the process of connecting object boundaries. The algorithm has achieved high accuracy when tested on more than 100 low DOF images, many with inhomogeneous foreground or background distractions. Compared with he state of the art algorithms, this new algorithm provides better accuracy at higher speed. James Z. Wang 0001, Jia Li 0001, Robert M. Gray, Gio Wiederhold |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | SIMPLIcity: Semantics-Sensitive Integrated Matching for Picture LIbrariesabstractWe present here SIMPLIcity (semantics-sensitive integrated matching for picture libraries), an image retrieval system, which uses semantics classification methods, a wavelet-based approach for feature extraction, and integrated region matching based upon image segmentation. An image is represented by a set of regions, roughly corresponding to objects, which are characterized by color, texture, shape, and location. The system classifies images into semantic categories. Potentially, the categorization enhances retrieval by permitting semantically-adaptive searching methods and narrowing down the searching range in a database. A measure for the overall similarity between images is developed using a region-matching scheme that integrates properties of all the regions in the images. The application of SIMPLIcity to several databases has demonstrated that our system performs significantly better and faster than existing ones. The system is fairly robust to image alterations. James Z. Wang 0001, Jia Li 0001, Gio Wiederhold |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Pathfinder: multiresolution region-based searching of pathology images using IRM
James Z. Wang 0001 |
AMIA | 1 |
| 2000 | Classification of Textured and Non-Textured Images Using Region SegmentationabstractThe classification of general-purpose photographs into textured and non-textured images is critical for developing accurate content-based image retrieval systems for large-scale image databases. With the accurate detection of textured images, we may retrieve images based on features tailored for the corresponding image type. In this paper, we present an algorithm to classify a photographic image as textured or non-textured using region segmentation and statistical testing. The application of the system to a database of about 60,000 general-purpose images shows much improved accuracy in retrieval. Jia Li 0001, James Z. Wang 0001, Gio Wiederhold |
ICIP | 2 |
| 2000 | IRM: integrated region matching for image retrievalabstractContent-based image retrieval using region segmentation has been an active research area. We present IRM (Integrated Region Matching), a novel similarity measure for region-based image similarity comparison. The targeted image retrieval systems represent an image by a set of regions, roughly corresponding to objects, which are characterized by features reflecting color, texture, shape, and location properties. The IRM measure for evaluating overall similarity between images incorporates properties of all the regions in the images by a region-matching scheme. Compared with retrieval based on individual regions, the overall similarity approach reduces the influence of inaccurate segmentation, helps to clarify the semantics of a particular region, and enables a simple querying interface for region-based image retrieval systems. The IRM has been implemented as a part of our experimental SIMPLIcity image retrieval system. The application to a database of about 200,000 general-purpose images shows exceptional robustness to image alterations such as intensity variation, sharpness variation, color distortions, shape distortions, cropping, shifting, and rotation. Compared with several existing systems, our system in general achieves more accurate retrieval at higher speed. Jia Li 0001, James Z. Wang 0001, Gio Wiederhold |
ACM Multimedia | 2 |
| 2000 | SIMPLIcity: a region-based retrieval system for picture libraries and biomedical image databases
James Z. Wang 0001 |
ACM Multimedia | 1 |
| 2000 | Region-based retrieval of biomedical images
James Z. Wang 0001 |
ACM Multimedia | 1 |
| 1998 | System for efficient and secure distribution of medical images on the Internet
James Z. Wang 0001, Gio Wiederhold |
AMIA | 1 |
| 1998 | System for screening objectionable imagesabstractAs computers and the Internet become more and more available to families, access of objectionable graphics by children is increasingly a problem that many parents are concerned about. This paper describes WIPE™ (Wavelet Image Pornography Elimination), a system capable of classifying an image as objectionable or benign. The algorithm uses a combination of an icon filter, a graph–photo detector, a color histogram filter, a texture filter and a wavelet-based shape matching algorithm to provide robust screening of on-line objectionable images. Semantically-meaningful feature vector matching is carried out so that comparisons between a given on-line image and images in a pre-marked training data set can be performed efficiently and effectively. The system is practical for real-world applications, processing queries at a speed of less than 2 s each, including the time taken to compute the feature vector for the query, on a Pentium Pro PC. Besides its exceptional speed, it has demonstrated 96% sensitivity over a test set of 1076 digital photographs found on objectionable news groups. It wrongly classified 9% of a set of 10,809 benign photographs obtained from various sources. The specificity in real-world applications is expected to be much higher because benign on-line graphs can be filtered out with our graph–photo detector with 100% sensitivity and nearly 100% specificity, and surrounding text can be used to assist the classification process. James Z. Wang 0001, Jia Li 0001, Gio Wiederhold, Oscar Firschein |
Comput. Commun. | 1 |
| 1997 | A Textual Information Detection and Elimination System for Secure Medical Image Distribution
James Z. Wang 0001, Michel Bilello, Gio Wiederhold |
AMIA | 1 |