Yan-Ying Chen

dblp:66/9376 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
6since 2021 · last 2026
0009-0000-5901-1538ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Generative modeling · 62% Language models and text generation · 12% Face, body and person analysis · 8%
Human-computer interaction and pervasive computing
5 papers
Human-AI interaction · 68% User interface design and tools · 18% Learning and educational technologies · 8%
Databases, data mining, and information retrieval
12 papers
Information retrieval · 55% Web and social media mining · 24% Recommender systems · 21%
Computer graphics and multimedia
8 papers
Multimedia analysis and retrieval · 73% Visualization and visual analytics · 27%

Topics — the 30 heaviest of 56, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.922026
ShaLa: Multimodal Shared Latent Generative Modelling · AAAI 2026
Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder · ICCV 2025
Machine learning › Generative modeling
variational autoencoder
1.322026
ShaLa: Multimodal Shared Latent Generative Modelling · AAAI 2026
Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder · ICCV 2025
Machine learning › Generative modeling › multimodal generation
multimodal generative model
1.012026
ShaLa: Multimodal Shared Latent Generative Modelling · AAAI 2026
Machine learning › Generative modeling › variational autoencoder
multimodal variational autoencoder
1.012026
ShaLa: Multimodal Shared Latent Generative Modelling · AAAI 2026
Machine learning › Generative modeling › diffusion model › controllable generation
attribute control
0.912025
Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder · ICCV 2025
Human-AI interaction
co-creative design
0.912025
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching · CHI 2025
User interface design and tools
creativity support tools
0.912025
BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer · CHI 2025
Human-AI interaction › AI-assisted creativity
LLM-assisted design
0.912025
BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer · CHI 2025
Human-AI interaction › generative AI
text-to-image generation
0.912025
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching · CHI 2025
Visualization and visual analytics › visual analytics
visual analytics for machine learning
0.812024
VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-Making · UIST 2024
Human-AI interaction › explainable AI
explainable AI interfaces
0.812024
VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-Making · UIST 2024
Information retrieval › image retrieval › object retrieval
face image retrieval
0.542015
Scalable Face Image Retrieval Using Attribute-Enhanced Sparse Codewords · IEEE Trans. Multim. 2013
Where is who: large-scale photo retrieval by facial attributes and canvas layout · SIGIR 2012
Semi-supervised face image retrieval using sparse coding with identity constraint · ACM Multimedia 2011
Computer vision › Face, body and person analysis
facial attribute analysis
0.422015
Visually Interpreting Names as Demographic Attributes by Exploiting Click-Through Data · AAAI 2015
Facial Attribute Space Compression by Latent Human Topic Discovery · ACM Multimedia 2014
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Natural language and speech › Language models and text generation › text summarization
title generation
0.412019
Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation · ACL (1) 2019
Learning and educational technologies › student knowledge modeling
knowledge tracing
0.412019
Augmenting Knowledge Tracing by Considering Forgetting Behavior · WWW 2019
Natural language and speech › Language models and text generation › text summarization
extractive summarization
0.312018
Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations · EMNLP 2018
Natural language and speech › Language models and text generation
text summarization
0.312018
Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations · EMNLP 2018
Robotics › Robot navigation and mapping › localization
vision-based localization
0.312018
ContextualNet: Exploiting Contextual Information Using LSTMs to Improve Image-Based Localization · ICRA 2018
Computer vision › 3D vision
visual localization
0.312018
ContextualNet: Exploiting Contextual Information Using LSTMs to Improve Image-Based Localization · ICRA 2018
Recommender systems › domain-specific recommendation
travel recommendation
0.322013
Travel Recommendation by Mining People Attributes and Travel Group Types From Community-Contributed Photos · IEEE Trans. Multim. 2013
Personalized travel recommendation by mining people attributes from community-contributed photos · ACM Multimedia 2011
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition
0.322012
People search and activity mining in large-scale community-contributed photos · ACM Multimedia 2012
Photo search by face positions and facial attributes on touch devices · ACM Multimedia 2011
Information retrieval
image retrieval
0.322012
Where is who: large-scale photo retrieval by facial attributes and canvas layout · SIGIR 2012
Semi-supervised face image retrieval using sparse coding with identity constraint · ACM Multimedia 2011
Machine learning › Transfer learning and domain adaptation › knowledge transfer
analogical transfer
0.312025
BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer · CHI 2025
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder
0.312025
Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder · ICCV 2025
Interaction techniques and input
sketching
0.312025
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching · CHI 2025
Machine learning › Trustworthy machine learning
interpretability
0.212024
VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-Making · UIST 2024
Recommender systems
user profiling
0.212015
Visually Interpreting Names as Demographic Attributes by Exploiting Click-Through Data · AAAI 2015

Methods — techniques the papers use, named apart from their topics

what-if analysis · 2.3visual interactive model exploration · 2.3tree-of-life approach · 1.7large language model · 1.7variational inference · 1.0multimodal VAE · 1.0diffusion prior · 1.0within-subjects study · 0.9text-to-image models · 0.9cross-attention · 0.9conditional variational autoencoder · 0.9adapter · 0.9deep knowledge tracing · 0.4sentiment analysis · 0.4hierarchy cross-media learning · 0.4distant labeling · 0.3disjunctive model · 0.3sparse coding · 0.3
YearPublicationVenuePosition
2026 ShaLa: Multimodal Shared Latent Generative Modelling
abstract
This paper presents a novel generative framework for learning shared latent representations across multimodal data. Many advanced multimodal methods focus on capturing all combinations of modality-specific details across inputs, which can inadvertently obscure the high-level semantic concepts that are shared across modalities. Notably, Multimodal VAEs with low-dimensional latent variables are designed to capture shared representations, enabling various tasks such as joint multimodal synthesis and cross-modal inference. However, multimodal VAEs often struggle to design expressive joint variational posteriors and suffer from low-quality synthesis. In this work, ShaLa addresses these challenges by integrating a novel architectural inference model and a second-stage expressive diffusion prior, which not only facilitates effective inference of shared latent representation but also significantly improves the quality of downstream multimodal synthesis. We validate ShaLa extensively across multiple benchmarks, demonstrating superior coherence and synthesis quality compared to state-of-the-art multimodal VAEs. Furthermore, ShaLa scales to many more modalities while prior multimodal VAEs have fallen short in capturing the increasing complexity of the shared latent space.
Jiali Cui, Yan-Ying Chen, Matthew Klenk 0001
AAAI2
2025 BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer
abstract
We present BioSpark, a system for analogical innovation designed to act as a creativity partner in reducing the cognitive effort in finding, mapping, and creatively adapting diverse inspirations. While prior approaches have focused on initial stages of finding inspirations, BioSpark uses LLMs embedded in a familiar, visual, Pinterest-like interface to go beyond inspiration to supporting users in identifying the key solution mechanisms, transferring them to the problem domain, considering tradeoffs, and elaborating on details and characteristics. To accomplish this BioSpark introduces several novel contributions, including a tree-of-life enabled approach for generating relevant and diverse inspirations, as well as AI-powered cards including 'Sparks' for analogical transfer; 'Trade-offs' for considering pros and cons; and 'Q&A' for deeper elaboration. We evaluated BioSpark through workshops with professional designers and a controlled user study, finding that using BioSpark led to a greater number of generated ideas; those ideas being rated higher in creative quality; and more diversity in terms of biological inspirations used than a control condition. Our results suggest new avenues for creativity support tools embedding AI in familiar interaction paradigms for designer workflows.
Hyeonsu B. Kang, David Chuan-En Lin, Yan-Ying Chen, Matthew K. Hong, Nikolas Martelaro, Aniket Kittur
CHI3
2025 Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching
abstract
With recent advancements in the capabilities of Text-to-Image (T2I) AI models, product designers have begun experimenting with them in their work. However, T2I models struggle to interpret abstract language and the current user experience of T2I tools can induce design fixation rather than a more iterative, exploratory process. To address these challenges, we developed Inkspire, a sketch-driven tool that supports designers in prototyping product design concepts with analogical inspirations and a complete sketch-to-design-to-sketch feedback loop. To inform the design of Inkspire, we conducted an exchange session with designers and distilled design goals for improving T2I interactions. In a within-subjects study comparing Inkspire to ControlNet, we found that Inkspire supported designers with more inspiration and exploration of design ideas, and improved aspects of the co-creative process by allowing designers to effectively grasp the current state of the AI to guide it towards novel design intentions.
David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro, Aniket Kittur, Yan-Ying Chen, Matthew K. Hong
CHI5
2025 Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational Autoencoder
abstract
Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of continuous attributes, especially multiple attributes simultaneously, in a new domain (e.g., numeric values like eye openness or car width) with text-only guidance remains a significant challenge. To address this, we introduce the Attribute (Att) Adapter, a novel plug-and-play module designed to enable fine-grained, multi-attributes control in pretrained diffusion models. Our approach learns a single control adapter from a set of sample images that can be unpaired and contain multiple visual attributes. The Att-Adapter leverages the decoupled cross attention module to naturally harmonize the multiple domain attributes with text conditioning. We further introduce Conditional Variational Autoencoder (CVAE) to the Att-Adapter to mitigate overfitting, matching the diverse nature of the visual world. Evaluations on two public datasets show that Att-Adapter outperforms all LoRA-based baselines in controlling continuous attributes. Additionally, our method enables a broader control range and also improves disentanglement across multiple attributes, surpassing StyleGAN-based techniques. Notably, Att-Adapter is flexible, requiring no paired synthetic data for training, and is easily scalable to multiple attributes within a single model.
Wonwoong Cho, Yan-Ying Chen, Matthew Klenk 0001, David I. Inouye
ICCV2
2024 VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-Making
abstract
Ensuring that Machine Learning (ML) models make correct and meaningful inferences is necessary for the broader adoption of such models into high-stakes decision-making scenarios. Thus, ML model engineers increasingly use eXplainable AI (XAI) tools to investigate the capabilities and limitations of their ML models before deployment. However, explaining sequential ML models, which make a series of decisions at each timestep, remains challenging. We present Visual Interactive Model Explorer (VIME), an XAI toolbox that enables ML model engineers to explain decisions of sequential models in different “what-if” scenarios. Our evaluation with 14 ML experts, who investigated two existing sequential ML models using VIME and a baseline XAI toolbox to explore “what-if” scenarios, showed that VIME made it easier to identify and explain instances when the models made wrong decisions compared to the baseline. Our work informs the design of future interactive XAI mechanisms for evaluating sequential ML-based decision support systems.
Anindya Das Antar, Somayeh Molaei, Yan-Ying Chen, Matthew L. Lee, Nikola Banovic 0001
UIST3
2023 Machine learning-based measure of cognitive complexity explains variance in rank-ordered preference
Shabnam Hakimi, Yan-Ying Chen, Monica P. Van, Scott A. Carter, Emily S. Sumner, Nayeli Bravo, Kalani Murakami, Charlene C. Wu, Matthew Klenk 0001
CogSci2
2020 Thoracic Disease Identification and Localization using Distance Learning and Region Verification
Cheng Zhang 0014, Francine Chen 0001, Yan-Ying Chen
BMVC3
2020 Tackling challenges of neural purchase stage identification from imbalanced twitter data
abstract
Abstract Twitter and other social media platforms are often used for sharing interest in products. The identification of purchase decision stages, such as in the AIDA model (Awareness, Interest, Desire, and Action), can enable more personalized e-commerce services and a finer-grained targeting of advertisements than predicting purchase intent only. In this paper, we propose and analyze neural models for identifying the purchase stage of single tweets in a user’s tweet sequence. In particular, we identify three challenges of purchase stage identification: imbalanced label distribution with a high number of non-purchase-stage instances, limited amount of training data, and domain adaptation with no or only little target domain data. Our experiments reveal that the imbalanced label distribution is the main challenge for our models. We address it with ranking loss and perform detailed investigations of the performance of our models on the different output classes. In order to improve the generalization of the models and augment the limited amount of training data, we examine the use of sentiment analysis as a complementary, secondary task in a multitask framework. For applying our models to tweets from another product domain, we consider two scenarios: for the first scenario without any labeled data in the target product domain, we show that learning domain-invariant representations with adversarial training is most promising, while for the second scenario with a small number of labeled target examples, fine-tuning the source model weights performs best. Finally, we conduct several analyses, including extracting attention weights and representative phrases for the different purchase stages. The results suggest that the model is learning features indicative of purchase stages and that the confusion errors are sensible.
Heike Adel, Francine Chen 0001, Yan-Ying Chen
Nat. Lang. Eng.3
2019 Adversarial Domain Adaptation Using Artificial Titles for Abstractive Title Generation
abstract
A common issue in training a deep learning, abstractive summarization model is lack of a large set of training summaries.This paper examines techniques for adapting from a labeled source domain to an unlabeled target domain in the context of an encoder-decoder model for text generation.In addition to adversarial domain adaptation (ADA), we introduce the use of artificial titles and sequential training to capture the grammatical style of the unlabeled target domain.Evaluation on adapting to/from news articles and Stack Exchange posts indicates that the use of these techniques can boost performance for both unsupervised adaptation as well as fine-tuning with limited target data.
Francine Chen 0001, Yan-Ying Chen
ACL (1)2
2019 Addressing Data Bias Problems for Chest X-ray Image Report Generation
Philipp Harzig, Yan-Ying Chen, Francine Chen 0001, Rainer Lienhart
BMVC2
2019 Sensory Media Association through Reciprocating Training
abstract
Machine learning achieved great progress in recent years. However, state-of-the-art machine learning systems are still far behind biological learning systems on learning directly from sensors without offline labeling. This paper proposes an approach for automating machine learning from multi-modal sensors. In this learning setup, the system has no access to any human labeling tool which is not available to a biological learning system such as a dog or a newborn baby. We tested the learning proposal with audiovisual data. The testing system contains two deep autoencoders, one for learning speech representations and another for learning image representations. Two deep networks are trained to bridge the latent spaces of two autoencoders, yielding representation mappings for both speech-to-image and image-to-speech. To improve feature clustering in both latent spaces, the system alternately uses one modality to guide the learning of another modality. Different from traditional technology that uses a fixed modality for supervision (e.g. using text labels for image classification), the proposed approach facilitates a machine to learn from sensory inputs of two or more modalities through alternating guidance among these modalities. We evaluate the proposed model with MNIST digit images and corresponding digit speeches in the Google Command Digit Dataset (GCDD) and got very promising results.
Qiong Liu 0003, Ray Yuan, Yan-Ying Chen
ISM5
2019 Augmenting Knowledge Tracing by Considering Forgetting Behavior
abstract
Computer-aided education systems are now seeking to provide each student with personalized materials based on a student's individual knowledge. To provide suitable learning materials, tracing each student's knowledge over a period of time is important. However, predicting each student's knowledge is difficult because students tend to forget. The forgetting behavior is mainly because of two reasons: the lag time from the previous interaction, and the number of past trials on a question. Although there are a few studies that consider forgetting while modeling a student's knowledge, some models consider only partial information about forgetting, whereas others consider multiple features about forgetting, ignoring a student's learning sequence. In this paper, we focus on modeling and predicting a student's knowledge by considering their forgetting behavior. We extend the deep knowledge tracing model [17], which is a state-of-the-art sequential model for knowledge tracing, to consider forgetting by incorporating multiple types of information related to forgetting. Experiments on knowledge tracing datasets show that our proposed model improves the predictive performance as compared to baselines. Moreover, we also examine that the combination of multiple types of information that affect the behavior of forgetting results in performance improvement.
Koki Nagatani, Qian Zhang 0061, Masahiro Sato, Yan-Ying Chen, Francine Chen 0001, Tomoko Ohkuma
WWW4
2018 Harnessing Popularity in Social Media for Extractive Summarization of Online Conversations
abstract
We leverage a popularity measure in social media as a distant label for extractive summarization of online conversations.In social media, users can vote, share, or bookmark a post they prefer.The number of these actions is regarded as a measure of popularity.However, popularity is not determined solely by content of a post, e.g., a text or an image it contains, but is highly based on its contexts, e.g., timing, and authority.We propose Disjunctive model that computes the contribution of content and context separately.For evaluation, we build a dataset where the informativeness of comments is annotated.We evaluate the results with ranking metrics, and show that our model outperforms the baseline models which directly use popularity as a measure of informativeness.
Ryuji Kano, Yasuhide Miura, Motoki Taniguchi, Yan-Ying Chen, Francine Chen 0001, Tomoko Ohkuma
EMNLP4
2018 ContextualNet: Exploiting Contextual Information Using LSTMs to Improve Image-Based Localization
abstract
Convolutional Neural Networks (CNN) have successfully been utilized for localization using a single monocular image [1]. Most of the work to date has either focused on reducing the dimensionality of data for better learning of parameters during training or on developing different variations of CNN models to improve pose estimation. Many of the best performing works solely consider the content in a single image, while the context from historical images is ignored. In this paper, we propose a combined CNN-LSTM which is capable of incorporating contextual information from historical images to better estimate the current pose. Experimental results achieved using a dataset collected in an indoor office space improved the overall system results to 0.8 m & 2.5° at the third quartile of the cumulative distribution as compared with 1.5 m & 3.0° achieved by PoseNet [1]. Furthermore, we demonstrate how the temporal information exploited by the CNN-LSTM model assists in localizing the robot in situations where image content does not have sufficient features.
Brendan Emery, Yan-Ying Chen
ICRA3
2018 Learning to Disentangle Interleaved Conversational Threads with a Siamese Hierarchical Network and Similarity Ranking
abstract
Jyun-Yu Jiang, Francine Chen, Yan-Ying Chen, Wei Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Jyun-Yu Jiang, Francine Chen 0001, Yan-Ying Chen, Wei Wang 0010
NAACL-HLT3
2017 Video to Text Summary: Joint Video Summarization and Captioning with Recurrent Neural Networks
Bor-Chun Chen, Yan-Ying Chen, Francine Chen 0001
BMVC2
2017 Multi-task learning for face identification and attribute estimation
abstract
Convolution neural network (CNN) has been shown as one of state-of-the-art approaches for learning face representations. However, previous works only utilized identity information instead of leveraging human attributes (e.g., gender and age) which contain high-level semantic meaning. In this work, we aim to incorporate identity and human attributes in learning discriminative face representations through multi-task learning. In our experiments, we learn face representation by using the largest publicly face dataset CASIA-WebFace with gender and age labels, and then evaluate learned features on widely-used LFW benchmark for face verification and identification. We also compare the effectiveness of different attributes for improving face identification. The results show that the proposed model outperforms the baseline CNN method without using multi-task learning and hand-crafted features such as high-dimensional LBP. We also do experiments on gender and age estimation on Adience benchmark to demonstrate that human attribute prediction can also benefit from the proposed multi-task representation learning.
Hui-Lan Hsieh, Winston H. Hsu, Yan-Ying Chen
ICASSP3
2017 Image-based user profiling of frequent and regular venue categories
abstract
The availability of mobile access has shifted social media use. With that phenomenon, what users shared on social media and where they visited is naturally an excellent resource to learn their visiting behavior. Knowing visit behaviors would help market survey and customer relationship management, e.g., sending customers coupons of the businesses that they visit frequently. Most prior studies leverage meta-data e.g., check-in locations to profile visiting behavior but neglect important information from user-contributed content, e.g., images. This work addresses a novel use of image content for predicting the user visit behavior, i.e., the frequent and regular business venue categories that the content owner would visit. To collect training data, we propose a strategy to use geo-metadata associated with images for deriving the labels of an image owner's visit behavior. Moreover, we model a user's sequential images by using an end-to-end learning framework to reduce the optimization loss. That helps improve the prediction accuracy against the baseline as demonstrated in our experiments. The prediction is completely based on image content that is more available in social media than geo-metadata, and thus allows coverage in profiling a wider set of users.
Ryosuke Shigenaka, Yan-Ying Chen, Francine Chen 0001, Dhiraj Joshi, Yukihiro Tsuboshita
ICME2
2017 Scalable Face Track Retrieval in Video Archives Using Bag-of-Faces Sparse Representation
abstract
Huge video archives consisting of news programs, dramas, movies, and Web videos (e.g., YouTube) are available in our daily life. In all these videos, human is usually one of the most important subjects. Using state-of-the-art techniques, we can efficiently detect and track faces in the videos. In order to organize large-scale face tracks, containing sequences of (detected) consecutive faces in the videos, we propose an efficient method to retrieve human face tracks using bag-of-faces sparse representation (BoF-SR). Using the proposed method, a face track is encoded as a single BoF-SR, therefore allowing an efficient indexing method to handle large-scale data. To further consider the possible variations in face tracks, we generalize our method to find multiple SRs, in an unsupervised manner, to represent a bag of faces and balance the tradeoff between performance and retrieval time. The experimental results on two real-world (million-scale) data sets confirm that the proposed methods achieve significant performance gains compared with different state-of-the-art methods.
Bor-Chun Chen, Yan-Ying Chen, Yin-Hsi Kuo, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Winston H. Hsu
IEEE Trans. Circuits Syst. Video Technol.2
2016 Business-Aware Visual Concept Discovery from Social Media for Multimodal Business Venue Recognition
abstract
Image localization is important for marketing and recommendation of local business; however, the level of granularity is still a critical issue. Given a consumer photo and its rough GPS information, we are interested in extracting the fine-grained location information, i.e. business venues, of the image. To this end, we propose a novel framework for business venue recognition. The framework mainly contains three parts. First, business-aware visual concept discovery: we mine a set of concepts that are useful for business venue recognition based on three guidelines including business awareness, visually detectable, and discriminative power. We define concepts that satisfy all of these three criteria as business-aware visual concept. Second, business-aware concept detection by convolutional neural networks (BA-CNN): we propose a new network configuration that can incorporate semantic signals mined from business reviews for extracting semantic concept features from a query image. Third, multimodal business venue recognition: we extend visually detected concepts to multimodal feature representations that allow a test image to be associated with business reviews and images from social media for business venue recognition. The experiments results show the visual concepts detected by BA-CNN can achieve up to 22.5% relative improvement for business venue recognition compared to the state-of-the-art convolutional neural network features. Experiments also show that by leveraging multimodal information from social media we can further boost the performance, especially when the database images belonging to each business venue are scarce.
Bor-Chun Chen, Yan-Ying Chen, Francine Chen 0001, Dhiraj Joshi
AAAI2
2016 Corpus for Customer Purchase Behavior Prediction in Social Media
Shigeyuki Sakaki, Francine Chen 0001, Mandy Korpusik, Yan-Ying Chen
LREC4
2015 Visually Interpreting Names as Demographic Attributes by Exploiting Click-Through Data
abstract
Name of an identity is strongly influenced by his/her cultural background such as gender and ethnicity, both vital attributes for user profiling, attribute-based retrieval, etc. Typically, the associations between names and attributes (e.g., people named "Amy" are mostly females) are annotated manually or provided by the census data of governments. We propose to associate a name and its likely demographic attributes by exploiting click-throughs between name queries and images with automatically detected facial attributes. This is the first work attempting to translate an abstract name to demographic attributes in visual-data-driven manner, and it is adaptive to incremental data, more countries and even unseen names (the names out of click-through data) without additional manual labels. In the experiments, the automatic name-attribute associations can help gender inference with competitive accuracy by using manual labeling. It also benefits profiling social media users and keyword-based face image retrieval, especially for contributing 12% relative improvement of accuracy in adapting to unseen names.
Yan-Ying Chen, Yin-Hsi Kuo, Chun-Che Wu, Winston H. Hsu
AAAI1
2015 Identify Visual Human Signature in community via wearable camera
abstract
With the increasing popularity of wearable devices, information becomes much easily available. However, personal information sharing still poses great challenges because of privacy issues. We propose an idea of Visual Human Signature (VHS) which can represent each person uniquely even captured in different views/poses by wearable cameras. We evaluate the performance of multiple effective modalities for recognizing an identity, including facial appearance, visual patches, facial attributes and clothing attributes. We propose to emphasize significant dimensions and do weighted voting fusion for incorporating the modalities to improve the VHS recognition. By jointly considering multiple modalities, the VHS recognition rate can reach by 51% in frontal images and 48% in the more challenging environment and our approach can surpass the baseline with average fusion by 25% and 16%. We also introduce Multiview Celebrity Identity Dataset (MCID), a new dataset containing hundreds of identities with different view and clothing for comprehensive evaluation.
Chia-Chin Tsao, Yan-Ying Chen, Yu-Lin Hou, Winston H. Hsu
ICASSP2
2015 Assistive Image Comment Robot - A Novel Mid-Level Concept-Based Representation
abstract
We present a general framework and working system for predicting likely affective responses of the viewers in the social media environment after an image is posted online. Our approach emphasizes a mid-level concept representation, in which intended affects of the image publisher is characterized by a large pool of visual concepts (termed PACs) detected from image content directly instead of textual metadata, evoked viewer affects are represented by concepts (termed VACs) mined from online comments, and statistical methods are used to model the correlations among these two types of concepts. We demonstrate the utilities of such approaches by developing an end-to-end Assistive Comment Robot application, which further includes components for multi-sentence comment generation, interactive interfaces, and relevance feedback functions. Through user studies, we showed machine suggested comments were accepted by users for online posting in 90 percent of completed user sessions, while very favorable results were also observed in various dimensions (plausibility, preference, and realism) when assessing the quality of the generated image comments.
Yan-Ying Chen, Tao Chen 0015, Taikun Liu, Hong-Yuan Mark Liao, Shih-Fu Chang
IEEE Trans. Affect. Comput.1
2014 Predicting Viewer Affective Comments Based on Image Content in Social Media
abstract
Visual sentiment analysis is getting increasing attention because of the rapidly growing amount of images in online social interactions and several emerging applications such as online propaganda and advertisement. Recent studies have shown promising progress in analyzing visual affect concepts intended by the media content publisher. In contrast, this paper focuses on predicting what viewer affect concepts will be triggered when the image is perceived by the viewers. For example, given an image tagged with "yummy food," the viewers are likely to comment "delicious" and "hungry," which we refer to as viewer affect concepts (VAC) in this paper. To the best of our knowledge, this is the first work explicitly distinguishing intended publisher affect concepts and induced viewer affect concepts associated with social visual content, and aiming at understanding their correlations. We present around 400 VACs automatically mined from million-scale real user comments associated with images in social media. Furthermore, we propose an automatic visual based approach to predict VACs by first detecting publisher affect concepts in image content and then applying statistical correlations between such publisher affect concepts and the VACs. We demonstrate major benefits of the proposed methods in several real-world tasks - recommending images to invoke certain target VACs among viewers, increasing the accuracy of predicting VACs by 20.1% and finally developing a social assistant tool that may suggest plausible, content-specific and desirable comments when users view new images.
Yan-Ying Chen, Tao Chen 0015, Winston H. Hsu, Hong-Yuan Mark Liao, Shih-Fu Chang
ICMR1
2014 Object-Based Visual Sentiment Concept Analysis and Application
abstract
This paper studies the problem of modeling object-based visual concepts such as "crazy car" and "shy dog" with a goal to extract emotion related information from social multimedia content. We focus on detecting such adjective-noun pairs because of their strong co-occurrence relation with image tags about emotions. This problem is very challenging due to the highly subjective nature of the adjectives like "crazy" and "shy" and the ambiguity associated with the annotations. However, associating adjectives with concrete physical nouns makes the combined visual concepts more detectable and tractable. We propose a hierarchical system to handle the concept classification in an object specific manner and decompose the hard problem into object localization and sentiment related concept modeling. In order to resolve the ambiguity of concepts we propose a novel classification approach by modeling the concept similarity, leveraging on online commonsense knowledgebase. The proposed framework also allows us to interpret the classifiers by discovering discriminative features. The comparisons between our method and several baselines show great improvement in classification performance. We further demonstrate the power of the proposed system with a few novel applications such as sentiment-aware music slide shows of personal albums.
Tao Chen 0015, Felix X. Yu, Yin Cui, Yan-Ying Chen, Shih-Fu Chang
ACM Multimedia5
2014 Automatic Facial Image Annotation and Retrieval by Integrating Voice Label and Visual Appearance
abstract
Annotation is important for managing and retrieving a large amount of photos, but it is generally labor-intensive and time-consuming. However, speaking while taking photos is straightforward and effortless, and using voice for annotation is faster than typing words. To best reduce the manual cost of annotating photos, we propose a novel framework which utilizes the scarce spoken annotations recorded while capturing as voice labels and automatically label every facial image in the photo collection. To accomplish this goal, we employ a probabilistic graphical model which integrates voice labels and visual appearances for inference. Combined with group prior estimation and gender attribute association, we can achieve an outstanding performance on the proposed synthesized group photo collections.
Hong-Wun Jheng, Bor-Chun Chen, Yan-Ying Chen, Winston H. Hsu
ACM Multimedia3
2014 Discovering the City by Mining Diverse and Multimodal Data Streams
abstract
This work attempts to tackle the IBM grand challenge - seeing the daily life of New York City (NYC) in various perspectives by exploring rich and diverse social media content. Most existing works address this problem relying on single media source and covering limited life aspects. Because different social media are usually chosen for specific purposes, multiple social media mining and integration are essential to understand a city comprehensively. In this work, we first discover the similar and unique natures (e.g., attractions, topics) across social media in terms of visual and semantic perceptions. For example, Instagram users share more food and travel photos while Twitter users discuss more about sports and news. Based on these characteristics, we analyze a broad spectrum of life aspects - trends, events, food, wearing and transportation in NYC by mining a huge amount of diverse and freely available media (e.g., 1.6M Instagram photos, 5.3M Twitter posts). Because transportation logs are hardly available in social media, the NYC Open Data (e.g., 6.5B subway station transactions) is leveraged to visualize temporal traffic patterns. Furthermore, the experiments demonstrate that our approaches can effectively overview urban life with considerable technical improvement, e.g., having 16% relative gains in food recognition accuracy by a hierarchy cross-media learning strategy, reducing the feature dimensions of sentiment analysis by 10 times without sacrificing precision.
Yin-Hsi Kuo, Yan-Ying Chen, Bor-Chun Chen, Wen-Yu Lee, Chun-Che Wu, Yu-Lin Hou, Wen-Feng Cheng, Yi-Chih Tsai, Chung-Yen Hung, Liang-Chi Hsieh, Winston H. Hsu
ACM Multimedia2
2014 Facial Attribute Space Compression by Latent Human Topic Discovery
abstract
Facial attribute is important information for a variety of machine vision tasks including recognition, classification, and retrieval. There arises a strong need for detecting various facial attributes such as gender, age and more which consume more computation and storage resources. Therefore, we propose a compression framework to find fewer significant Latent Human Topics (LHT) to approximate more facial attributes. LHT is a combination of attribute correlation by transferring facial attribute space to compressional space with Singular Value Decomposition (SVD). Using the proposed scheme, we can easily detect the facial attributes from a face image via fast reconstructing the compressed labels automatically detected by a few LHT classifiers. Experimental results show that our system can achieve similar performance with substantially fewer dimensions compared to the original number of facial attributes, and it even shows slight improvements because LHT carry informative attribute correlations learned from data.
Yan-Ying Chen, Bor-Chun Chen, Yu-Lin Hou, Winston H. Hsu
ACM Multimedia2
2013 Enabling low bitrate mobile visual recognition: a performance versus bandwidth evaluation
abstract
The rapid development of technologies in both hardware and software have made content-based multimedia services feasible on mobile devices such as smartphones and tablets; and the strong needs for mobile visual search and recognition have been emerging. While many real applications of visual recognition require a large scale recognition systems, the same technologies that support server-based scalable visual recognition may not be feasible on mobile devices due to the resource constraints. Although the client-server framework ensures the scalability, the real-time response subjects to the limitation on network bandwidth. Therefore, the main challenge for mobile visual recognition system should be the recognition bitrate, which is the amount of data transmission under the same recognition performance. For this work, we exploit and compare various strategies such as compact features, feature compression, feature signatures by hashing, image scaling, etc., to enable low bitrate mobile visual recognition. We argue that thumbnail image is a competitive candidate for low bitrate visual recognition because it carries multiple features at once and multi-feature fusion is important as the size of semantic space increases. Our evaluations on two subsets of ImageNet, both contain more than 10,000 images with 19 and 137 categories, verify the efficacy of thumbnail images. We further suggest a new strategy that combines single (local) feature signature and the thumbnail image, which achieves significant bitrate reduction from (average) 102,570 to 4,661 bytes with merely (overall) 10% performance degradation.
Yu-Chuan Su, Tzu-Hsuan Chiu, Yan-Ying Chen, Chun-Yen Yeh, Winston H. Hsu
ACM Multimedia3
2013 Search-based relevance association with auxiliary contextual cues
abstract
In this work, we target at solving the Bing challenge provided by Microsoft. The task is to design an effective and efficient measurement of query terms in describing the images (image-query pairs) crawled from the web. We observe that the provided image-query pairs (e.g., text-based image retrieval results) are usually related to their surrounding text; however, the relationship between image content seems to be ignored. Hence, we attempt to integrate the visual information for better ranking results. In addition, we found that plenty of query terms are related to people (e.g., celebrity) and user might have similar queries (click logs) in the search engine. Therefore, in this work, we propose a relevance association by investigating the effectiveness of different auxiliary contextual cues (i.e., face, click logs, visual similarity). Experimental results show that the proposed method can have 16% relative improvement compared to the original ranking results. Especially, for people-related queries, we can further have 45.7% relative improvement.
Chun-Che Wu, Kuan-Yu Chu, Yin-Hsi Kuo, Yan-Ying Chen, Wen-Yu Lee, Winston H. Hsu
ACM Multimedia4
2013 Travel Recommendation by Mining People Attributes and Travel Group Types From Community-Contributed Photos
abstract
Leveraging community-contributed data (e.g., blogs, GPS logs, and geo-tagged photos) for personalized recommendation is one of the active research problems since there are rich contexts and human activities in such explosively growing data. In this work, we focus on personalized travel recommendation and show promising applications by leveraging the freely available community-contributed photos. We propose to conduct personalized travel recommendation by further considering specific user profiles or attributes (e.g., gender, age, race) as well as travel group types (e.g., family, friends, couple). Instead of mining photo logs only, we exploit the automatically detected people attributes and travel group types in the photo contents. By information-theoretic measures, we demonstrate that such detected user profiles are informative and effective for travel recommendation-especially providing a promising aspect for personalization. We effectively mine the demographics of individual and group travelers for different locations (or landmarks) and their travel paths. A probabilistic Bayesian learning framework which further entails mobile recommendation on the spot is introduced as well. We experiment on more than 10 million photos collected from 19 major cities worldwide and conduct the extensive investigation of profiling activities in communities according to temporal and spatial information. Note that the photos in the paper attribute to various Flickr users under the Creative Commons License. The experiments confirm that people attributes of individuals and groups are promising and orthogonal to prior works using travel logs only and can further improve prior travel recommendation methods especially for difficult predictions by further leveraging user contexts via mobile devices.
Yan-Ying Chen, An-Jung Cheng, Winston H. Hsu
IEEE Trans. Multim.1
2013 Scalable Face Image Retrieval Using Attribute-Enhanced Sparse Codewords
abstract
Photos with people (e.g., family, friends, celebrities, etc.) are the major interest of users. Thus, with the exponentially growing photos, large-scale content-based face image retrieval is an enabling technology for many emerging applications. In this work, we aim to utilize automatically detected human attributes that contain semantic cues of the face photos to improve content-based face retrieval by constructing semantic codewords for efficient large-scale face retrieval. By leveraging human attributes in a scalable and systematic framework, we propose two orthogonal methods named attribute-enhanced sparse coding and attribute-embedded inverted indexing to improve the face retrieval in the offline and online stages. We investigate the effectiveness of different attributes and vital factors essential for face retrieval. Experimenting on two public datasets, the results show that the proposed methods can achieve up to 43.5% relative improvement in MAP compared to the existing methods.
Bor-Chun Chen, Yan-Ying Chen, Yin-Hsi Kuo, Winston H. Hsu
IEEE Trans. Multim.2
2013 Automatic Training Image Acquisition and Effective Feature Selection From Community-Contributed Photos for Facial Attribute Detection
abstract
Facial attributes are shown effective for mining specific persons and profiling human activities in large-scale media such as surveillance videos or photo-sharing services. For comprehensive analysis, a rich number of facial attributes is required. Generally, each attribute detector is obtained by supervised learning via the use of large training data. It is promising to leverage the exponentially growing community contributed photos and the associated informative contexts to ease the burden of manual annotation; however, such huge noisy data from the Internet still pose great challenges. We propose to measure the quality of training images by discriminable visual features, which are verified with the relative discrimination between the unlabeled images and the pseudo-positives (pseudo-negatives) retrieved by textual relevance. The proposed feature selection requires no heuristic threshold, therefore, can be generalized to multiple feature modalities. We further exploit the rich context cues (e.g., tags, geo-locations, etc.) associated with the publicly available photos for mining more semantically consistent but visually diverse training images around the world. Experimenting in the benchmarks, we demonstrate that our work can successfully acquire effective training images for learning generic facial attributes, where the classification error is relatively reduced up to 23.35% compared with that of the text-based approach and shown comparable with that of costly manual annotations. (All of the face images presented in this paper except the training images in Fig. 8 attribute to various Flickr users under Creative Commons Licenses).
Yan-Ying Chen, Winston H. Hsu, Hong-Yuan Mark Liao
IEEE Trans. Multim.1
2012 People search and activity mining in large-scale community-contributed photos
abstract
A growing number of users are contributing a huge amount of photos containing people (e.g., family, classmates, colleagues, etc.) to social media for the purpose of photo sharing and social communication. There arises a strong need for automatically analyzing the people shown in large-scale photos because these visual data comprise abundant consumer activities which greatly benefit demographic analysis and enhance marketing research. In this work, we aim at learning facial attributes (gender, race, age, etc.) by these publicly available photos and exploiting the detected facial attributes for locating designated persons, profiling user preferences and predicting social group types. In addition, community-contributed data possess rich contexts such as tags, geo-locations and time stamps, which strongly correlate with user intentions and preferences. The knowledge would be informative to actively refine the recognition models and promising towards improvement of photo management, personalized recommendation and social networking. Most importantly, this framework effectively relieves costly annotation efforts and ensures scalability for large-scale media.
Yan-Ying Chen
ACM Multimedia1
2012 Discovering informative social subgraphs and predicting pairwise relationships from group photos
abstract
An increasing number of users are contributing the sheer amount of group photos (e.g., for family, classmates, colleagues, etc.) on social media for the purpose of photo sharing and social communication. There arise strong needs for automatically understanding the group types (e.g., family vs. classmates) for recommendation services (e.g., recommending a family-friendly restaurant) and even predicting the pairwise relationships (e.g., mother-child) between the people in the photo for mining implicit social connections. Interestingly, we observe that the group photos are composed of atomic subgroups corresponding to certain social relationships. For this work, we propose a novel framework to (1) connect faces of different attributes and positions as a face graph and (2) discover informative subgraphs to represent social subgroups in group photos. A group photo can be further represented by a bag-of-face-subgraphs (BoFG) -- the occurring frequency of social subgroups, which is informative to categorize specific group types or events. We demonstrate the effectiveness of BoFG in recognizing family photos and achieve 30.5% relative improvement over the state-of-the-art low-level features. Moreover, we propose to predict the pairwise relationships (e.g., husband-wife) in a face graph by the co-occurrence information (e.g., co-occurring with a child) in the mined subgraphs. The experiments demonstrate that the informative social subgroups significantly outperform prior work (36% relatively) which considers merely facial attributes for determining pairwise relationships.
Yan-Ying Chen, Winston H. Hsu, Hong-Yuan Mark Liao
ACM Multimedia1
2012 Where is who: large-scale photo retrieval by facial attributes and canvas layout
abstract
The ubiquitous availability of digital cameras has made it easier than ever to capture moments of life, especially the ones accompanied with friends and family. It is generally believed that most family photos are with faces that are sparsely tagged. Therefore, a better solution to manage and search in the tremendously growing personal or group photos is highly anticipated. In this paper, we propose a novel way to search for face photos by simultaneously considering attributes (e.g., gender, age, and race), positions, and sizes of the target faces. To better match the content and layout of the multiple faces in mind, our system allows the user to graphically specify the face positions and sizes on a query "canvas," where each attribute combination is defined as an icon for easier representation. As a secondary feature, the user can even place specific faces from the previous search results for appearance-based retrieval. The scenario has been realized on a tablet device with an intuitive touch interface. Experimenting with a large-scale Flickr dataset of more than 200k faces, the proposed formulation and joint ranking have made us achieve a hit rate of 0.420 at rank 100, significantly improving from 0.036 of the prior search scheme using attributes alone. We have also achieved an average running time of 0.0558 second by the proposed block-based indexing approach.
Yu-Heng Lei, Yan-Ying Chen, Bor-Chun Chen, Lime Iida, Winston H. Hsu
SIGIR2
2012 Learning by expansion: Exploiting social media for image classification with few training examples
Sheng-Yuan Wang, Wei-Shing Liao, Liang-Chi Hsieh, Yan-Ying Chen, Winston H. Hsu
Neurocomputing4
2011 Semi-supervised face image retrieval using sparse coding with identity constraint
abstract
We aim to develop a scalable face image retrieval system which can integrate with partial identity information to improve the retrieval result. To achieve this goal, we first apply sparse coding on local features extracted from face images combining with inverted indexing to construct an efficient and scalable face retrieval system. We then propose a novel coding scheme that refines the representation of the original sparse coding by using identity information. Using the proposed coding scheme, face images with large intra-class variances will still be quantized into similar visual words if they share the same identity. Experimental results show that our system can achieve salient retrieval results on LFW dataset (13K faces) and outperform linear search methods using well known face recognition feature descriptors.
Bor-Chun Chen, Yin-Hsi Kuo, Yan-Ying Chen, Kuan-Yu Chu, Winston H. Hsu
ACM Multimedia3
2011 Personalized travel recommendation by mining people attributes from community-contributed photos
abstract
Leveraging community-contributed data (e.g., blogs, GPS logs, and geo-tagged photos) for travel recommendation is one of the active researches since there are rich contexts and trip activities in such explosively growing data. In this work, we focus on personalized travel recommendation by leveraging the freely available community-contributed photos. We propose to conduct personalized travel recommendation by further considering specific user profiles or attributes (e.g., gender, age, race). In stead of mining photo logs only, we argue to leverage the automatically detected people attributes in the photo contents. By information-theoretic measures, we will demonstrate that such people attributes are informative and effective for travel recommendation -- especially providing a promising aspect for personalization. We effectively mine the demographics for different locations (or landmarks) and travel paths. A probabilistic Bayesian learning framework which further entails mobile recommendation on the spot is introduced. We experiment on four million photos collected for eight major worldwide cities. The experiments confirm that people attributes are promising and orthogonal to prior works using travel logs only and can further improve prior travel recommendation methods especially in difficult predictions by further leveraging user contexts in mobile devices.
An-Jung Cheng, Yan-Ying Chen, Yen-Ta Huang, Winston H. Hsu, Hong-Yuan Mark Liao
ACM Multimedia2
2011 Photo search by face positions and facial attributes on touch devices
abstract
With the explosive growth of camera devices, people can freely take photos to capture moments of life, especially the ones accompanied with friends and family. Therefore, a better solution to organize the increasing number of personal or group photos is highly required. In this paper, we propose a novel way to search for face images according facial attributes and face similarity of the target persons. To better match the face layout in mind, our system allows the user to graphically specify the face positions and sizes on a query "canvas," where each attribute or identity is defined as an "icon" for easier representation. Moreover, we provide aesthetics filtering to enhance visual experience by removing candidates of poor photographic qualities. The scenario has been realized on a touch device with an intuitive user interface. With the proposed block-based indexing approach, we can achieve near real-time retrieval (0.1 second on average) in a large-scale dataset (more than 200k faces in Flickr images).
Yu-Heng Lei, Yan-Ying Chen, Lime Iida, Bor-Chun Chen, Hsiao-Hang Su, Winston H. Hsu
ACM Multimedia2