Mei-Chen Yeh

dblp:89/3060 · DBLP profile ↗
← Back
38ranked-venue papers
14as first author
8since 2021 · last 2025
0000-0001-8665-7860ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 Self-supervised learning of pseudo classes for generalized zero-shot fine-grained recognition
Yan-He Chen, Mei-Chen Yeh
Multim. Tools Appl.2
2024 Self-Supervised Multi-Label Classification with Global Context and Local Attention
abstract
Self-supervised learning has proven highly effective across various tasks, showcasing its versatility in different applications. Despite these achievements, the challenges inherent in multi-label classification have seen limited attention. This paper introduces GAELLE, a novel self-supervised multi-label classification framework that simultaneously captures image context and object information. GAELLE employs a combination of global context and local attention mechanisms to discern diverse levels of semantic information in images. The global component comprehensively learns image content while local attention eliminates object-irrelevant nuances by aligning embeddings with a projection head. The integration of global and local features in GAELLE effectively captures intricate object-scene relationships. To further enhance this capability, we introduce a global and local swap prediction technique, facilitating the nuanced interplay between various objects and scenes within images. Experimental results showcase GAELLE's state-of-the-art performance in self-supervised multi-label classification tasks, highlighting its effectiveness in uncovering complex relationships between multiple objects and scenes in images.
Mei-Chen Yeh
ICMR2
2024 Indirect visual-semantic alignment for generalized zero-shot recognition
Yan-He Chen, Mei-Chen Yeh
Multim. Syst.2
2023 Weakly- and Semi-Supervised Object Localization
abstract
Weakly supervised object localization deals with the lack of location-level labels to train localization models. Recently a new evaluation protocol is proposed in which full supervision is available but limited to only a small validation set. It motives us to explore semi-supervised learning for addressing this problem. In particular, the localization model is developed via self-training: we use a small amount of data with full supervision to train a class-agnostic detector, and use it to generate pseudo bounding boxes for data with weak supervision. Furthermore, we propose a selection algorithm to discover high-quality pseudo labels, and deal with data imbalance caused by pseudo labeling. We demonstrate the superiority of the proposed method with performance on par with the state of the art on two benchmarks.
Zhen-Tang Huang, Yan-He Chen, Mei-Chen Yeh
ICASSP3
2022 Generative and Adaptive Multi-Label Generalized Zero-Shot Learning
abstract
We address the problem of multi-label generalized zero-shot learning where the task is to predict the labels (usually more than one) of a target image whether each of its labels belongs to the seen or unseen category. To alleviate the extreme data-imbalance problem, in which no annotated images are available for unseen classes during training, state-of-the-art single-label zero-shot learning methods learn to synthesize the class-specific visual features from seen classes. However, synthesizing multi-label visual features from multi-label images has not been extensively studied. By exploring the relationship between an image and its labels, we address the multi-label generalized zero-shot learning problem via a hybrid framework of generative and adaptive learning. We convert an image into a label classifier, which can vary among intra-class samples. The adaptive mechanism facilitates the usage of a single-label feature generating model for creating multi-label features from multi-label images. We show that the proposed method improves the state of the art ZSL/GZSL methods on two benchmark datasets.
Kuan-Ying Chen, Mei-Chen Yeh
ICME2
2022 A Semi-Supervised Learning Approach for Traditional Chinese Scene Text Detection
abstract
With the advancement of multimedia technology, the information in surrounding environment has becoming accessible. In particular, automatic scene text detection is essential for subsequent text recognition, understanding and analysis. However, most existing methods are primarily designed for English, while those for other languages are scarce. In this paper we present a traditional Chinese scene text detector, built upon a robust object detector trained with labeled and unlabeled data via semi-supervised learning. Moreover, we expand the limited labeled data by data synthesis and a data augmentation method. We demonstrate the effectiveness of the proposed method through extensive experiments, and examine the design choices in developing a practical system that can instantly and accurately detect traditional Chinese texts in complex scenes.
Chia-Fu Yeh, Mei-Chen Yeh
MMSP2
2021 Generalized Zero-Shot Recognition Through Image-Guided Semantic Classification
abstract
We present a new visual-semantic embedding method for generalized zero-shot learning. Different to existing embedding-based methods that learn the correspondence between an image classifier and its class prototype for each class, we learn the mapping between an image and its semantic classifier. Given an input image, the proposed method creates a label classifier and applies it to all label embeddings to determine whether a label belongs to the input image. Therefore, a semantic classifier is image conditioned and is generated during inference. We validate our approach with four standard benchmark datasets.
Mei-Chen Yeh
ICIP2
2021 Text-Enhanced Attribute-Based Attention for Generalized Zero-Shot Fine-Grained Image Classification
abstract
We address the generalized zero-shot fine-grained image classification problem, in which classes are visually similar and training images for some classes are not available. We leverage auxiliary information in the form of textual descriptions to facilitate the task. Specifically, we propose a text-enhanced attribute-based attention mechanism to compute features from the most relevant image regions guided from the most relevant attributes. Experiments on two popular datasets of CUB and AWA2 show the effectiveness of the proposed method.
Yan-He Chen, Mei-Chen Yeh
ICMR2
2020 Unsupervised Multi-Task Domain Adaptation
abstract
With abundant labeled data, deep convolutional neural networks have shown great success in various image recognition tasks. However, these models are often less powerful when applied to novel datasets due to a phenomenon known as domain shift. Unsupervised domain adaptation methods aim to address this problem, allowing deep models trained on the labeled source domain to be used on a different target domain (without labels). In this paper, we investigate whether the generalization ability of an unsupervised domain adaptation method can be improved through multi-task learning, with learned features required to be both domain invariant and discriminative for multiple different but relevant tasks. Experiments evaluating two fundamental recognition tasks-image recognition and segmentation-show that the generalization ability empowered by multi-task learning may not benefit recognition when the model is directly applied on the target domain, but the multi-task learning setting can boost the performance of state-of-the-art unsupervised domain adaptation methods by a non-negligible margin.
Shih-Min Yang, Mei-Chen Yeh
ICPR2
2020 Artist-based painting classification using Markov random fields with convolution neural network
Kai-Lung Hua, Trang-Thi Ho, Kevin Alfianto Jangtjik, Mei-Chen Yeh
Multim. Tools Appl.5
2020 Multilabel Deep Visual-Semantic Embedding
abstract
Inspired by the great success from deep convolutional neural networks (CNNs) for single-label visual-semantic embedding, we exploit extending these models for multilabel images. We propose a new learning paradigm for multilabel image classification, in which labels are ranked according to its relevance to the input image. In contrast to conventional CNN models that learn a latent vector representation (i.e., the image embedding vector), the developed visual model learns a mapping (i.e., a transformation matrix) from an image in an attempt to differentiate between its relevant and irrelevant labels. Despite the conceptual simplicity of our approach, the proposed model achieves state-of-the-art results on three public benchmark datasets.
Mei-Chen Yeh, Yi-Nan Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Early Detection of Vacant Parking Spaces Using Dashcam Videos
abstract
A major problem in metropolitan areas is finding parking spaces. Existing parking guidance systems often adopt fixed sensors or cameras that cannot provide information from the driver’s point of view. Motivated by the advent of dashboard cameras (dashcams), we develop neural-network-based methods for detecting vacant parking spaces in videos recorded by a dashcam. Detecting vacant parking spaces in dashcam videos enables early detection of spaces. Different from conventional object detection methods, we leverage the monotonicity of the detection confidence with respect to the distance away of the approaching target parking space and propose a new loss function, which can not only yield improved detection results but also enable early detection. To evaluate our detection method, we create a new large dataset containing 5,800 dashcam videos captured from 22 indoor and outdoor parking lots. To the best of our knowledge, this is the first and largest driver’s view video dataset that supports parking space detection and provides parking space occupancy annotations.
Ming-Che Wu, Mei-Chen Yeh
AAAI2
2019 A Topological Data Analysis Approach to Video Summarization
abstract
This paper explores the use of simplicial complex to construct a new structure-wise representation for videos. Complementary to the appearance-based representation, which usually involves feature extraction from video content, the proposed method captures the structural properties of a video and can be used for various video processing tasks. We demonstrate a case study of the proposed approach to automated video summarization, which relies only on the developed topological structures and requires no training phase. Experimental results show the potential of this approach to facilitate tasks involving the understanding and analysis of structural information inherent in videos.
Chuan-Shen Hu, Mei-Chen Yeh
ICIP2
2019 Virtual Portraitist: An Intelligent Tool for Taking Well-Posed Selfies
abstract
Smart photography carries the promise of quality improvement and functionality extension in making aesthetically appealing pictures. In this article, we focus on self-portrait photographs and introduce new methods that guide a user in how to best pose while taking a selfie. While most of the current solutions use a post processing procedure to beautify a picture, the developed tool enables a novel function of recommending a good look before the photo is captured. Given an input face image, the tool automatically estimates the pose-based aesthetic score, finds the most attractive angle of the face, and suggests how the pose should be adjusted. The recommendation results are determined adaptively to the appearance and initial pose of the input face. We apply a data mining approach to find distinctive, frequent itemsets and association rules from online profile pictures, upon which the aesthetic estimation and pose recommendation methods are developed. A simulated and a real image set are used for experimental evaluation. The results show the proposed aesthetic estimation method can effectively select user-favorable photos. Moreover, the recommendation performance for the vertical adjustment is moderately related to the degree of conformity among the professional photographers’ recommendations. This study echoes the trend of instant photo sharing, in which a user takes a picture and then immediately shares it on a social network without engaging in tedious editing.
Chuan-Shen Hu, Yi-Tsung Hsieh, Hsiao-Wei Lin, Mei-Chen Yeh
ACM Trans. Multim. Comput. Commun. Appl.4
2017 Granularity-based interactive image display
abstract
This paper presents a prototype system that assists users in accessing an unstructured image set (e.g., search results of a query). The system provides a spectrum of overviews, each of which is determined by the display granularity (i.e., the level of summary) an user desires. This new functionality enables a new granularity-based interactive image browsing experience.
Geng-Zhi Fann, Mei-Chen Yeh
ICIP2
2017 A CNN-LSTM framework for authorship classification of paintings
abstract
The authenticity of digital painting image is an urgent demand in the field of art. Yet, determining the authorship of a certain painting is a challenging task due to two reasons: (1) various artists might share similar painting styles; and (2) an artist could create different styles. In this paper, we present a novel method for authorship classification of paintings based on a CNN-LSTM framework. First, a multiscale pyramid is constructed from a painting image. Second, a CNN-LSTM model is learned and it returns possibly multiple labels for one image. To aggregate the final classification result, an adaptive fusion method is employed. Experimental results show that the proposed method has superior classification performance compared with the state-of-the-art techniques.
Kevin Alfianto Jangtjik, Trang-Thi Ho, Mei-Chen Yeh, Kai-Lung Hua
ICIP3
2017 The Impact of Feng Shui on House Price: A Data Perspective
abstract
In this paper, we address interesting questions about how feng shui influences house price from a data perspective. First, is feng shui likely to influence house price? Second, how do different feng shui features, e.g., house shape, master bedroom location, and other interior room arrangements, influence the price? Third, can we automatically diagnose the feng shui problems of a house? From a dataset of 4,579 items of information on houses in Taipei City, we automatically infer what factors determine the house price and identify the feng shui problems of a house that may lower a home's economic value.
Tsuo-Chen Wu, Mei-Chen Yeh
ISM2
2017 Human Pose Tracking Using Online Latent Structured Support Vector Machine
Kai-Lung Hua, Irawati Nurmala Sari, Mei-Chen Yeh
MMM (1)3
2016 A Computational Approach to Finding Facial Patterns of a Babyface
abstract
Facial babyishness has a strong impact on social perceptions and interactions; however, the components constituting a babyface remain unclear. In this paper, we present a computational approach for identifying important but less apparent facial patterns of a babyface, using voluminous face images on the web. The proposed approach is built upon computationally efficient data mining techniques. A new image set with ground truth data collected from users and an evaluation approach based on age estimation are presented in the experiment. The results show that the mined patterns are effective for understanding and determining babyfaces. The findings of this study should provide information for future investigations on the prediction and analysis of trait impressions using the patterns.
Zi-Yi Ke, Mei-Chen Yeh
ICMR2
2016 Artist-based Classification via Deep Learning with Multi-scale Weighted Pooling
abstract
For analyzing digital images of paintings we propose a new approach to categorize them based on artist. Determining the authorship of a painting is challenging because common subjects are illustrated in paintings, and paintings of an artist may not have a unique style. The proposed approach is built upon convolutional neural networks (CNN)---a class of biologically inspired vision model that recently demonstrates near-human performance on several visual recognition tasks. However, training a CNN model requires large scale training data of a fixed input image size (e.g. 224 * 224). In this paper, we propose to construct a multi-layer pyramid from an image, providing 21X more features than using a single layer (i.e., the original image) alone. We train a CNN model for each layer, and propose a new weighted fusion scheme to adaptively combine the decision results. To evaluate the proposed methods, we collect a new painting image dataset, categorized into 13 artists. As demonstrated in the experimental results, the proposed method achieves a promising result---88.08% recall rate in top-2 retrieval on the challenging classification task.
Kevin Alfianto Jangtjik, Mei-Chen Yeh, Kai-Lung Hua
ACM Multimedia2
2016 A content-based approach for detecting highlights in action movies
Mei-Chen Yeh, Yen-Wei Tsai, Hao-Chen Hsu
Multim. Syst.1
2016 Fast medium-scale multiperson identification in aerial videos
Mei-Chen Yeh, Han-Kuen Chiu, Jia-Shung Wang
Multim. Tools Appl.1
2014 Virtual Portraitist: Aesthetic Evaluation of Selfies Based on Angle
abstract
This work addresses the Huawei Grand Challenge that seeks solutions of quality improvement and functionality extension in computational photography. We propose virtual portraitist-a new method that helps users take good selfies in angle. Dissimilar to current solutions that mostly use a post-processing step to fix a photograph, the proposed method enables a novel function of recommending a good look before the photo is captured. This is achieved by using an automatic approach for estimating the aesthetic quality score of a selfie based on angle. In particular, a set of distinctive patterns discovered from a collection of online profile pictures are combined with head pose and camera orientation to rate the quality of a selfie. Experiments validate the effectiveness of the approach.
Mei-Chen Yeh, Hsiao-Wei Lin
ACM Multimedia1
2014 Taking good selfies on your phone
abstract
Selfies are in vogue because smartphones and social media websites have helped popularize the phenomenon whereby people can easily take and share self-portrait photographs. However, taking good selfies is not always a trivial task for end users. In this demonstration we present a tool for guiding users on which angles to take selfies from. Dissimilar to existing techniques that enhance the aesthetic quality of selfies through a post-processing procedure, the proposed approach focuses on the creation of a good look before the photo is captured. A few distinctive patterns automatically discovered from a collection of online profile pictures are used to score a selfie in angle. The technique should enable new features such as "virtual portraitist" for current consumer cameras and smartphones, providing an interactive approach for capturing selfies.
Mei-Chen Yeh, Hsiao-Wei Lin
ACM Multimedia1
2014 Fast salient object detection through efficient subwindow search
Mei-Chen Yeh, Chih-Fan Hsu, Chia-Ju Lu
Pattern Recognit. Lett.1
2013 Real-time salient object detection
abstract
Salient object detection techniques have a variety of multimedia applications of broad interest. However, the detection must be fast to truly aid in these processes. There exist many robust algorithms tackling the salient object detection problem but most of them are computationally demanding. In this demonstration we show a fast salient object detection system implemented in a conventional PC environment. We examine the challenges faced in the design and development of a practical system that can achieve accurate detection in real-time.
Chia-Ju Lu, Chih-Fan Hsu, Mei-Chen Yeh
ACM Multimedia3
2012 Relative features for photo quality assessment
abstract
Automatic evaluation of photo aesthetic quality is a challenging problem in multimedia computing. Numerous aesthetic features have been proposed in previous works but the features are extracted solely from the photo under evaluation. In this paper, we explore the use of multiple images, and present the relative features that can be easily computed from any score-based features. We show that evaluation on a group basis can facilitate the quality assessment problem. Although the extraction of the new feature is extremely simple, computationally efficient, and requires no training phase, experimental results validate the effectiveness of the proposed approach.
Mei-Chen Yeh, Yu-Chen Cheng
ICIP1
2012 Automatic cinemagraphs for ranking beautiful scenes
abstract
This work addresses the NHK challenge that aims at automatic recognition of beautiful scenes in broadcast programs. We propose a method that can be used to automatically extract beautiful scenes from videos and rank the scenes in terms of beauty. In particular, we introduce cinemagraphs as an alternative manner for presenting beautiful scenes, and a notion of beauty based on the presence of interesting motions. The method is fully automatic, requires no training phase, and produces beautiful scene cinemagraphs as by-products.
Yin-Tzu Chan, Hao-Chen Hsu, Po-Yi Li, Mei-Chen Yeh
ACM Multimedia4
2012 An approach to automatic creation of cinemagraphs
abstract
A cinemagraph is a new type of medium that infuses a static image with the dynamics of one particular region. It is in many ways intermediate between a photograph and a video, and has a number of attractive potential applications, such as the creation of dynamic scenes for games and interactive environments. However, creating cinemagraphs is time consuming and requires certain level of proficiency on photo editing techniques. In this paper, we present a fully automatic approach that creates cinemagraphs from video sequences. Specifically, we view cinemagraph construction as a constrained optimization problem that seeks a sub-volume in video with the maximum cumulative flow fields. The problem can be efficiently solved by a branch and-bound search scheme. A user survey is conducted to understand user preferences and demonstrate the performance of the proposed approach. The findings of this study should provide information for various design choices for an easy and versatile authoring tool for cinemagraphs.
Mei-Chen Yeh, Po-Yi Li
ACM Multimedia1
2012 A tool for automatic cinemagraphs
abstract
A cinemagraph is a new type of medium that infuses a static image with the dynamics of one or a few particular regions. It is in many ways intermediate between a photograph and a video, and provides a simple, yet expressive way to mix static and dynamic elements from a video clip. However, the process of creating cinemagraphs is usually tedious for end users and requires serious photo editing skills. In this demonstration we show a tool that creates cinemagraphs in a fully automatic manner. The technique should enable new features such as "intelligent cinemagraph mode" for digital cameras that provides an alternative method to capture the moment.
Mei-Chen Yeh, Po-Yi Li
ACM Multimedia1
2011 A Hierarchical Approach to Practical Beverage Package Recognition
Mei-Chen Yeh, Jason Tai
PSIVT (1)1
2011 Fast Visual Retrieval Using Accelerated Sequence Matching
abstract
We present an approach to represent, match, and index various types of visual data, with the primary goal of enabling effective and computationally efficient searches. In this approach, an image/video is represented by an ordered list of feature descriptors. Similarities between such representations are then measured by the approximate string matching technique. This approach unifies visual appearance and the ordering information in a holistic manner with joint consideration of visual-order consistency between the query and the reference instances, and can be used for automatically identifying local alignments between two pieces of visual data. This capability is essential for tasks such as video copy detection where only small portions of the query and the reference videos are similar. To deal with large volumes of data, we further show that this approach can be significantly accelerated along with a dedicated indexing structure. Extensive experiments on various visual retrieval and classification tasks demonstrate the superior performance of the proposed techniques compared to existing solutions.
Mei-Chen Yeh, Kwang-Ting Cheng
IEEE Trans. Multim.1
2009 A compact, effective descriptor for video copy detection
abstract
Large scale video copy detection tasks require a compact and computational-efficient descriptor that is robust to various transformations that are typically applied to generate copies. In this paper, we propose a new frame-level descriptor for such a task. The descriptor encodes the internal structure of a video frame by computing the pair-wise correlations between geometrically pre-indexed blocks. It is conceptually simple, small in size, and fast to compute. Experiments using the MUSCLE VCD benchmark show its superior performance compared to existing approaches.
Mei-Chen Yeh, Kwang-Ting Cheng
ACM Multimedia1
2008 A real-time, embedded face-annotation system
abstract
Face detection and recognition have numerous multimedia applications of broad interest, one of which is automatic face annotation. There exist many robust algorithms tackling these problems but most of these algorithms are computationally demanding and have only been implemented in PC- or server-based environments. In this demonstration we show a real-time face-annotation system on a commercial PDA development platform. We examine the challenges faced in the design and development of a practical system that can achieve detection and recognition in real-time using limited memory and computational resources which are common constraints for embedded applications.
Shih-Wei Chu, Mei-Chen Yeh, Kwang-Ting Cheng
ACM Multimedia2
2006 Fast Human Detection Using a Cascade of Histograms of Oriented Gradients
abstract
We integrate the cascade-of-rejectors approach with the Histograms of Oriented Gradients (HoG) features to achieve a fast and accurate human detection system. The features used in our system are HoGs of variable-size blocks that capture salient features of humans automatically. Using AdaBoost for feature selection, we identify the appropriate set of blocks, from a large set of possible blocks. In our system, we use the integral image representation and a rejection cascade which significantly speed up the computation. For a 320 × 280 image, the system can process 5 to 30 frames per second depending on the density in which we scan the image, while maintaining an accuracy level similar to existing methods.
Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng, Shai Avidan
CVPR (2)2
2006 Multimodal fusion using learned text concepts for image categorization
abstract
Conventional image categorization techniques primarily rely on low-level visual cues. In this paper, we describe a multimodal fusion scheme which improves the image classification accuracy by incorporating the information derived from the embedded texts detected in the image under classification. Specific to each image category, a text concept is first learned from a set of labeled texts in images of the target category using Multiple Instance Learning [1]. For an image under classification which contains multiple detected text lines, we calculate a weighted Euclidian distance between each text line and the learned text concept of the target category. Subsequently, the minimum distance, along with low-level visual cues, are jointly used as the features for SVM-based classification. Experiments on a challenging image database demonstrate that the proposed fusion framework achieves a higher accuracy than the state-of-art methods for image classification.
Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng
ACM Multimedia2
2005 Manifold learning, a promised land or work in progress?
abstract
Tasks of image clustering and classification often deal with data of very high dimensions. To alleviate the dimensionality curse, several methods, such as isomap, LLE and KPCA, have recently been proposed and applied to learn low-dimensional, non-linear embedded manifolds in high-dimensional spaces. Unfortunately, the scenarios in which these methods appear to be effective are very contrived. In this work, we empirically examine these methods on a realistic but not-so-difficult dataset. We discuss the promises and limitations of these dimension-reduction schemes.
Mei-Chen Yeh, I-Hsiang Lee, Gang Wu 0005, Yi Wu 0005, Edward Y. Chang
ICME1
2002 Scalable ideal-segmented chain coding
abstract
We present an optimal chain-code-like representation to code contour shapes; in addition, this representation can be easily extended to a scalable form which structures shape data as the base layer followed by two enhancement layers. A lossy coding scheme is also presented for low-bit-rate applications. Compared with the block-based CAE (context-based arithmetic encoding) method in MPEG-4 and DCC (differential chain coding) with an arithmetic coder, our method not only has a higher compression ratio with fewer computation steps, but also can be applied to layered transmission. The scalability achieved by our scheme can be recognized both spatially and quality-wise. Compared with other scalable shape coding methods, our scheme is simple but more efficient than progressive polygon encoding methods. In case of a base layer with distortion Dn=0.02, our scheme saves about 20-30% of the number of bits, that is, we have a smaller size base layer, which is important in layered transmission.
Mei-Chen Yeh, Yen-Lin Huang, Jia-Shung Wang
ICIP (1)1