EDBT 2026 Demo / reviewers in the wild / expert
Mei-Chen Yeh
dblp:89/3060
· DBLP profile ↗
38ranked-venue papers
14as first author
8since 2021 · last 2025
0000-0001-8665-7860ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-supervised learning of pseudo classes for generalized zero-shot fine-grained recognition
Yan-He Chen, Mei-Chen Yeh |
Multim. Tools Appl. | 2 |
| 2024 | Self-Supervised Multi-Label Classification with Global Context and Local AttentionabstractSelf-supervised learning has proven highly effective across various tasks, showcasing its versatility in different applications. Despite these achievements, the challenges inherent in multi-label classification have seen limited attention. This paper introduces GAELLE, a novel self-supervised multi-label classification framework that simultaneously captures image context and object information. GAELLE employs a combination of global context and local attention mechanisms to discern diverse levels of semantic information in images. The global component comprehensively learns image content while local attention eliminates object-irrelevant nuances by aligning embeddings with a projection head. The integration of global and local features in GAELLE effectively captures intricate object-scene relationships. To further enhance this capability, we introduce a global and local swap prediction technique, facilitating the nuanced interplay between various objects and scenes within images. Experimental results showcase GAELLE's state-of-the-art performance in self-supervised multi-label classification tasks, highlighting its effectiveness in uncovering complex relationships between multiple objects and scenes in images. Mei-Chen Yeh |
ICMR | 2 |
| 2024 | Indirect visual-semantic alignment for generalized zero-shot recognition
Yan-He Chen, Mei-Chen Yeh |
Multim. Syst. | 2 |
| 2023 | Weakly- and Semi-Supervised Object LocalizationabstractWeakly supervised object localization deals with the lack of location-level labels to train localization models. Recently a new evaluation protocol is proposed in which full supervision is available but limited to only a small validation set. It motives us to explore semi-supervised learning for addressing this problem. In particular, the localization model is developed via self-training: we use a small amount of data with full supervision to train a class-agnostic detector, and use it to generate pseudo bounding boxes for data with weak supervision. Furthermore, we propose a selection algorithm to discover high-quality pseudo labels, and deal with data imbalance caused by pseudo labeling. We demonstrate the superiority of the proposed method with performance on par with the state of the art on two benchmarks. Zhen-Tang Huang, Yan-He Chen, Mei-Chen Yeh |
ICASSP | 3 |
| 2022 | Generative and Adaptive Multi-Label Generalized Zero-Shot LearningabstractWe address the problem of multi-label generalized zero-shot learning where the task is to predict the labels (usually more than one) of a target image whether each of its labels belongs to the seen or unseen category. To alleviate the extreme data-imbalance problem, in which no annotated images are available for unseen classes during training, state-of-the-art single-label zero-shot learning methods learn to synthesize the class-specific visual features from seen classes. However, synthesizing multi-label visual features from multi-label images has not been extensively studied. By exploring the relationship between an image and its labels, we address the multi-label generalized zero-shot learning problem via a hybrid framework of generative and adaptive learning. We convert an image into a label classifier, which can vary among intra-class samples. The adaptive mechanism facilitates the usage of a single-label feature generating model for creating multi-label features from multi-label images. We show that the proposed method improves the state of the art ZSL/GZSL methods on two benchmark datasets. Kuan-Ying Chen, Mei-Chen Yeh |
ICME | 2 |
| 2022 | A Semi-Supervised Learning Approach for Traditional Chinese Scene Text DetectionabstractWith the advancement of multimedia technology, the information in surrounding environment has becoming accessible. In particular, automatic scene text detection is essential for subsequent text recognition, understanding and analysis. However, most existing methods are primarily designed for English, while those for other languages are scarce. In this paper we present a traditional Chinese scene text detector, built upon a robust object detector trained with labeled and unlabeled data via semi-supervised learning. Moreover, we expand the limited labeled data by data synthesis and a data augmentation method. We demonstrate the effectiveness of the proposed method through extensive experiments, and examine the design choices in developing a practical system that can instantly and accurately detect traditional Chinese texts in complex scenes. Chia-Fu Yeh, Mei-Chen Yeh |
MMSP | 2 |
| 2021 | Generalized Zero-Shot Recognition Through Image-Guided Semantic ClassificationabstractWe present a new visual-semantic embedding method for generalized zero-shot learning. Different to existing embedding-based methods that learn the correspondence between an image classifier and its class prototype for each class, we learn the mapping between an image and its semantic classifier. Given an input image, the proposed method creates a label classifier and applies it to all label embeddings to determine whether a label belongs to the input image. Therefore, a semantic classifier is image conditioned and is generated during inference. We validate our approach with four standard benchmark datasets. Mei-Chen Yeh |
ICIP | 2 |
| 2021 | Text-Enhanced Attribute-Based Attention for Generalized Zero-Shot Fine-Grained Image ClassificationabstractWe address the generalized zero-shot fine-grained image classification problem, in which classes are visually similar and training images for some classes are not available. We leverage auxiliary information in the form of textual descriptions to facilitate the task. Specifically, we propose a text-enhanced attribute-based attention mechanism to compute features from the most relevant image regions guided from the most relevant attributes. Experiments on two popular datasets of CUB and AWA2 show the effectiveness of the proposed method. Yan-He Chen, Mei-Chen Yeh |
ICMR | 2 |
| 2020 | Unsupervised Multi-Task Domain AdaptationabstractWith abundant labeled data, deep convolutional neural networks have shown great success in various image recognition tasks. However, these models are often less powerful when applied to novel datasets due to a phenomenon known as domain shift. Unsupervised domain adaptation methods aim to address this problem, allowing deep models trained on the labeled source domain to be used on a different target domain (without labels). In this paper, we investigate whether the generalization ability of an unsupervised domain adaptation method can be improved through multi-task learning, with learned features required to be both domain invariant and discriminative for multiple different but relevant tasks. Experiments evaluating two fundamental recognition tasks-image recognition and segmentation-show that the generalization ability empowered by multi-task learning may not benefit recognition when the model is directly applied on the target domain, but the multi-task learning setting can boost the performance of state-of-the-art unsupervised domain adaptation methods by a non-negligible margin. Shih-Min Yang, Mei-Chen Yeh |
ICPR | 2 |
| 2020 | Artist-based painting classification using Markov random fields with convolution neural network
Kai-Lung Hua, Trang-Thi Ho, Kevin Alfianto Jangtjik, Mei-Chen Yeh |
Multim. Tools Appl. | 5 |
| 2020 | Multilabel Deep Visual-Semantic EmbeddingabstractInspired by the great success from deep convolutional neural networks (CNNs) for single-label visual-semantic embedding, we exploit extending these models for multilabel images. We propose a new learning paradigm for multilabel image classification, in which labels are ranked according to its relevance to the input image. In contrast to conventional CNN models that learn a latent vector representation (i.e., the image embedding vector), the developed visual model learns a mapping (i.e., a transformation matrix) from an image in an attempt to differentiate between its relevant and irrelevant labels. Despite the conceptual simplicity of our approach, the proposed model achieves state-of-the-art results on three public benchmark datasets. Mei-Chen Yeh, Yi-Nan Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Early Detection of Vacant Parking Spaces Using Dashcam VideosabstractA major problem in metropolitan areas is finding parking spaces. Existing parking guidance systems often adopt fixed sensors or cameras that cannot provide information from the driver’s point of view. Motivated by the advent of dashboard cameras (dashcams), we develop neural-network-based methods for detecting vacant parking spaces in videos recorded by a dashcam. Detecting vacant parking spaces in dashcam videos enables early detection of spaces. Different from conventional object detection methods, we leverage the monotonicity of the detection confidence with respect to the distance away of the approaching target parking space and propose a new loss function, which can not only yield improved detection results but also enable early detection. To evaluate our detection method, we create a new large dataset containing 5,800 dashcam videos captured from 22 indoor and outdoor parking lots. To the best of our knowledge, this is the first and largest driver’s view video dataset that supports parking space detection and provides parking space occupancy annotations. Ming-Che Wu, Mei-Chen Yeh |
AAAI | 2 |
| 2019 | A Topological Data Analysis Approach to Video SummarizationabstractThis paper explores the use of simplicial complex to construct a new structure-wise representation for videos. Complementary to the appearance-based representation, which usually involves feature extraction from video content, the proposed method captures the structural properties of a video and can be used for various video processing tasks. We demonstrate a case study of the proposed approach to automated video summarization, which relies only on the developed topological structures and requires no training phase. Experimental results show the potential of this approach to facilitate tasks involving the understanding and analysis of structural information inherent in videos. Chuan-Shen Hu, Mei-Chen Yeh |
ICIP | 2 |
| 2019 | Virtual Portraitist: An Intelligent Tool for Taking Well-Posed SelfiesabstractSmart photography carries the promise of quality improvement and functionality extension in making aesthetically appealing pictures. In this article, we focus on self-portrait photographs and introduce new methods that guide a user in how to best pose while taking a selfie. While most of the current solutions use a post processing procedure to beautify a picture, the developed tool enables a novel function of recommending a good look before the photo is captured. Given an input face image, the tool automatically estimates the pose-based aesthetic score, finds the most attractive angle of the face, and suggests how the pose should be adjusted. The recommendation results are determined adaptively to the appearance and initial pose of the input face. We apply a data mining approach to find distinctive, frequent itemsets and association rules from online profile pictures, upon which the aesthetic estimation and pose recommendation methods are developed. A simulated and a real image set are used for experimental evaluation. The results show the proposed aesthetic estimation method can effectively select user-favorable photos. Moreover, the recommendation performance for the vertical adjustment is moderately related to the degree of conformity among the professional photographers’ recommendations. This study echoes the trend of instant photo sharing, in which a user takes a picture and then immediately shares it on a social network without engaging in tedious editing. Chuan-Shen Hu, Yi-Tsung Hsieh, Hsiao-Wei Lin, Mei-Chen Yeh |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Granularity-based interactive image displayabstractThis paper presents a prototype system that assists users in accessing an unstructured image set (e.g., search results of a query). The system provides a spectrum of overviews, each of which is determined by the display granularity (i.e., the level of summary) an user desires. This new functionality enables a new granularity-based interactive image browsing experience. Geng-Zhi Fann, Mei-Chen Yeh |
ICIP | 2 |
| 2017 | A CNN-LSTM framework for authorship classification of paintingsabstractThe authenticity of digital painting image is an urgent demand in the field of art. Yet, determining the authorship of a certain painting is a challenging task due to two reasons: (1) various artists might share similar painting styles; and (2) an artist could create different styles. In this paper, we present a novel method for authorship classification of paintings based on a CNN-LSTM framework. First, a multiscale pyramid is constructed from a painting image. Second, a CNN-LSTM model is learned and it returns possibly multiple labels for one image. To aggregate the final classification result, an adaptive fusion method is employed. Experimental results show that the proposed method has superior classification performance compared with the state-of-the-art techniques. Kevin Alfianto Jangtjik, Trang-Thi Ho, Mei-Chen Yeh, Kai-Lung Hua |
ICIP | 3 |
| 2017 | The Impact of Feng Shui on House Price: A Data PerspectiveabstractIn this paper, we address interesting questions about how feng shui influences house price from a data perspective. First, is feng shui likely to influence house price? Second, how do different feng shui features, e.g., house shape, master bedroom location, and other interior room arrangements, influence the price? Third, can we automatically diagnose the feng shui problems of a house? From a dataset of 4,579 items of information on houses in Taipei City, we automatically infer what factors determine the house price and identify the feng shui problems of a house that may lower a home's economic value. Tsuo-Chen Wu, Mei-Chen Yeh |
ISM | 2 |
| 2017 | Human Pose Tracking Using Online Latent Structured Support Vector Machine
Kai-Lung Hua, Irawati Nurmala Sari, Mei-Chen Yeh |
MMM (1) | 3 |
| 2016 | A Computational Approach to Finding Facial Patterns of a BabyfaceabstractFacial babyishness has a strong impact on social perceptions and interactions; however, the components constituting a babyface remain unclear. In this paper, we present a computational approach for identifying important but less apparent facial patterns of a babyface, using voluminous face images on the web. The proposed approach is built upon computationally efficient data mining techniques. A new image set with ground truth data collected from users and an evaluation approach based on age estimation are presented in the experiment. The results show that the mined patterns are effective for understanding and determining babyfaces. The findings of this study should provide information for future investigations on the prediction and analysis of trait impressions using the patterns. Zi-Yi Ke, Mei-Chen Yeh |
ICMR | 2 |
| 2016 | Artist-based Classification via Deep Learning with Multi-scale Weighted PoolingabstractFor analyzing digital images of paintings we propose a new approach to categorize them based on artist. Determining the authorship of a painting is challenging because common subjects are illustrated in paintings, and paintings of an artist may not have a unique style. The proposed approach is built upon convolutional neural networks (CNN)---a class of biologically inspired vision model that recently demonstrates near-human performance on several visual recognition tasks. However, training a CNN model requires large scale training data of a fixed input image size (e.g. 224 * 224). In this paper, we propose to construct a multi-layer pyramid from an image, providing 21X more features than using a single layer (i.e., the original image) alone. We train a CNN model for each layer, and propose a new weighted fusion scheme to adaptively combine the decision results. To evaluate the proposed methods, we collect a new painting image dataset, categorized into 13 artists. As demonstrated in the experimental results, the proposed method achieves a promising result---88.08% recall rate in top-2 retrieval on the challenging classification task. Kevin Alfianto Jangtjik, Mei-Chen Yeh, Kai-Lung Hua |
ACM Multimedia | 2 |
| 2016 | A content-based approach for detecting highlights in action movies
Mei-Chen Yeh, Yen-Wei Tsai, Hao-Chen Hsu |
Multim. Syst. | 1 |
| 2016 | Fast medium-scale multiperson identification in aerial videos
Mei-Chen Yeh, Han-Kuen Chiu, Jia-Shung Wang |
Multim. Tools Appl. | 1 |
| 2014 | Virtual Portraitist: Aesthetic Evaluation of Selfies Based on AngleabstractThis work addresses the Huawei Grand Challenge that seeks solutions of quality improvement and functionality extension in computational photography. We propose virtual portraitist-a new method that helps users take good selfies in angle. Dissimilar to current solutions that mostly use a post-processing step to fix a photograph, the proposed method enables a novel function of recommending a good look before the photo is captured. This is achieved by using an automatic approach for estimating the aesthetic quality score of a selfie based on angle. In particular, a set of distinctive patterns discovered from a collection of online profile pictures are combined with head pose and camera orientation to rate the quality of a selfie. Experiments validate the effectiveness of the approach. Mei-Chen Yeh, Hsiao-Wei Lin |
ACM Multimedia | 1 |
| 2014 | Taking good selfies on your phoneabstractSelfies are in vogue because smartphones and social media websites have helped popularize the phenomenon whereby people can easily take and share self-portrait photographs. However, taking good selfies is not always a trivial task for end users. In this demonstration we present a tool for guiding users on which angles to take selfies from. Dissimilar to existing techniques that enhance the aesthetic quality of selfies through a post-processing procedure, the proposed approach focuses on the creation of a good look before the photo is captured. A few distinctive patterns automatically discovered from a collection of online profile pictures are used to score a selfie in angle. The technique should enable new features such as "virtual portraitist" for current consumer cameras and smartphones, providing an interactive approach for capturing selfies. Mei-Chen Yeh, Hsiao-Wei Lin |
ACM Multimedia | 1 |
| 2014 | Fast salient object detection through efficient subwindow search
Mei-Chen Yeh, Chih-Fan Hsu, Chia-Ju Lu |
Pattern Recognit. Lett. | 1 |
| 2013 | Real-time salient object detectionabstractSalient object detection techniques have a variety of multimedia applications of broad interest. However, the detection must be fast to truly aid in these processes. There exist many robust algorithms tackling the salient object detection problem but most of them are computationally demanding. In this demonstration we show a fast salient object detection system implemented in a conventional PC environment. We examine the challenges faced in the design and development of a practical system that can achieve accurate detection in real-time. Chia-Ju Lu, Chih-Fan Hsu, Mei-Chen Yeh |
ACM Multimedia | 3 |
| 2012 | Relative features for photo quality assessmentabstractAutomatic evaluation of photo aesthetic quality is a challenging problem in multimedia computing. Numerous aesthetic features have been proposed in previous works but the features are extracted solely from the photo under evaluation. In this paper, we explore the use of multiple images, and present the relative features that can be easily computed from any score-based features. We show that evaluation on a group basis can facilitate the quality assessment problem. Although the extraction of the new feature is extremely simple, computationally efficient, and requires no training phase, experimental results validate the effectiveness of the proposed approach. Mei-Chen Yeh, Yu-Chen Cheng |
ICIP | 1 |
| 2012 | Automatic cinemagraphs for ranking beautiful scenesabstractThis work addresses the NHK challenge that aims at automatic recognition of beautiful scenes in broadcast programs. We propose a method that can be used to automatically extract beautiful scenes from videos and rank the scenes in terms of beauty. In particular, we introduce cinemagraphs as an alternative manner for presenting beautiful scenes, and a notion of beauty based on the presence of interesting motions. The method is fully automatic, requires no training phase, and produces beautiful scene cinemagraphs as by-products. Yin-Tzu Chan, Hao-Chen Hsu, Po-Yi Li, Mei-Chen Yeh |
ACM Multimedia | 4 |
| 2012 | An approach to automatic creation of cinemagraphsabstractA cinemagraph is a new type of medium that infuses a static image with the dynamics of one particular region. It is in many ways intermediate between a photograph and a video, and has a number of attractive potential applications, such as the creation of dynamic scenes for games and interactive environments. However, creating cinemagraphs is time consuming and requires certain level of proficiency on photo editing techniques. In this paper, we present a fully automatic approach that creates cinemagraphs from video sequences. Specifically, we view cinemagraph construction as a constrained optimization problem that seeks a sub-volume in video with the maximum cumulative flow fields. The problem can be efficiently solved by a branch and-bound search scheme. A user survey is conducted to understand user preferences and demonstrate the performance of the proposed approach. The findings of this study should provide information for various design choices for an easy and versatile authoring tool for cinemagraphs. Mei-Chen Yeh, Po-Yi Li |
ACM Multimedia | 1 |
| 2012 | A tool for automatic cinemagraphsabstractA cinemagraph is a new type of medium that infuses a static image with the dynamics of one or a few particular regions. It is in many ways intermediate between a photograph and a video, and provides a simple, yet expressive way to mix static and dynamic elements from a video clip. However, the process of creating cinemagraphs is usually tedious for end users and requires serious photo editing skills. In this demonstration we show a tool that creates cinemagraphs in a fully automatic manner. The technique should enable new features such as "intelligent cinemagraph mode" for digital cameras that provides an alternative method to capture the moment. Mei-Chen Yeh, Po-Yi Li |
ACM Multimedia | 1 |
| 2011 | A Hierarchical Approach to Practical Beverage Package Recognition
Mei-Chen Yeh, Jason Tai |
PSIVT (1) | 1 |
| 2011 | Fast Visual Retrieval Using Accelerated Sequence MatchingabstractWe present an approach to represent, match, and index various types of visual data, with the primary goal of enabling effective and computationally efficient searches. In this approach, an image/video is represented by an ordered list of feature descriptors. Similarities between such representations are then measured by the approximate string matching technique. This approach unifies visual appearance and the ordering information in a holistic manner with joint consideration of visual-order consistency between the query and the reference instances, and can be used for automatically identifying local alignments between two pieces of visual data. This capability is essential for tasks such as video copy detection where only small portions of the query and the reference videos are similar. To deal with large volumes of data, we further show that this approach can be significantly accelerated along with a dedicated indexing structure. Extensive experiments on various visual retrieval and classification tasks demonstrate the superior performance of the proposed techniques compared to existing solutions. Mei-Chen Yeh, Kwang-Ting Cheng |
IEEE Trans. Multim. | 1 |
| 2009 | A compact, effective descriptor for video copy detectionabstractLarge scale video copy detection tasks require a compact and computational-efficient descriptor that is robust to various transformations that are typically applied to generate copies. In this paper, we propose a new frame-level descriptor for such a task. The descriptor encodes the internal structure of a video frame by computing the pair-wise correlations between geometrically pre-indexed blocks. It is conceptually simple, small in size, and fast to compute. Experiments using the MUSCLE VCD benchmark show its superior performance compared to existing approaches. Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 1 |
| 2008 | A real-time, embedded face-annotation systemabstractFace detection and recognition have numerous multimedia applications of broad interest, one of which is automatic face annotation. There exist many robust algorithms tackling these problems but most of these algorithms are computationally demanding and have only been implemented in PC- or server-based environments. In this demonstration we show a real-time face-annotation system on a commercial PDA development platform. We examine the challenges faced in the design and development of a practical system that can achieve detection and recognition in real-time using limited memory and computational resources which are common constraints for embedded applications. Shih-Wei Chu, Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 2 |
| 2006 | Fast Human Detection Using a Cascade of Histograms of Oriented GradientsabstractWe integrate the cascade-of-rejectors approach with the Histograms of Oriented Gradients (HoG) features to achieve a fast and accurate human detection system. The features used in our system are HoGs of variable-size blocks that capture salient features of humans automatically. Using AdaBoost for feature selection, we identify the appropriate set of blocks, from a large set of possible blocks. In our system, we use the integral image representation and a rejection cascade which significantly speed up the computation. For a 320 × 280 image, the system can process 5 to 30 frames per second depending on the density in which we scan the image, while maintaining an accuracy level similar to existing methods. Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng, Shai Avidan |
CVPR (2) | 2 |
| 2006 | Multimodal fusion using learned text concepts for image categorizationabstractConventional image categorization techniques primarily rely on low-level visual cues. In this paper, we describe a multimodal fusion scheme which improves the image classification accuracy by incorporating the information derived from the embedded texts detected in the image under classification. Specific to each image category, a text concept is first learned from a set of labeled texts in images of the target category using Multiple Instance Learning [1]. For an image under classification which contains multiple detected text lines, we calculate a weighted Euclidian distance between each text line and the learned text concept of the target category. Subsequently, the minimum distance, along with low-level visual cues, are jointly used as the features for SVM-based classification. Experiments on a challenging image database demonstrate that the proposed fusion framework achieves a higher accuracy than the state-of-art methods for image classification. Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng |
ACM Multimedia | 2 |
| 2005 | Manifold learning, a promised land or work in progress?abstractTasks of image clustering and classification often deal with data of very high dimensions. To alleviate the dimensionality curse, several methods, such as isomap, LLE and KPCA, have recently been proposed and applied to learn low-dimensional, non-linear embedded manifolds in high-dimensional spaces. Unfortunately, the scenarios in which these methods appear to be effective are very contrived. In this work, we empirically examine these methods on a realistic but not-so-difficult dataset. We discuss the promises and limitations of these dimension-reduction schemes. Mei-Chen Yeh, I-Hsiang Lee, Gang Wu 0005, Yi Wu 0005, Edward Y. Chang |
ICME | 1 |
| 2002 | Scalable ideal-segmented chain codingabstractWe present an optimal chain-code-like representation to code contour shapes; in addition, this representation can be easily extended to a scalable form which structures shape data as the base layer followed by two enhancement layers. A lossy coding scheme is also presented for low-bit-rate applications. Compared with the block-based CAE (context-based arithmetic encoding) method in MPEG-4 and DCC (differential chain coding) with an arithmetic coder, our method not only has a higher compression ratio with fewer computation steps, but also can be applied to layered transmission. The scalability achieved by our scheme can be recognized both spatially and quality-wise. Compared with other scalable shape coding methods, our scheme is simple but more efficient than progressive polygon encoding methods. In case of a base layer with distortion Dn=0.02, our scheme saves about 20-30% of the number of bits, that is, we have a smaller size base layer, which is important in layered transmission. Mei-Chen Yeh, Yen-Lin Huang, Jia-Shung Wang |
ICIP (1) | 1 |