VLDB 2026 Research / reviewers in the wild / expert
Wei-Ta Chu
dblp:57/5913
· DBLP profile ↗
101ranked-venue papers
60as first author
30since 2021 · last 2026
0000-0001-5722-7239ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 56 first-author · 23 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion
Jui-Feng Chi, Wei-Ta Chu, Sheng-Long Lin |
MMM (1) | 2 |
| 2026 | Human-Object Interaction detection enhanced by common sense and contextual cues
Cheng-Kang Tan, Wei-Ta Chu |
Pattern Recognit. | 2 |
| 2025 | Multimodal Fusion for Dementia Detection using Voice and Facial FeaturesabstractEarly detection of dementia is important in the aging society. Research has shown that early diagnosis and treatment can effectively slow down cognitive decline in the elderly. Our goal is to develop a low-cost dementia detection method by analyzing videos capturing the progress of potential patients taking the Short Portable Mental Status Questionnaire (SPMSQ). We propose a multimodal fusion method that effectively predicts the degree of dementia based on both voice and facial features. Because of lacking enough training data, we propose a feature augmentation method based on the mix-up technique to train a more effective model. Experimental results demonstrate the effectiveness of the proposed method. Jia-Yi Chen, Wei-Ta Chu |
ISCAS | 2 |
| 2025 | MapLlama: A Two-Stage Approach for Map Question Answering Using a Fine-tuned Large Language ModelabstractA map is a common medium to convey rich information. People present data on maps and answer spatial-related questions, usually requiring the ability to analyze the map and work on reasoning. In this work, we propose a two-stage map question-answering (MapQA) system. First, map images are converted into structured data based on a query-based model. A fine-tuned large language model (LLM) is employed with the structured text data to answer the given questions. We point out several issues in the current MapQA, propose a realistic scenario, and tackle it with a fine-tuned LLM. The experimental results show that the proposed method works well for diverse questions and performs better than the state-of-the-art techniques. Zong-Lin Li, Wei-Ta Chu |
ISCAS | 2 |
| 2025 | SuPACape: Graph-based Category-Agnostic Pose Estimation with Super-Category and Pose AdaptivityabstractGiven a support image with keypoint annotations, a category-agnostic pose estimation (CAPE) method aims to predict keypoints in a query image from the same category. The relationship between keypoint features of the support image and the visual features of the query image is discovered to achieve CAPE. Although many CAPE methods have been proposed, two challenges remain: 1) Support information is limited, especially when only a few shots of support images are given; 2) the pose in the support image may be significantly different from that in the query image. For the first issue, we propose to consider keypoint features from the categories in the same super-category as the support image to enhance the support representations. For the second issue, we propose a query-adaptive adjacency matrix generated based on the similarity between features of support keypoints and the query image. Experimental results verify that both ideas bring performance gain and make the proposed method achieve state-of-the-art results on the MP-100 benchmark. Yi-Hsuan Lu, Wei-Ta Chu |
MMAsia | 2 |
| 2025 | HCV: Lightweight Hybrid CNN-Vision Transformer for Visual Object Tracking
Liang-Chia Chen, Wei-Ta Chu |
MMM (2) | 2 |
| 2025 | ALSA-UAD: Unsupervised anomaly detection on histopathology images using adversarial learning and simulated anomaly
Yu-Chen Lai, Wei-Ta Chu |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Multiple Player Tracking With 3D Projection and Spatio-Temporal Information In Multi-View Sports VideosabstractPlayer tracking is a fundamental task in sports video understanding. Many technical challenges should be addressed due to irregular movement, occlusion between players, and complex background. In this work, we present a framework that utilizes synchronized videos captured from multiple view-points. We construct 2D player trajectories from each video, construct 3D player trajectories based on multiple videos, and then associate 2D and 3D trajectories to achieve multiple player tracking. Experimental results show that the proposed method achieves SOTA performance on a volleyball dataset and a basketball dataset. Yi-Peng Wang, Wei-Ta Chu |
ICASSP | 2 |
| 2024 | Transformer-Based Clipped Contrastive Quantization Learning For Unsupervised Image RetrievalabstractUnsupervised image retrieval aims to learn the important visual characteristics without any given level to retrieve the similar images for a given query image. The Convolutional Neural Network (CNN)-based approaches have been extensively exploited with self-supervised contrastive learning for image hashing. However, the existing approaches suffer due to lack of effective utilization of global features by CNNs and biased-ness created by false negative pairs in the contrastive learning. In this paper, we propose a TransClippedCLR model by encoding the global context of an image using Transformer having local context through patch based processing, by generating the hash codes through product quantization and by avoiding the potential false negative pairs through clipped contrastive learning. The proposed model is tested with superior performance for unsupervised image retrieval on benchmark datasets, including CIFAR10, NUS-Wide and Flickr25K, as compared to the recent state-of-the-art deep models. The results using the proposed clipped contrastive learning are greatly improved on all datasets as compared to same backbone network with vanilla contrastive learning. Ayush Dubey, Shiv Ram Dubey, Satish Kumar Singh, Wei-Ta Chu |
ICIP | 4 |
| 2024 | CS-HOI: Human Object Interaction Detection Enhanced by Common Sense
Cheng-Kang Tan, Wei-Ta Chu |
MMAsia | 2 |
| 2024 | Incremental Few-Shot Object Detection by Leveraging External Information from Large Multimodal Models
Guan-Yu Wu, Wei-Ta Chu |
MMAsia | 2 |
| 2024 | Frequency disentangled residual network
Satya Rajendra Singh, Roshan Reddy Yedla, Shiv Ram Dubey, Rakesh Kumar Sanodiya, Wei-Ta Chu |
Multim. Syst. | 5 |
| 2024 | Overall positive prototype for few-shot open-set recognition
Liang-Yu Sun, Wei-Ta Chu |
Pattern Recognit. | 2 |
| 2024 | Positive and Negative Set Designs in Contrastive Feature Learning for Temporal Action SegmentationabstractWhen data labels are scarce, contrastive learning is often used to learn representations in a weakly-supervised or unsupervised way. In contrastive learning, not only the learning mechanism, but also the designs of positive and negative sets are critical. While most previous works of Temporal Action Segmentation (TAS) focus on designing new segmentation methods, we investigate the importance of positive and negative set designs in contrastive learning and verify that better representations can be learned to enhance performance of existing TAS methods. Specific to timestamp-supervised TAS and unsupervised TAS, respectively, we propose positive/negative set designs, associated with the ideas of ambiguous frames and the set expansion process to make learned representations more effective. In the evaluation, we demonstrate that performance of timestamp-supervised TAS can be boosted by 8% to 15% in terms of F1@10 across three different datasets, and the performance of unsupervised TAS can be boosted by 3% to 5% in terms of F1 scores, achieving new state-of-the-art TAS results. Wei-Ta Chu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A Trajectory-based Statistics and Tactics Analysis System for Table TennisabstractFor table tennis videos, we develop a system to analyze and generate statistics based on ball trajectories. By a real-time ball detector, the ball trajectory is constructed based on the tracking by detection scheme. Landing points on the table are estimated. Based on moving direction and the sequence of landing points, three-stage analysis can be achieved. We also analyze how a point starts (serving type classification) and how a point ends (point loss classification). Guan-Yu Wu, Chun-Ho Hung, Hsuan-Wei Chen, Wei-Ta Chu |
MMAsia | 4 |
| 2023 | Occlusion-Aware Manga Character Re-identification with Self-Paced Contrastive LearningabstractExisting methods for manga character re-identification primarily rely on facial information, overlooking the unique characteristics of characters’ bodies and failing to address common challenges like occlusion by speech balloons and incomplete body parts. To tackle these issues, we propose a method called Occlusion-Aware Manga Character Re-identification (OAM-ReID) with self-paced contrastive learning, which leverages annotated body data from the Manga109 dataset for training. By synthesizing data with occluded speech balloons and incomplete bodies, we empower the framework to be aware of occlusion, so that more effective feature representations are learnt. Experimental results show that this approach outperforms the state-of-the-art person ReID method. Ci-Yin Zhang, Wei-Ta Chu |
MMAsia | 2 |
| 2023 | The NCKU-VTF Dataset and a Multi-scale Thermal-to-Visible Face Synthesis System
Tsung-Han Ho, Chen-Yin Yu, Tsai-Yen Ko, Wei-Ta Chu |
MMM (1) | 4 |
| 2023 | Manga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion Analysis
Yi-Ting Yang, Wei-Ta Chu |
MMM (2) | 2 |
| 2023 | SSSD: Self-Supervised Self DistillationabstractWith labeled data, self distillation (SD) has been proposed to develop compact but effective models without a complex teacher model available in advance. Such approaches need labeled data to guide the self distillation process. Inspired by self-supervised (SS) learning, we propose a self-supervised self distillation (SSSD) approach in this work. Based on an unlabeled image dataset, a model is constructed to learn visual representations in a self-supervised manner. This pre-trained model is then adopted to extract visual representations of the target dataset and generates pseudo labels via clustering. The pseudo labels guide the SD process, and thus enable SD to proceed in an unsupervised way (no data labels are required at all). We verify this idea based on evaluations on the CIFAR-10, CIFAR-100, and ImageNet-1K datasets, and demonstrate the effectiveness of this unsupervised SD approach. Performance outperforming similar frameworks is also shown. Wei-Chi Chen, Wei-Ta Chu |
WACV | 2 |
| 2022 | Vision Transformer Hashing for Image RetrievalabstractRecently, Transformer has emerged as a new architecture in deep learning by utilizing self-attention without convolution. Transformer is also extended to Vision Transformer (ViT) for the visual recognition with a promising performance on ImageNet. In this paper, we propose a Vision Transformer Hashing (VTS) for image retrieval. We utilize the pre-trained ViT on ImageNet as the backbone network and add the hashing head. The proposed VTS model is fine tuned for hashing under six different image retrieval frameworks with their objective functions. We perform the extensive experiments on CIFAR10, ImageNet, NUS-Wide, and COCO datasets. The proposed VTS based image retrieval outperforms the recent state-of-the-art hashing techniques with a significant margin. We also find the proposed VTS model as the backbone network is better than the existing networks, such as AlexNet and ResNet. The code is released at https://github.com/shivram1987/VisionTransformerHashing. Shiv Ram Dubey, Satish Kumar Singh, Wei-Ta Chu |
ICME | 3 |
| 2022 | Multimodal Fusion with Cross-Modal Attention for Action Recognition in Still ImagesabstractWe propose a cross-modal attention module to combine information from different cues and different modalities, to achieve action recognition in still images. Feature maps are extracted from the entire image, the detected human bounding box, and the detected human skeleton, respectively. Inspired by the transformer structure, we design the processing between the query vector from one cue/modality, and the key vector from another cue/modality. Feature maps from different cues/modalities are cross-referred so that better representations can be obtained to yield better performance. We show that the proposed framework outperforms the state-of-the-art systems without the requirement of an extra training dataset. We also conduct ablation studies to investigate how different settings impact the final results. Jia-Hua Tsai, Wei-Ta Chu |
MMAsia | 2 |
| 2022 | Indie Games Popularity Prediction by Considering Multimodal Features
Yu-Heng Huang, Wei-Ta Chu |
MMM (2) | 2 |
| 2022 | An imitation learning framework for generating multi-modal trajectories from unstructured demonstrations
Jain-Wei Peng, Min-Chun Hu 0001, Wei-Ta Chu |
Neurocomputing | 3 |
| 2022 | Instant Basketball Defensive Trajectory GenerationabstractTactic learning in virtual reality (VR) has been proven to be effective for basketball training. Endowed with the ability of generating virtual defenders in real time according to the movement of virtual offenders controlled by the user, a VR basketball training system can bring more immersive and realistic experiences for the trainee. In this article, an autoregressive generative model for instantly producing basketball defensive trajectory is introduced. We further focus on the issue of preserving the diversity of the generated trajectories. A differentiable sampling mechanism is adopted to learn the continuous Gaussian distribution of player position. Moreover, several heuristic loss functions based on the domain knowledge of basketball are designed to make the generated trajectories assemble real situations in basketball games. We compare the proposed method with the state-of-the-art works in terms of both objective and subjective manners. The objective manner compares the average position, velocity, and acceleration of the generated defensive trajectories with the real ones to evaluate the fidelity of the results. In addition, more high-level aspects such as the empty space for offender and the defensive pressure of the generated trajectory are also considered in the objective evaluation. As for the subjective manner, visual comparison questionnaires on the proposed and other methods are thoroughly conducted. The experimental results show that the proposed method can achieve better performance than previous basketball defensive trajectory generation works in terms of different evaluation metrics. Wen-Cheng Chen, Wan-Lun Tsai, Huan-Hua Chang, Min-Chun Hu 0001, Wei-Ta Chu |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2021 | Multi-Class Novelty Detection with Generated Hard Novel Features
Wei-Ta Chu, Wei-Ting Cao |
BMVC | 1 |
| 2021 | Searching by Generating: Flexible and Efficient One-Shot NAS With Architecture GeneratorabstractIn one-shot NAS, sub-networks need to be searched from the supernet to meet different hardware constraints. However, the search cost is high and N times of searches are needed for N different constraints. In this work, we propose a novel search strategy called architecture generator to search sub-networks by generating them, so that the search process can be much more efficient and flexible. With the trained architecture generator, given target hardware constraints as the input, N good architectures can be generated for N constraints by just one forward pass without re-searching and supernet retraining. Moreover, we propose a novel single-path supernet, called unified supernet, to further improve search efficiency and reduce GPU memory consumption of the architecture generator. With the architecture generator and the unified supernet, we propose a flexible and efficient one-shot NAS framework, called Searching by Generating NAS (SGNAS). With the pre-trained supernt, the search time of SGNAS for N different hardware constraints is only 5 GPU hours, which is 4N times faster than previous SOTA single-path methods. After training from scratch, the top1-accuracy of SGNAS on ImageNet is 77.1%, which is comparable with the SOTAs. The code is available at: https://github.com/eric8607242/SGNAS. Sian-Yao Huang, Wei-Ta Chu |
CVPR | 2 |
| 2021 | PONAS: Progressive One-shot Neural Architecture Search for Very Efficient DeploymentabstractWe propose a Progressive One-Shot Neural Architecture Search (PONAS) method to achieve a very efficient model searching for various hardware constraints. Given a constraint, most neural architecture search (NAS) methods either sample a set of sub-networks according to a pre-trained accuracy predictor, or adopt the evolutionary algorithm to evolve specialized networks from the supernet. Both approaches are time consuming. Here our key idea for very efficient deployment is, when searching the architecture space, constructing a table that stores the validation accuracy of all candidate blocks at all layers. For a stricter hardware constraint, the architecture of a specialized network can be efficiently determined based on this table by picking the best candidate blocks that yield the least accuracy loss. To accomplish this idea, we propose the PONAS method to combine advantages of progressive NAS and one-shot methods. A two-stage training scheme, including the meta training stage and the fine-tuning stage, is proposed to make the search process efficient and stable. During search, we evaluate candidate blocks in different layers and construct an accuracy table that is to be used in architecture searching. Comprehensive experiments verify that PONAS is extremely flexible, and is able to find architecture of a specialized network in around 10 seconds. In ImageNet classification, 76.29% top-1 accuracy can be obtained, which is comparable with the state of the arts. Sian-Yao Huang, Wei-Ta Chu |
IJCNN | 2 |
| 2021 | Automatic Baseball Pitch OverlayabstractTo provide rich viewing experience and assist pitcher training, we propose an automatic baseball pitch overlay system in this paper. Given multiple pitching video sequences, this system detects and tracks the ball to construct ball trajectories. Because of occlusion, motion blur, and background noise, the ball usually cannot be detected successfully. We propose a series of processes like initial compensation and polynomial fitting to construct complete trajectories. To make the overlay results more appealing, different sequences are weighted differently, and different trajectories are intentionally drawn in different colors. We believe this would be the first fully-automatic pitch overlay system that only takes pitching videos as inputs. Source code is at \\https://github.com/chonyy/ML-auto-baseball-pitching-overlay. Ting-Hsuan Chou, Wei-Ta Chu |
ICMR | 2 |
| 2021 | Thermal Face Recognition Based on Multi-scale Image Synthesis
Wei-Ta Chu, Ping-Shen Huang |
MMM (1) | 1 |
| 2021 | Multi-label image recognition by using semantics consistency, object correlation, and multiple samples
Wei-Ta Chu, Si-Heng Huang |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020abstractThe International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173 Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki |
ICMR | 1 |
| 2020 | An autoregressive generation model for producing instant basketball defensive trajectoryabstractLearning basketball tactic via virtual reality environment requires real-time feedback to improve the realism and interactivity. For example, the virtual defender should move immediately according to the player's movement. In this paper, we proposed an autoregressive generative model for basketball defensive trajectory generation. To learn the continuous Gaussian distribution of player position, we adopt a differentiable sampling process to sample the candidate location with a standard deviation loss, which can preserve the diversity of the trajectories. Furthermore, we design several additional loss functions based on the domain knowledge of basketball to make the generated trajectories match the real situation in basketball games. The experimental results show that the proposed method can achieve better performance than previous works in terms of different evaluation metrics. Huan-Hua Chang, Wen-Cheng Chen, Wan-Lun Tsai, Min-Chun Hu 0001, Wei-Ta Chu |
MMAsia | 5 |
| 2020 | Thermal Face Recognition Based on Transformation by Residual U-Net and Pixel Shuffle Upsampling
Soumya Chatterjee 0002, Wei-Ta Chu |
MMM (1) | 2 |
| 2019 | A Genetic Programming Approach to Integrate Multilayer CNN Features for Image Classification
Wei-Ta Chu, Hao-An Chu |
MMM (1) | 1 |
| 2019 | Photo Filter Classification and Filter Recommendation without Much Manual LabelingabstractBecause how users employ filters to photos may reveal user's preference or mental state, a photo filter classification method is potentially demanded to enable future large-scale analysis. We adopt the transfer learning technique to transform deep models pre-trained for object classification into models suitable for photo filter classification. Based on accurate classification results, we build a filter recommendation approach without much manual labeling. It can be easily extended when more training data are available. A series of experimental studies are conducted to demonstrate effectiveness of filter classification with transfer learning. We also demonstrate the proposed filter recommendation achieves encouraging performance. Wei-Ta Chu, Yu-Tzu Fan |
MMSP | 1 |
| 2019 | Thermal Facial Landmark Detection by Deep Multi-Task LearningabstractWe present a neural network to jointly consider facial landmark detection and emotion recognition for thermal face images. The first part of this network is based on the U-Net structure, targeting at extracting good features for advanced analysis. Using U - Net as the basic structure enables modeling context information based on a limited number of training data. The second part of this network contains two branches that are designed for landmark detection and emotion recognition, respectively. We propose a two-stage training mechanism to learn this network, and demonstrate the effectiveness of the proposed approach. This work is believed to be one of the few studies on thermal face image analysis. Wei-Ta Chu, Yu-Hui Liu |
MMSP | 1 |
| 2019 | Spatiotemporal Modeling and Label Distribution Learning for Video SummarizationabstractFor a video which content does not follow specific production rules, or without professional editing, at least two problems should be solved to generate a good video summary. First, the summarization system should jointly model visual content in the spatial domain and visual dynamics in the temporal domain. Second, the system should consider the inconsistency between users, i.e., different users may annotate the same video segment with different importance scores. In this paper, we present a video summarization system that models spatiotemporal information of video segments, and predicts the distribution of importance scores for each segment. Based on the estimated importance scores, video summaries are generated by picking the ones with higher scores. We especially demonstrate the effectiveness of label distribution learning based on two video benchmarks. Wei-Ta Chu, Yu-Hsin Liu |
MMSP | 1 |
| 2019 | Manga face detection based on deep neural networks fusing global and local information
Wei-Ta Chu |
Pattern Recognit. | 1 |
| 2018 | A Parametric Study of Deep Perceptual Model on Visible to Thermal Face RecognitionabstractRecently deep perceptual mapping (DPM) based on auto-encoder provides the state-the-art thermal to visible face recognition. Features extracted from patches of a long-wave infra-red (LWIR) face image are transformed into a space by an auto-encoder, such that features from infra-red images are comparable with features from visible images. In this paper, we comprehensively evaluate DPM with different settings, in order to build a reference study for future research. Wei-Ta Chu, Jo-Ning Wu |
VCIP | 1 |
| 2018 | Text Detection in Manga by Deep Region Proposal, Classification, and RegressionabstractText in manga presents high variations and different contextual information, and existing scene text detection methods are not directly applicable. We propose two approaches based on deep networks to detect text in manga. In the first approach, features extracted from multiple CNNs are joined and then fed to a combination of a classification network and a regression network. In the second approach, region proposal, feature extraction, and classification/regression, are taken together in a single deep network. The evaluation results show that the first approach achieves performance comparable to the current state of the art, while the second approach yields a big performance leap over existing ones. Wei-Ta Chu, Chih-Chi Yu |
VCIP | 1 |
| 2018 | Visual Weather Temperature PredictionabstractIn this paper, we attempt to employ convolutional recurrent neural networks for weather temperature estimation using only image data. We study ambient temperature estimation based on deep neural networks in two scenarios a) estimating temperature of a single outdoor image, and b) predicting temperature of the last image in an image sequence. In the first scenario, visual features are extracted by a convolutional neural network trained on a large-scale image dataset. We demonstrate that promising performance can be obtained, and analyze how volume of training data influences performance. In the second scenario, we consider the temporal evolution of visual appearance, and construct a recurrent neural network to predict the temperature of the last image in a given image sequence. We obtain better prediction accuracy compared to the state-of-the-art models. Further, we investigate how performance varies when information is extracted from different scene regions, and when images are captured in different daytime hours. Our approach further reinforces the idea of using only visual information for cost efficient weather prediction in the future. Wei-Ta Chu, Kai-Chia Ho, Ali Borji |
WACV | 1 |
| 2018 | Image Style Classification Based on Learnt Deep Correlation FeaturesabstractThis paper presents a comprehensive study of deep correlation features on image style classification. Inspired by that, correlation between feature maps can effectively describe image texture, and we design various correlations and transform them into style vectors, and investigate classification performance brought by different variants. In addition to intralayer correlation, interlayer correlation is proposed as well, and its effectiveness is verified. After showing the effectiveness of deep correlation features, we further propose a learning framework to automatically learn correlations between feature maps. Through extensive experiments on image style classification and artist classification, we demonstrate that the proposed learnt deep correlation features outperform several variants of convolutional neural network features by a large margin, and achieve the state-of-the-art performance. Wei-Ta Chu, Yi-Ling Wu |
IEEE Trans. Multim. | 1 |
| 2017 | Blog Article Summarization with Image-Text Alignment TechniquesabstractWe propose an image-text alignment framework to match images with text, and take blog article summarization as the main application. Objects in an image are first detected, from them deep features are extracted and transformed into a space commonly shared with the text. On the other hand, sentences of a blog article are represented as vectors, and are also embedded into the common space. With these processes, cross-modal matching can be achieved. A blog article is then summarized in the representation of images and their matched sentences. In evaluation, we demonstrate the effectiveness of the proposed method, and show that the generated summary makes more sense. Wei-Ta Chu, Ming-Chih Kao |
ISM | 1 |
| 2017 | Manga FaceNet: Face Detection in Manga based on Deep Neural NetworkabstractAmong various elements of manga, character's face plays one of the most important role in access and retrieval. We propose a DNN-based method to do manga face detection, which is a challenging but relatively unexplored topic. Given a manga page, we first find candidate regions based on the selective search scheme. A deep neural network is then proposed to detect manga faces of various appearance. We evaluate the proposed method based on a large-scale benchmark, and show performance comparison and convincing evaluation results that have rarely done before. Wei-Ta Chu |
ICMR | 1 |
| 2017 | Badminton Video Analysis based on Spatiotemporal and Stroke FeaturesabstractMost of the broadcasted sports events nowadays present game statistics to the viewers which can be used to design the gameplay strategy, improve player's performance, or improve accessing the point of interest of a sport game. However, few studies have been proposed for broadcasted badminton videos. In this paper, we integrate several visual analysis techniques to detect the court, detect players, classify strokes, and classify the player's strategy. Based on visual analysis, we can get some insights about the common strategy of a certain player. We evaluate performance of stroke classification, strategy classification, and show game statistics based on classification results. Wei-Ta Chu, Samuel I. G. Situmeang |
ICMR | 1 |
| 2017 | Camera as weather sensor: Estimating weather information from single images
Wei-Ta Chu, Xiang-You Zheng, Ding-Shiuan Ding |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | On broadcasted game video analysis: event detection, highlight detection, and highlight forecast
Wei-Ta Chu, Yung-Chieh Chou |
Multim. Tools Appl. | 1 |
| 2017 | Cultural difference and visual information on hotel rating prediction
Wei-Ta Chu, Wei-Han Huang |
World Wide Web | 1 |
| 2017 | A hybrid recommendation system considering visual information for predicting favorite restaurants
Wei-Ta Chu, Ya-Lun Tsai |
World Wide Web | 1 |
| 2016 | Manga-specific features and latent style model for manga style analysisabstractA latent style model describing manga styles based on the proposed manga-specific features is constructed to facilitate novel style-based applications. Two manga-specific features, i.e., screentone features showing texture and shade, and panel features showing panel arrangement, are firstly proposed to describe manga pages. Based on the latent Dirichlet allocation technique, we discover latent style elements embedded in manga documents, which are described by visual words derived from manga-specific features. Distributions of style elements are then used to measure similarity between manga documents, and facilitate the development ofvarious style-based applications. Experimental results show that the features and models especially designed for describing manga styles yield promising performance and could bring many potential extensions. Wei-Ta Chu, Wei-Chung Cheng |
ICASSP | 1 |
| 2016 | News story clustering with fisher embeddingabstractAn automatic news story clustering system is presented to facilitate efficient news browsing and summarization. We describe news content by considering both what objects appear and how these objects move in news stories. With Fisher embedding, we respectively encode local features, semantics features, and dense trajectories as Fisher vectors, based on which similarity between news stories can be well evaluated and thus better clustering performance can be obtained. We verify the effectiveness of Fisher encoding, and further show that motion-based features are more effective than appearance-based features through feature analysis. Wei-Ta Chu, Han-Nung Hsu |
ICASSP | 1 |
| 2016 | Deep Correlation Features for Image Style ClassificationabstractThis paper presents a comprehensive study of deep correlation features on image style classification. Inspired by that correlation between feature maps can effectively describe image texture, we design and transform various such correlations into style vectors, and investigate classification performance brought by different variants. In addition to intra-layer correlation, we also propose inter-layer correlation and verify its benefit. Through extensive experiments on image style classification and artist classification, we demonstrate that the proposed style vectors significantly outperforms CNN features coming from fully-connected layers, as well as outperforms the state-of-the-art deep representation. Wei-Ta Chu, Yi-Ling Wu |
ACM Multimedia | 1 |
| 2016 | Predicting Occupation from Images by Combining Face and Body Context InformationabstractFacial images embed age, gender, and other rich information that is implicitly related to occupation. In this work, we advocate that occupation prediction from a single facial image is a doable computer vision problem. We extract multilevel hand-crafted features associated with locality-constrained linear coding and convolutional neural network features as image occupation descriptors. To avoid the curse of dimensionality and overfitting, a boost strategy called multichannel SVM is used to integrate features from face and body. Intra- and interclass visual variations are jointly considered in the boosting framework to further improve performance. In the evaluation, we verify the effectiveness of predicting occupation from face and demonstrate promising performance obtained by combining face and body information. More importantly, our work further integrates deep features into the multichannel SVM framework and shows significantly better performance over the state of the art. Wei-Ta Chu, Chih-Hao Chiu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2015 | A Privacy-Preserving Bipartite Graph Matching Framework for Multimedia Analysis and RetrievalabstractThe emergence of cloud computing provides an unlimited computation/storage for users, and yields new opportunities for multimedia analysis and retrieval research. However, privacy of users, e.g., search intention, may be leaked to the server and maliciously utilized by companies or individuals with animus. This paper presents a privacy-preserving multimedia analysis framework based on a widely-adopted structure, i.e., bipartite graph, so that multimedia analysis and retrieval in the encrypted domain is enabled. This work aims to keep the server unaware of what the user wants to retrieve, and at the same time take advantage of the server's computation power. Homomorphic encryption schemes and communication protocols in the encrypted domain are integrated to facilitate bipartite graph construction and implement the Hungarian algorithm to find the best matching. Two applications, video tag suggestion and video copy detection, are developed on top of the privacy-preserving framework, and the evaluation results demonstrate that performance obtained in the encrypted domain is comparable with that obtained in the plain text domain. Wei-Ta Chu, Feng-Chi Chang |
ICMR | 1 |
| 2015 | Street sweeper: detecting and removing cars in street view images
Wei-Ta Chu, Ying-Chieh Chao, Yi-Sheng Chang |
Multim. Tools Appl. | 1 |
| 2015 | Optimized Comics-Based Storytelling for Temporal Image SequencesabstractWe propose a system to transform any temporal image sequence into a comics-based presentation, as an effective and interesting storytelling manner. Three main components, including page allocation, layout selection, and speech balloon placement, are respectively formulated as optimization problems, and systematic approaches are proposed to find solutions. Page allocation is viewed as a labeling problem, and the best solution is determined by the genetic algorithm. Importance values of images and predefined layouts are both represented in vector forms, and the best layout is selected by finding the best match between vectors. Feasible solutions of speech balloons constitute a solution space, and the best solution that jointly describes the best locations of all balloons in a page is determined by the particle swarm optimization algorithm. Objective evaluation and subjective evaluation are designed from various perspectives to demonstrate effectiveness and superiority of the proposed system. Wei-Ta Chu, Chia-Hsiang Yu, Hsin-Han Wang |
IEEE Trans. Multim. | 1 |
| 2014 | Predicting Occupation from Single Facial ImagesabstractFacial images embed age, gender, and other rich information that is implicitly related to occupation. In this work, we advocate that occupation prediction from a single facial image is a doable research direction. We first extract visual features from multiple levels of patches and describe them by locality-constrained linear coding. To avoid the curse of dimensionality and over fitting, a boost strategy called multi-feature SVM is used to integrate features. Intra-class and inter-class visual variations are jointly considered in the boosting framework to further improve performance. In the evaluation, we verify that this is a promising research topic with encouraging performance, and also discuss interesting issues from various perspectives. Wei-Ta Chu, Chih-Hao Chiu |
ISM | 1 |
| 2014 | Fast Object Detection Using Multistage Particle Window Deformable Part ModelabstractFor object detection, evaluating all sliding windows at various scales draws a computational efficiency issue. In this paper, we propose a fast object detection framework using the multistage particle window strategy to accelerate the cascade deformable part model (DPM). Coupling this strategy with the proposed early jump scheme, adaptive particle window generation, and efficient preprocessing, we demonstrate that the proposed method runs 34.5 times faster than the conventional DPM to detect objects in images, and is able to efficiently detect vehicles and pedestrians in on-road videos. Wei-Ta Chu, Ming-Hung Hsu |
ISM | 1 |
| 2014 | Line-Based Drawing Style Description for Manga ClassificationabstractDiversity of drawing styles of mangas can be easily perceived by humans, but are hard to be described in text. We design computational features derived from line segments to describe drawing styles, enabling style classification such as discriminating mangas targeting youth boys and youth girls, and discriminating artworks produced by different artists. With statistical analysis, we found that drawing styles can be effectively characterized by the proposed features such that explicit (e.g., density of line segments) or implicit (e.g., included angles between lines) observations can be made to facilitate various manga style classification. Wei-Ta Chu, Ying-Chieh Chao |
ACM Multimedia | 1 |
| 2014 | Color CENTRIST: Embedding color information in scene categorization
Wei-Ta Chu, Chih-Hao Chen, Han-Nung Hsu |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | ACM multimedia 2013 workshop on crowdsourcing for multimediaabstractThe topic "Crowdsourcing for Multimedia" encompasses the full range of techniques that combine human intelligence and a large number of individual contributors to advance the state of the art in multimedia research. The ACM Multimedia 2013 Workshop on Crowdsourcing for Multimedia (CrowdMM 2013) provided a forum for presenting new crowdsourcing techniques, exchanging innovative crowdsourcing ideas, and discussing crowdsourcing best practices for multimedia. The workshop program consisted of presented papers, a keynote speech and a panel discussion. A special feature of this year's workshop was the "Crowdsourcing for Multimedia Ideas Competition", the results of which were presented at the workshop. Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson |
ACM Multimedia | 2 |
| 2013 | Size does matter: how image size affects aesthetic perception?abstractThere is no doubt that an image's content determines how people assess the image aesthetically. Previous works have shown that image contrast, saliency features, and the composition of objects may jointly determine whether or not an image is perceived as aesthetically pleasing. In addition to an image's content, the way the image is presented may affect how much viewers appreciate it. For example, it may be assumed that a picture will always look better when it is displayed in a larger size. Is this "the-bigger-the-better" rule always valid? If not, in what situations is it invalid? Wei-Ta Chu, Yu-Kuang Chen, Kuan-Ta Chen |
ACM Multimedia | 1 |
| 2013 | Evaluation of Product Quantization for Image Search
Wei-Ta Chu, Chun-Chang Huang, Jen-Yu Yu |
MMM (1) | 1 |
| 2012 | Logo recognition and localization in real-world images by using visual patternsabstractBy describing spatial relationships between feature points, we present promising logo recognition and localization, which are verified based on two state-of-the-art datasets. Given features points on the query logo, similar features on test images are efficiently found by locality sensitive hashing. After filtering out outliers, candidate regions are found by the mean-sift algorithm, and each region is compared with the logo by jointly considering visual word histogram and visual patterns. Evaluation results show that visual patterns more appropriately describe logos and provide better performance than previous approaches. Wei-Ta Chu, Tsung-Che Lin |
ICASSP | 1 |
| 2012 | GPU-accelerated scene categorization under multiscale category-specific visual word strategyabstractWe utilize GPU to accelerate an essential component for computer vision and multimedia information retrieval, i.e. scene categorization. To construct bag of word models, we modify calculation of Euclidean distance so that feature clustering and visual word quantization can be processed in a parallel manner. We provide details of GPU implementations and conduct comprehensive experiments to verify the efficiency of GPU on multimedia analysis. Wei-Ta Chu, Sheng-Chun Tseng |
ICASSP | 1 |
| 2012 | Color CENTRIST: a color descriptor for scene categorizationabstractWe design a method to incorporate color information into the framework of CENsus Transform histogram (CENTRIST), a state-of-the-art visual descriptor for scene categorization. The newly proposed color CENTRIST descriptor describes global shape information by not only gradient derived from intensity values but also color variations between pixels in local image patches. Through extensive evaluations on various datasets, we demonstrate that the color CENTRIST descriptor is not only easily to be implemented, but also reliably achieves performance over that of CENTRIST. Wei-Ta Chu, Chih-Hao Chen |
ICMR | 1 |
| 2012 | Visual pattern discovery for architecture image classification and product image searchabstractMany objects have repetitive elements, and finding repetitive patterns facilitates object recognition and numerous applications. We devise a representation to describe configurations of repetitive elements. By modeling spatial configurations, visual patterns are more discriminative than local features, and are able to tackle with object scaling, rotation, and deformation. We transfer the pattern discovery problem into finding frequent subgraphs from a graph, and exploit a graph mining algorithm to solve this problem. Visual patterns are then exploited in architecture image classification and product image retrieval, based on the idea that visual pattern can describe elements conveying architecture styles and emblematic motifs of brands. Experimental results show that our pattern discovery approach has promising performance and is superior to the conventional bag-of-words approach. Wei-Ta Chu, Ming-Hung Tsai |
ICMR | 1 |
| 2012 | ACM multimedia 2012 workshop on crowdsourcing for multimediaabstractCrowdsourcing for multimedia involves exploiting both human intelligence and the combination of a large number of individual human contributions (i.e., the 'wisdom of the crowd') to develop techniques, systems and data sets that advance the state of the art. The ACM Multimedia 2012 Workshop on Crowdsourcing for Multimedia (CrowdMM 2012) provides a forum presenting crowdsourcing techniques for multimedia, as well as innovative ideas exemplifying how multimedia research can benefit from crowdsourcing. Through presented papers, invited talks and a panel, the workshop will promote interactive discussion on the scope and research potentials of crowdsourcing. The goal is to provide information to the multimedia research community on the principles of crowdsourcing and to inspire researchers to address the limitations of current studies by innovative use of human computation and collective intelligence. The workshop views crowdsourcing in the broad sense: it encompasses both unsolicited human contributions, e.g., tags assigned by users to images, and also solicited contributions, e.g., annotations gathered by making use of crowdsourcing platforms that micro-outsource tasks to a large pool of human workers. Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2012 | Somebody helps me: Travel video scene detection using web-based context
Wei-Ta Chu, Cheng-Jung Li |
Neurocomputing | 1 |
| 2012 | Rhythm of Motion Extraction and Rhythm-Based Cross-Media Alignment for Dance VideosabstractWe present how to extract rhythm information in dance videos and music, and accordingly correlate them based on rhythmic representation. From dancer's movement, we construct motion trajectories, detect turnings, and stops of trajectories, and then estimate rhythm of motion (ROM). For music, beats are detected to describe rhythm of music. Two modalities are therefore represented as sequences of rhythm information to facilitate finding cross-media correspondence. Two applications, i.e., background music replacement and music video generation, are developed to demonstrate the practicality of cross-media correspondence. We evaluate performance of ROM extraction, and conduct subjective/objective evaluation to show that rich browsing experience can be provided by the proposed applications. Wei-Ta Chu, Shang-Yin Tsai |
IEEE Trans. Multim. | 1 |
| 2011 | Score Following and Retrieval Based on Chroma and Octave Representation
Wei-Ta Chu, Meng-Luen Li |
MMM (1) | 1 |
| 2011 | Travelmedia: An intelligent management system for media captured in travel
Wei-Ta Chu, Cheng-Jung Li, Sheng-Chun Tseng |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | Editing by Viewing: Automatic Home Video Summarization by Viewing Behavior AnalysisabstractIn this paper, we propose the Interest Meter (IM), a system making the computer conscious of user's reactions to measure user's interest and thus use it to conduct video summarization. The IM takes account of users' spontaneous reactions when they view videos. To estimate user's viewing interest, quantitative interest measures are devised based on the perspectives of attention and emotion. For estimating attention states, variations of user's eye movement, blink, and head motion are considered. For estimating emotion states, facial expression is recognized as positive or neural emotion. By combining characteristics of attention and emotion by a fuzzy fusion scheme, we transform users' viewing behaviors into quantitative interest scores, determine interesting parts of videos, and finally concatenate them as video summaries. Experimental results show that the proposed concept “editing by viewing” works well and may provide a promising direction to consider the human factor in video summarization. Wei-Ting Peng, Wei-Ta Chu, Chia-Han Chang, Chien-Nan Chou, Wei-Jia Huang, Wen-Yan Chang, Yi-Ping Hung |
IEEE Trans. Multim. | 2 |
| 2010 | A real-time user Interest Meter and its applications in home video summarizingabstractIn this paper, we propose the Interest Meter (IM), a system making computer conscious of user's reactions, to measure user's interest in real time. The Interest Meter takes account of users' spontaneous reactions when users interact with computers. In this work, we analyze variations of user's eye movement, blink, head motion, and facial expression. Furthermore, we propose an algorithm to combine those signals into interest score and determine important parts of video shots when people watch raw home videos. Experimental result shows that this new type of editing mechanism can effectively generate home video summaries. Wei-Ting Peng, Chia-Han Chang, Wei-Ta Chu, Wei-Jia Huang, Chien-Nan Chou, Wen-Yan Chang, Yi-Ping Hung |
ICME | 3 |
| 2010 | Age classification for pose variant and occluded facesabstractWe extend the object class invariant (OCI) model to age classification, for pose variant and occluded faces. With the OCI model, we first localize faces from images captured in arbitrary views, and then determine the most distinctive features. Relationships between feature points and the invariant vector are described in terms of geometry and appearance information, in the form of a probabilistic model. In contrast to previous works on age classification/estimation, we emphasize that this method is especially useful for faces captured in real-world situations. Wei-Ta Chu, Jen-Yu Yu |
ACM Multimedia | 1 |
| 2010 | Travel Photo and Video Summarization with Cross-Media Correlation and Mutual Influence
Wei-Ta Chu, Che-Cheng Lin, Jen-Yu Yu |
MMM | 1 |
| 2010 | Consumer photo management and browsing facilitated by near-duplicate detection with feature filtering
Wei-Ta Chu |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Modeling spatiotemporal relationships between moving objects for event tactics analysis in tennis videos
Wei-Ta Chu, Wen-Ho Tsai |
Multim. Tools Appl. | 1 |
| 2009 | Using context information and local feature points in face clustering for consumer photosabstractWe introduce local feature points to achieve face clustering for consumer photos. After combining eigenfaces with context information like clothes, we further investigate the usage of local feature points to match face images. The relationships between face images are constructed by feature matching and then described as a graph. Outliers in the results of preliminary clustering are detected and are re-clustered according to matching characteristics. We report complete performance comparison for different datasets and show that the proposed method has superior performance than conventional approaches. Wei-Ta Chu, Ya-Lin Lee, Jen-Yu Yu |
ICASSP | 1 |
| 2009 | Automatic summarization of travel photos using near-duplication detection and feature filteringabstractWe try to address part of the challenge proposed by CeWe. A photo summarization method is developed to select representative photos. Based on the observation that the most important objects/views are often captured several times, we exploit near-duplicate detection techniques to represent a sequence of photo as a graph, and then the graph structure is analyzed to facilitate importance ranking. We focus on summarizing hundreds of photos taken in journeys lasting for one or two weeks. The qualitative and quantitative measurement results demonstrate the effectiveness of the proposed method. Wei-Ta Chu |
ACM Multimedia | 1 |
| 2009 | Feature classification for representative photo selectionabstractThis paper points out that different local feature points provide different impacts to near-duplicate detection and related applications. Aiming to automatic representative photo selection, we develop three feature classification methods, i.e., point-based, region-based, and pLSA-based classification, to differentiate local feature points described by SIFT descriptors. We investigate the performance of these classification methods, and discuss how they influence near-duplicate detection and extended applications. Experiments show that, with effective feature classification, more accurate representative selection results can be achieved. Wei-Ta Chu, Jen-Yu Yu |
ACM Multimedia | 1 |
| 2009 | Visual language model for face clustering in consumer photosabstractFor consumer photos, this work clusters faces with large variations in lighting, pose, and expression. After matching face images by local feature points, we transform matching situations into a novel representation called visual sentences. Then, visual language models are constructed to describe the dependency of image patches on faces. With the probabilistic framework, we develop a clustering algorithm to group the same individual's face images into the same cluster. An interesting observation about evaluating face clustering performance is proposed, and we demonstrate the superiority of the proposed visual language model approach. Wei-Ta Chu, Ya-Lin Lee, Jen-Yu Yu |
ACM Multimedia | 1 |
| 2009 | A User Experience Model for Home Video Summarization
Wei-Ting Peng, Wei-Jia Huang, Wei-Ta Chu, Chien-Nan Chou, Wen-Yan Chang, Chia-Han Chang, Yi-Ping Hung |
MMM | 3 |
| 2009 | RoleNet: Movie Analysis from the Perspective of Social NetworksabstractWith the idea of social network analysis, we propose a novel way to analyze movie videos from the perspective of social relationships rather than audiovisual features. To appropriately describe role's relationships in movies, we devise a method to quantify relations and construct role's social networks, called RoleNet. Based on RoleNet, we are able to perform semantic analysis that goes beyond conventional feature-based approaches. In this work, social relations between roles are used to be the context information of video scenes, and leading roles and the corresponding communities can be automatically determined. The results of community identification provide new alternatives in media management and browsing. Moreover, by describing video scenes with role's context, social-relation-based story segmentation method is developed to pave a new way for this widely-studied topic. Experimental results show the effectiveness of leading role determination and community identification. We also demonstrate that the social-based story segmentation approach works much better than the conventional tempo-based method. Finally, we give extensive discussions and state that the proposed ideas provide insights into context-based video analysis. Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu |
IEEE Trans. Multim. | 2 |
| 2008 | Event detection in tennis matches based on video data miningabstractThis paper proposes a mining-based method to achieve event detection for broadcasting tennis videos. Utilizing visual and aural information, we extract some high-level features to describe video segments. The audiovisual features are further transformed to symbolic streams and an efficient mining technique is applied to derive all frequent patterns that characterize tennis events. After mining, we categorize frequent patterns into several kinds of events and therefore achieve event detection for tennis videos by checking the correspondence between mined patterns and events. The experimental results show that the proposed approach is a promising way to detect events in broadcasting tennis video. Min-Chun Hu 0001, Yi-Tang Wang, Chen-Wei Chou, Kuei-Yi Hsieh, Wei-Ta Chu, Ja-Ling Wu |
ICME | 5 |
| 2008 | Automatic selection of representative photo and smart thumbnailing using near-duplicate detectionabstractThis paper presents two applications about representative photo selection and smart thumbnailing using the results of near-duplicate detection. For a given photo cluster, near-duplicate photo pairs are first determined, and the relationships between them are modeled by a graph. The most typical one is then automatically selected by examining the mutual relation between them. For smart thumbnailing, we determine the region-of-interest of the selected representative photo based on locally matched feature points, which is a view different from conventional saliency-based approaches. The experiments show satisfactory performance in representative selection and promising results in ROI determination. Wei-Ta Chu |
ACM Multimedia | 1 |
| 2008 | Aesthetics-Based Automatic Home Video Skimming System
Wei-Ting Peng, Yueh-Hsuan Chiang, Wei-Ta Chu, Wei-Jia Huang, Wei-Lun Chang, Po-Chung Huang, Yi-Ping Hung |
MMM | 3 |
| 2008 | Explicit semantic events detection and development of realistic applications for broadcasting baseball videos
Wei-Ta Chu, Ja-Ling Wu |
Multim. Tools Appl. | 1 |
| 2007 | Exploring Broadcasting Baseball Videos Based on Multimodal and Multidisciplinary StudyabstractThis demonstration presents a comprehensive work covering semantic event detection, summary/highlight generation, query answering, and efficient browsing, for broadcasting baseball videos. This system integrates the techniques of content-based analysis, key-phrase spotting, natural language processing, data mining, and human-computer interface. It presents a multimodal and multidisciplinary work to show an instance for a media-rich life. Wei-Ta Chu, Ja-Ling Wu |
ICME | 1 |
| 2007 | Movie Analysis Based on Roles' Social NetworkabstractRoles in a movie form a small society and their interrelationship provides clues for movie understanding. Based on this observation, we present a new viewpoint to perform semantic movie analysis. Through checking the co-occurrence of roles in different scenes, we construct a roles' social network to describe their relationships. We introduce the concept of social network analysis to elaborately identify leading roles and the hidden communities. With the results of community identification, we perform storyline detection that facilitates more flexible movie browsing and higher-level movie analysis. The experimental results show that the proposed community identification method is accurate and is robust to errors. Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu |
ICME | 2 |
| 2006 | Extraction of Baseball Trajectory and Physics-Based Validation for Single-View Baseball Video SequencesabstractTo enrich the viewing experience of baseball games and provide some clues for enhancing pitcher's performance, we propose a Kalman filter-based approach to track ball trajectory from single-view pitching sequences. Without setting extraordinary equipments in stadiums or other sensing instruments, this approach robustly extracts ball trajectory for pitching sequences captured from TV channels or downloaded from the Internet. To validate the detected ball trajectories, we investigate the characteristics of ball trajectories on the basis of a baseball physical model. The effectiveness of ball trajectory extraction and ball position detection are presented Wei-Ta Chu, Chia-Wei Wang, Ja-Ling Wu |
ICME | 1 |
| 2006 | Tiling slideshowabstractThis paper presents a new medium, called tiling slideshow, to display photos in a tile-like manner, coordinating with the pace of background music. In contrast to the conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Motivated by the concepts of technical writing, each displaying layout is composed of a larger topic photo and several small-size supportive photos. Based on this idea, the proposed tiling slideshow system consists of three major components: image clustering, music analyzer, and layout organizer. Given the limited displaying space, we consider the context and relationship between photos and model the layout organization as a constrainted optimization problem. Experiments on real consumer photograph collections show that the novel displaying method gives users more pleasant browsing experience than the methods that focus only on single photograph display. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 2 |
| 2006 | Audiovisual slideshow: present your journey by photosabstractThis demonstration presents a novel way to systematically display photos and enhance the viewing experience of photo browsing. In contrast to conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Moreover, the displaying pace is coordinated with the beat of the user-selected incidental music. To automatically generate the audiovisual slideshow, we develop a system that consists of three main components: photo analysis, music analysis, and audiovisual composition. Audiovisual content analysis and cross-media synchronization issues are addressed in this work. This novel demonstration is especially suitable to present photos taken in a journey. It vigorously presents the delights of traveling and helps us recall or experience the trip. Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu |
ACM Multimedia | 2 |
| 2006 | Development of realistic applications based on explicit event detection in broadcasting baseball videosabstractThis paper presents a framework that explicitly detects events in broadcasting baseball videos and facilitates the development of various extended applications. Three phases are included: reliable shot classification, explicit event detection, and elaborate applications. In the shot classification stage, color and geometric information are utilized to classify shots into several canonical views. To explicitly detect semantic events, rule-based decision and model-based decision methods are developed. We emphasize that this system efficiently and exactly identifies what happened in baseball games rather than roughly finding some interesting parts. Based on explicit event detection, many accurate and practical applications such as box score generation and game summarization can be built. The evaluation results show the effectiveness of the proposed methods and demonstrate some insights about bridging semantic gaps for sports videos. Wei-Ta Chu, Ja-Ling Wu |
MMM | 1 |
| 2005 | Integration of rule-based and model-based decision methods for baseball event detectionabstractTo exactly detect what events occur in baseball games, a framework that integrates rule-based and model-based decision methods is proposed. The rule-based decision module infers what happened by checking the information changes in the caption. The model-based decision module further classifies the events that could not be explicitly determined by checking caption information only. Thirteen events, including hit, double, home run, and so on, are considered in this work. The promising experimental results show the effectiveness of the proposed framework and facilitate the development of advanced video applications. Wei-Ta Chu, Ja-Ling Wu |
ICME | 1 |
| 2005 | Generative and Discriminative Modeling toward Semantic Context Detection in Audio TracksabstractSemantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. We propose a hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. For action movies we focused in this work, hidden Markov models (HMMs) are used to model four representative audio events, i.e. gunshot, explosion, car-braking, and engine sounds. At the semantic context level, generative (ergodic hidden Markov model) and discriminative (support vector machine, SVM) approaches are investigated to fuse the characteristics and correlations among various audio events, which provide cues for detecting gunplay and car-chasing scenes. The experimental results demonstrate the effectiveness of the proposed approaches and draw a sketch for semantic indexing and retrieval. Moreover, the differences between two fusion schemes are discussed to be the reference for future research. Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu |
MMM | 1 |
| 2005 | Toward better retrieval and presentation by exploring cross-media correlations
Wei-Ta Chu, Herng-Yow Chen |
Multim. Syst. | 1 |
| 2005 | Toward semantic indexing and retrieval using hierarchical audio models
Wei-Ta Chu, Wen-Huang Cheng, Yung-Jen Hsu 0001, Ja-Ling Wu |
Multim. Syst. | 1 |
| 2004 | A study of semantic context detection by using SVM and GMM approachesabstractSemantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. In this paper, we propose an hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. HMM are used to model audio events, and SVM and GMM are used to fuse the characteristics of various audio events related to some specific semantic concepts. The experimental results show that the approach is effective in detecting semantic context. The comparison between SVM- and GMM-based approaches is also studied Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu, Yung-Jen Hsu 0001 |
ICME | 1 |
| 2002 | Cross-media correlation: a case study of navigated hypermedia documentsabstractThe research issues on multiple media correlation have arisen with more and more integrated multimedia applications. The multimedia correlation is used to coordinate different media and facilitate cross-media access. This paper presents our work on two types of multimedia correlation: explicit and implicit relations. We develop a system to carefully capture explicit relations and devise some computed synchronization processes to discover implicit relations between media objects. The proposed computed synchronization techniques, including speech-text alignment process in temporal domain, automatic scrolling process in spatial domain, and content dependency check process in content domain, will be addressed. Experimental results show that in the speech-text alignment process 80% of forced alignment are in-sync even the speech recognition accuracy is as low as 25%. The automatic scrolling process effectively maintains a resynchronization mechanism in different displaying environments. Wei-Ta Chu, Herng-Yow Chen |
ACM Multimedia | 1 |
| 2002 | The WSML system: web-based synchronization multimedia lecture systemabstractThis demonstration presents a web-based multimedia lecture system that perfectly integrates multimedia lecturing with static HTML pages and enables their presentation in synchronization. We develop key techniques to reproduce vivid web-teaching scenarios for on-demand access, including events capturing scheme and synchronized mechanism. To facilitate multimedia lectures creation, the easy-to-use authoring and managing tools specially designed for teachers have been also included in the development. In addition, integrated presentation and efficient access issues are also discussed for the student's aspect. Kuo-Yu Liu, Natalius Huang, Bo-Hung Wu, Wei-Ta Chu, Herng-Yow Chen |
ACM Multimedia | 4 |