Wei-Ta Chu

dblp:57/5913 · DBLP profile ↗
← Back
101ranked-venue papers
60as first author
30since 2021 · last 2026
0000-0001-5722-7239ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 90 · 56 first-author · 23 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion
Jui-Feng Chi, Wei-Ta Chu, Sheng-Long Lin
MMM (1)2
2026 Human-Object Interaction detection enhanced by common sense and contextual cues
Cheng-Kang Tan, Wei-Ta Chu
Pattern Recognit.2
2025 Multimodal Fusion for Dementia Detection using Voice and Facial Features
abstract
Early detection of dementia is important in the aging society. Research has shown that early diagnosis and treatment can effectively slow down cognitive decline in the elderly. Our goal is to develop a low-cost dementia detection method by analyzing videos capturing the progress of potential patients taking the Short Portable Mental Status Questionnaire (SPMSQ). We propose a multimodal fusion method that effectively predicts the degree of dementia based on both voice and facial features. Because of lacking enough training data, we propose a feature augmentation method based on the mix-up technique to train a more effective model. Experimental results demonstrate the effectiveness of the proposed method.
Jia-Yi Chen, Wei-Ta Chu
ISCAS2
2025 MapLlama: A Two-Stage Approach for Map Question Answering Using a Fine-tuned Large Language Model
abstract
A map is a common medium to convey rich information. People present data on maps and answer spatial-related questions, usually requiring the ability to analyze the map and work on reasoning. In this work, we propose a two-stage map question-answering (MapQA) system. First, map images are converted into structured data based on a query-based model. A fine-tuned large language model (LLM) is employed with the structured text data to answer the given questions. We point out several issues in the current MapQA, propose a realistic scenario, and tackle it with a fine-tuned LLM. The experimental results show that the proposed method works well for diverse questions and performs better than the state-of-the-art techniques.
Zong-Lin Li, Wei-Ta Chu
ISCAS2
2025 SuPACape: Graph-based Category-Agnostic Pose Estimation with Super-Category and Pose Adaptivity
abstract
Given a support image with keypoint annotations, a category-agnostic pose estimation (CAPE) method aims to predict keypoints in a query image from the same category. The relationship between keypoint features of the support image and the visual features of the query image is discovered to achieve CAPE. Although many CAPE methods have been proposed, two challenges remain: 1) Support information is limited, especially when only a few shots of support images are given; 2) the pose in the support image may be significantly different from that in the query image. For the first issue, we propose to consider keypoint features from the categories in the same super-category as the support image to enhance the support representations. For the second issue, we propose a query-adaptive adjacency matrix generated based on the similarity between features of support keypoints and the query image. Experimental results verify that both ideas bring performance gain and make the proposed method achieve state-of-the-art results on the MP-100 benchmark.
Yi-Hsuan Lu, Wei-Ta Chu
MMAsia2
2025 HCV: Lightweight Hybrid CNN-Vision Transformer for Visual Object Tracking
Liang-Chia Chen, Wei-Ta Chu
MMM (2)2
2025 ALSA-UAD: Unsupervised anomaly detection on histopathology images using adversarial learning and simulated anomaly
Yu-Chen Lai, Wei-Ta Chu
J. Vis. Commun. Image Represent.2
2024 Multiple Player Tracking With 3D Projection and Spatio-Temporal Information In Multi-View Sports Videos
abstract
Player tracking is a fundamental task in sports video understanding. Many technical challenges should be addressed due to irregular movement, occlusion between players, and complex background. In this work, we present a framework that utilizes synchronized videos captured from multiple view-points. We construct 2D player trajectories from each video, construct 3D player trajectories based on multiple videos, and then associate 2D and 3D trajectories to achieve multiple player tracking. Experimental results show that the proposed method achieves SOTA performance on a volleyball dataset and a basketball dataset.
Yi-Peng Wang, Wei-Ta Chu
ICASSP2
2024 Transformer-Based Clipped Contrastive Quantization Learning For Unsupervised Image Retrieval
abstract
Unsupervised image retrieval aims to learn the important visual characteristics without any given level to retrieve the similar images for a given query image. The Convolutional Neural Network (CNN)-based approaches have been extensively exploited with self-supervised contrastive learning for image hashing. However, the existing approaches suffer due to lack of effective utilization of global features by CNNs and biased-ness created by false negative pairs in the contrastive learning. In this paper, we propose a TransClippedCLR model by encoding the global context of an image using Transformer having local context through patch based processing, by generating the hash codes through product quantization and by avoiding the potential false negative pairs through clipped contrastive learning. The proposed model is tested with superior performance for unsupervised image retrieval on benchmark datasets, including CIFAR10, NUS-Wide and Flickr25K, as compared to the recent state-of-the-art deep models. The results using the proposed clipped contrastive learning are greatly improved on all datasets as compared to same backbone network with vanilla contrastive learning.
Ayush Dubey, Shiv Ram Dubey, Satish Kumar Singh, Wei-Ta Chu
ICIP4
2024 CS-HOI: Human Object Interaction Detection Enhanced by Common Sense
Cheng-Kang Tan, Wei-Ta Chu
MMAsia2
2024 Incremental Few-Shot Object Detection by Leveraging External Information from Large Multimodal Models
Guan-Yu Wu, Wei-Ta Chu
MMAsia2
2024 Frequency disentangled residual network
Satya Rajendra Singh, Roshan Reddy Yedla, Shiv Ram Dubey, Rakesh Kumar Sanodiya, Wei-Ta Chu
Multim. Syst.5
2024 Overall positive prototype for few-shot open-set recognition
Liang-Yu Sun, Wei-Ta Chu
Pattern Recognit.2
2024 Positive and Negative Set Designs in Contrastive Feature Learning for Temporal Action Segmentation
abstract
When data labels are scarce, contrastive learning is often used to learn representations in a weakly-supervised or unsupervised way. In contrastive learning, not only the learning mechanism, but also the designs of positive and negative sets are critical. While most previous works of Temporal Action Segmentation (TAS) focus on designing new segmentation methods, we investigate the importance of positive and negative set designs in contrastive learning and verify that better representations can be learned to enhance performance of existing TAS methods. Specific to timestamp-supervised TAS and unsupervised TAS, respectively, we propose positive/negative set designs, associated with the ideas of ambiguous frames and the set expansion process to make learned representations more effective. In the evaluation, we demonstrate that performance of timestamp-supervised TAS can be boosted by 8% to 15% in terms of F1@10 across three different datasets, and the performance of unsupervised TAS can be boosted by 3% to 5% in terms of F1 scores, achieving new state-of-the-art TAS results.
Wei-Ta Chu
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Trajectory-based Statistics and Tactics Analysis System for Table Tennis
abstract
For table tennis videos, we develop a system to analyze and generate statistics based on ball trajectories. By a real-time ball detector, the ball trajectory is constructed based on the tracking by detection scheme. Landing points on the table are estimated. Based on moving direction and the sequence of landing points, three-stage analysis can be achieved. We also analyze how a point starts (serving type classification) and how a point ends (point loss classification).
Guan-Yu Wu, Chun-Ho Hung, Hsuan-Wei Chen, Wei-Ta Chu
MMAsia4
2023 Occlusion­-Aware Manga Character Re­-identification with Self-Paced Contrastive Learning
abstract
Existing methods for manga character re-identification primarily rely on facial information, overlooking the unique characteristics of characters’ bodies and failing to address common challenges like occlusion by speech balloons and incomplete body parts. To tackle these issues, we propose a method called Occlusion-Aware Manga Character Re-identification (OAM-ReID) with self-paced contrastive learning, which leverages annotated body data from the Manga109 dataset for training. By synthesizing data with occluded speech balloons and incomplete bodies, we empower the framework to be aware of occlusion, so that more effective feature representations are learnt. Experimental results show that this approach outperforms the state-of-the-art person ReID method.
Ci-Yin Zhang, Wei-Ta Chu
MMAsia2
2023 The NCKU-VTF Dataset and a Multi-scale Thermal-to-Visible Face Synthesis System
Tsung-Han Ho, Chen-Yin Yu, Tsai-Yen Ko, Wei-Ta Chu
MMM (1)4
2023 Manga Text Detection with Manga-Specific Data Augmentation and Its Applications on Emotion Analysis
Yi-Ting Yang, Wei-Ta Chu
MMM (2)2
2023 SSSD: Self-Supervised Self Distillation
abstract
With labeled data, self distillation (SD) has been proposed to develop compact but effective models without a complex teacher model available in advance. Such approaches need labeled data to guide the self distillation process. Inspired by self-supervised (SS) learning, we propose a self-supervised self distillation (SSSD) approach in this work. Based on an unlabeled image dataset, a model is constructed to learn visual representations in a self-supervised manner. This pre-trained model is then adopted to extract visual representations of the target dataset and generates pseudo labels via clustering. The pseudo labels guide the SD process, and thus enable SD to proceed in an unsupervised way (no data labels are required at all). We verify this idea based on evaluations on the CIFAR-10, CIFAR-100, and ImageNet-1K datasets, and demonstrate the effectiveness of this unsupervised SD approach. Performance outperforming similar frameworks is also shown.
Wei-Chi Chen, Wei-Ta Chu
WACV2
2022 Vision Transformer Hashing for Image Retrieval
abstract
Recently, Transformer has emerged as a new architecture in deep learning by utilizing self-attention without convolution. Transformer is also extended to Vision Transformer (ViT) for the visual recognition with a promising performance on ImageNet. In this paper, we propose a Vision Transformer Hashing (VTS) for image retrieval. We utilize the pre-trained ViT on ImageNet as the backbone network and add the hashing head. The proposed VTS model is fine tuned for hashing under six different image retrieval frameworks with their objective functions. We perform the extensive experiments on CIFAR10, ImageNet, NUS-Wide, and COCO datasets. The proposed VTS based image retrieval outperforms the recent state-of-the-art hashing techniques with a significant margin. We also find the proposed VTS model as the backbone network is better than the existing networks, such as AlexNet and ResNet. The code is released at https://github.com/shivram1987/VisionTransformerHashing.
Shiv Ram Dubey, Satish Kumar Singh, Wei-Ta Chu
ICME3
2022 Multimodal Fusion with Cross-Modal Attention for Action Recognition in Still Images
abstract
We propose a cross-modal attention module to combine information from different cues and different modalities, to achieve action recognition in still images. Feature maps are extracted from the entire image, the detected human bounding box, and the detected human skeleton, respectively. Inspired by the transformer structure, we design the processing between the query vector from one cue/modality, and the key vector from another cue/modality. Feature maps from different cues/modalities are cross-referred so that better representations can be obtained to yield better performance. We show that the proposed framework outperforms the state-of-the-art systems without the requirement of an extra training dataset. We also conduct ablation studies to investigate how different settings impact the final results.
Jia-Hua Tsai, Wei-Ta Chu
MMAsia2
2022 Indie Games Popularity Prediction by Considering Multimodal Features
Yu-Heng Huang, Wei-Ta Chu
MMM (2)2
2022 An imitation learning framework for generating multi-modal trajectories from unstructured demonstrations
Jain-Wei Peng, Min-Chun Hu 0001, Wei-Ta Chu
Neurocomputing3
2022 Instant Basketball Defensive Trajectory Generation
abstract
Tactic learning in virtual reality (VR) has been proven to be effective for basketball training. Endowed with the ability of generating virtual defenders in real time according to the movement of virtual offenders controlled by the user, a VR basketball training system can bring more immersive and realistic experiences for the trainee. In this article, an autoregressive generative model for instantly producing basketball defensive trajectory is introduced. We further focus on the issue of preserving the diversity of the generated trajectories. A differentiable sampling mechanism is adopted to learn the continuous Gaussian distribution of player position. Moreover, several heuristic loss functions based on the domain knowledge of basketball are designed to make the generated trajectories assemble real situations in basketball games. We compare the proposed method with the state-of-the-art works in terms of both objective and subjective manners. The objective manner compares the average position, velocity, and acceleration of the generated defensive trajectories with the real ones to evaluate the fidelity of the results. In addition, more high-level aspects such as the empty space for offender and the defensive pressure of the generated trajectory are also considered in the objective evaluation. As for the subjective manner, visual comparison questionnaires on the proposed and other methods are thoroughly conducted. The experimental results show that the proposed method can achieve better performance than previous basketball defensive trajectory generation works in terms of different evaluation metrics.
Wen-Cheng Chen, Wan-Lun Tsai, Huan-Hua Chang, Min-Chun Hu 0001, Wei-Ta Chu
ACM Trans. Intell. Syst. Technol.5
2021 Multi-Class Novelty Detection with Generated Hard Novel Features
Wei-Ta Chu, Wei-Ting Cao
BMVC1
2021 Searching by Generating: Flexible and Efficient One-Shot NAS With Architecture Generator
abstract
In one-shot NAS, sub-networks need to be searched from the supernet to meet different hardware constraints. However, the search cost is high and N times of searches are needed for N different constraints. In this work, we propose a novel search strategy called architecture generator to search sub-networks by generating them, so that the search process can be much more efficient and flexible. With the trained architecture generator, given target hardware constraints as the input, N good architectures can be generated for N constraints by just one forward pass without re-searching and supernet retraining. Moreover, we propose a novel single-path supernet, called unified supernet, to further improve search efficiency and reduce GPU memory consumption of the architecture generator. With the architecture generator and the unified supernet, we propose a flexible and efficient one-shot NAS framework, called Searching by Generating NAS (SGNAS). With the pre-trained supernt, the search time of SGNAS for N different hardware constraints is only 5 GPU hours, which is 4N times faster than previous SOTA single-path methods. After training from scratch, the top1-accuracy of SGNAS on ImageNet is 77.1%, which is comparable with the SOTAs. The code is available at: https://github.com/eric8607242/SGNAS.
Sian-Yao Huang, Wei-Ta Chu
CVPR2
2021 PONAS: Progressive One-shot Neural Architecture Search for Very Efficient Deployment
abstract
We propose a Progressive One-Shot Neural Architecture Search (PONAS) method to achieve a very efficient model searching for various hardware constraints. Given a constraint, most neural architecture search (NAS) methods either sample a set of sub-networks according to a pre-trained accuracy predictor, or adopt the evolutionary algorithm to evolve specialized networks from the supernet. Both approaches are time consuming. Here our key idea for very efficient deployment is, when searching the architecture space, constructing a table that stores the validation accuracy of all candidate blocks at all layers. For a stricter hardware constraint, the architecture of a specialized network can be efficiently determined based on this table by picking the best candidate blocks that yield the least accuracy loss. To accomplish this idea, we propose the PONAS method to combine advantages of progressive NAS and one-shot methods. A two-stage training scheme, including the meta training stage and the fine-tuning stage, is proposed to make the search process efficient and stable. During search, we evaluate candidate blocks in different layers and construct an accuracy table that is to be used in architecture searching. Comprehensive experiments verify that PONAS is extremely flexible, and is able to find architecture of a specialized network in around 10 seconds. In ImageNet classification, 76.29% top-1 accuracy can be obtained, which is comparable with the state of the arts.
Sian-Yao Huang, Wei-Ta Chu
IJCNN2
2021 Automatic Baseball Pitch Overlay
abstract
To provide rich viewing experience and assist pitcher training, we propose an automatic baseball pitch overlay system in this paper. Given multiple pitching video sequences, this system detects and tracks the ball to construct ball trajectories. Because of occlusion, motion blur, and background noise, the ball usually cannot be detected successfully. We propose a series of processes like initial compensation and polynomial fitting to construct complete trajectories. To make the overlay results more appealing, different sequences are weighted differently, and different trajectories are intentionally drawn in different colors. We believe this would be the first fully-automatic pitch overlay system that only takes pitching videos as inputs. Source code is at \\https://github.com/chonyy/ML-auto-baseball-pitching-overlay.
Ting-Hsuan Chou, Wei-Ta Chu
ICMR2
2021 Thermal Face Recognition Based on Multi-scale Image Synthesis
Wei-Ta Chu, Ping-Shen Huang
MMM (1)1
2021 Multi-label image recognition by using semantics consistency, object correlation, and multiple samples
Wei-Ta Chu, Si-Heng Huang
J. Vis. Commun. Image Represent.1
2020 MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020
abstract
The International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173
Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki
ICMR1
2020 An autoregressive generation model for producing instant basketball defensive trajectory
abstract
Learning basketball tactic via virtual reality environment requires real-time feedback to improve the realism and interactivity. For example, the virtual defender should move immediately according to the player's movement. In this paper, we proposed an autoregressive generative model for basketball defensive trajectory generation. To learn the continuous Gaussian distribution of player position, we adopt a differentiable sampling process to sample the candidate location with a standard deviation loss, which can preserve the diversity of the trajectories. Furthermore, we design several additional loss functions based on the domain knowledge of basketball to make the generated trajectories match the real situation in basketball games. The experimental results show that the proposed method can achieve better performance than previous works in terms of different evaluation metrics.
Huan-Hua Chang, Wen-Cheng Chen, Wan-Lun Tsai, Min-Chun Hu 0001, Wei-Ta Chu
MMAsia5
2020 Thermal Face Recognition Based on Transformation by Residual U-Net and Pixel Shuffle Upsampling
Soumya Chatterjee 0002, Wei-Ta Chu
MMM (1)2
2019 A Genetic Programming Approach to Integrate Multilayer CNN Features for Image Classification
Wei-Ta Chu, Hao-An Chu
MMM (1)1
2019 Photo Filter Classification and Filter Recommendation without Much Manual Labeling
abstract
Because how users employ filters to photos may reveal user's preference or mental state, a photo filter classification method is potentially demanded to enable future large-scale analysis. We adopt the transfer learning technique to transform deep models pre-trained for object classification into models suitable for photo filter classification. Based on accurate classification results, we build a filter recommendation approach without much manual labeling. It can be easily extended when more training data are available. A series of experimental studies are conducted to demonstrate effectiveness of filter classification with transfer learning. We also demonstrate the proposed filter recommendation achieves encouraging performance.
Wei-Ta Chu, Yu-Tzu Fan
MMSP1
2019 Thermal Facial Landmark Detection by Deep Multi-Task Learning
abstract
We present a neural network to jointly consider facial landmark detection and emotion recognition for thermal face images. The first part of this network is based on the U-Net structure, targeting at extracting good features for advanced analysis. Using U - Net as the basic structure enables modeling context information based on a limited number of training data. The second part of this network contains two branches that are designed for landmark detection and emotion recognition, respectively. We propose a two-stage training mechanism to learn this network, and demonstrate the effectiveness of the proposed approach. This work is believed to be one of the few studies on thermal face image analysis.
Wei-Ta Chu, Yu-Hui Liu
MMSP1
2019 Spatiotemporal Modeling and Label Distribution Learning for Video Summarization
abstract
For a video which content does not follow specific production rules, or without professional editing, at least two problems should be solved to generate a good video summary. First, the summarization system should jointly model visual content in the spatial domain and visual dynamics in the temporal domain. Second, the system should consider the inconsistency between users, i.e., different users may annotate the same video segment with different importance scores. In this paper, we present a video summarization system that models spatiotemporal information of video segments, and predicts the distribution of importance scores for each segment. Based on the estimated importance scores, video summaries are generated by picking the ones with higher scores. We especially demonstrate the effectiveness of label distribution learning based on two video benchmarks.
Wei-Ta Chu, Yu-Hsin Liu
MMSP1
2019 Manga face detection based on deep neural networks fusing global and local information
Wei-Ta Chu
Pattern Recognit.1
2018 A Parametric Study of Deep Perceptual Model on Visible to Thermal Face Recognition
abstract
Recently deep perceptual mapping (DPM) based on auto-encoder provides the state-the-art thermal to visible face recognition. Features extracted from patches of a long-wave infra-red (LWIR) face image are transformed into a space by an auto-encoder, such that features from infra-red images are comparable with features from visible images. In this paper, we comprehensively evaluate DPM with different settings, in order to build a reference study for future research.
Wei-Ta Chu, Jo-Ning Wu
VCIP1
2018 Text Detection in Manga by Deep Region Proposal, Classification, and Regression
abstract
Text in manga presents high variations and different contextual information, and existing scene text detection methods are not directly applicable. We propose two approaches based on deep networks to detect text in manga. In the first approach, features extracted from multiple CNNs are joined and then fed to a combination of a classification network and a regression network. In the second approach, region proposal, feature extraction, and classification/regression, are taken together in a single deep network. The evaluation results show that the first approach achieves performance comparable to the current state of the art, while the second approach yields a big performance leap over existing ones.
Wei-Ta Chu, Chih-Chi Yu
VCIP1
2018 Visual Weather Temperature Prediction
abstract
In this paper, we attempt to employ convolutional recurrent neural networks for weather temperature estimation using only image data. We study ambient temperature estimation based on deep neural networks in two scenarios a) estimating temperature of a single outdoor image, and b) predicting temperature of the last image in an image sequence. In the first scenario, visual features are extracted by a convolutional neural network trained on a large-scale image dataset. We demonstrate that promising performance can be obtained, and analyze how volume of training data influences performance. In the second scenario, we consider the temporal evolution of visual appearance, and construct a recurrent neural network to predict the temperature of the last image in a given image sequence. We obtain better prediction accuracy compared to the state-of-the-art models. Further, we investigate how performance varies when information is extracted from different scene regions, and when images are captured in different daytime hours. Our approach further reinforces the idea of using only visual information for cost efficient weather prediction in the future.
Wei-Ta Chu, Kai-Chia Ho, Ali Borji
WACV1
2018 Image Style Classification Based on Learnt Deep Correlation Features
abstract
This paper presents a comprehensive study of deep correlation features on image style classification. Inspired by that, correlation between feature maps can effectively describe image texture, and we design various correlations and transform them into style vectors, and investigate classification performance brought by different variants. In addition to intralayer correlation, interlayer correlation is proposed as well, and its effectiveness is verified. After showing the effectiveness of deep correlation features, we further propose a learning framework to automatically learn correlations between feature maps. Through extensive experiments on image style classification and artist classification, we demonstrate that the proposed learnt deep correlation features outperform several variants of convolutional neural network features by a large margin, and achieve the state-of-the-art performance.
Wei-Ta Chu, Yi-Ling Wu
IEEE Trans. Multim.1
2017 Blog Article Summarization with Image-Text Alignment Techniques
abstract
We propose an image-text alignment framework to match images with text, and take blog article summarization as the main application. Objects in an image are first detected, from them deep features are extracted and transformed into a space commonly shared with the text. On the other hand, sentences of a blog article are represented as vectors, and are also embedded into the common space. With these processes, cross-modal matching can be achieved. A blog article is then summarized in the representation of images and their matched sentences. In evaluation, we demonstrate the effectiveness of the proposed method, and show that the generated summary makes more sense.
Wei-Ta Chu, Ming-Chih Kao
ISM1
2017 Manga FaceNet: Face Detection in Manga based on Deep Neural Network
abstract
Among various elements of manga, character's face plays one of the most important role in access and retrieval. We propose a DNN-based method to do manga face detection, which is a challenging but relatively unexplored topic. Given a manga page, we first find candidate regions based on the selective search scheme. A deep neural network is then proposed to detect manga faces of various appearance. We evaluate the proposed method based on a large-scale benchmark, and show performance comparison and convincing evaluation results that have rarely done before.
Wei-Ta Chu
ICMR1
2017 Badminton Video Analysis based on Spatiotemporal and Stroke Features
abstract
Most of the broadcasted sports events nowadays present game statistics to the viewers which can be used to design the gameplay strategy, improve player's performance, or improve accessing the point of interest of a sport game. However, few studies have been proposed for broadcasted badminton videos. In this paper, we integrate several visual analysis techniques to detect the court, detect players, classify strokes, and classify the player's strategy. Based on visual analysis, we can get some insights about the common strategy of a certain player. We evaluate performance of stroke classification, strategy classification, and show game statistics based on classification results.
Wei-Ta Chu, Samuel I. G. Situmeang
ICMR1
2017 Camera as weather sensor: Estimating weather information from single images
Wei-Ta Chu, Xiang-You Zheng, Ding-Shiuan Ding
J. Vis. Commun. Image Represent.1
2017 On broadcasted game video analysis: event detection, highlight detection, and highlight forecast
Wei-Ta Chu, Yung-Chieh Chou
Multim. Tools Appl.1
2017 Cultural difference and visual information on hotel rating prediction
Wei-Ta Chu, Wei-Han Huang
World Wide Web1
2017 A hybrid recommendation system considering visual information for predicting favorite restaurants
Wei-Ta Chu, Ya-Lun Tsai
World Wide Web1
2016 Manga-specific features and latent style model for manga style analysis
abstract
A latent style model describing manga styles based on the proposed manga-specific features is constructed to facilitate novel style-based applications. Two manga-specific features, i.e., screentone features showing texture and shade, and panel features showing panel arrangement, are firstly proposed to describe manga pages. Based on the latent Dirichlet allocation technique, we discover latent style elements embedded in manga documents, which are described by visual words derived from manga-specific features. Distributions of style elements are then used to measure similarity between manga documents, and facilitate the development ofvarious style-based applications. Experimental results show that the features and models especially designed for describing manga styles yield promising performance and could bring many potential extensions.
Wei-Ta Chu, Wei-Chung Cheng
ICASSP1
2016 News story clustering with fisher embedding
abstract
An automatic news story clustering system is presented to facilitate efficient news browsing and summarization. We describe news content by considering both what objects appear and how these objects move in news stories. With Fisher embedding, we respectively encode local features, semantics features, and dense trajectories as Fisher vectors, based on which similarity between news stories can be well evaluated and thus better clustering performance can be obtained. We verify the effectiveness of Fisher encoding, and further show that motion-based features are more effective than appearance-based features through feature analysis.
Wei-Ta Chu, Han-Nung Hsu
ICASSP1
2016 Deep Correlation Features for Image Style Classification
abstract
This paper presents a comprehensive study of deep correlation features on image style classification. Inspired by that correlation between feature maps can effectively describe image texture, we design and transform various such correlations into style vectors, and investigate classification performance brought by different variants. In addition to intra-layer correlation, we also propose inter-layer correlation and verify its benefit. Through extensive experiments on image style classification and artist classification, we demonstrate that the proposed style vectors significantly outperforms CNN features coming from fully-connected layers, as well as outperforms the state-of-the-art deep representation.
Wei-Ta Chu, Yi-Ling Wu
ACM Multimedia1
2016 Predicting Occupation from Images by Combining Face and Body Context Information
abstract
Facial images embed age, gender, and other rich information that is implicitly related to occupation. In this work, we advocate that occupation prediction from a single facial image is a doable computer vision problem. We extract multilevel hand-crafted features associated with locality-constrained linear coding and convolutional neural network features as image occupation descriptors. To avoid the curse of dimensionality and overfitting, a boost strategy called multichannel SVM is used to integrate features from face and body. Intra- and interclass visual variations are jointly considered in the boosting framework to further improve performance. In the evaluation, we verify the effectiveness of predicting occupation from face and demonstrate promising performance obtained by combining face and body information. More importantly, our work further integrates deep features into the multichannel SVM framework and shows significantly better performance over the state of the art.
Wei-Ta Chu, Chih-Hao Chiu
ACM Trans. Multim. Comput. Commun. Appl.1
2015 A Privacy-Preserving Bipartite Graph Matching Framework for Multimedia Analysis and Retrieval
abstract
The emergence of cloud computing provides an unlimited computation/storage for users, and yields new opportunities for multimedia analysis and retrieval research. However, privacy of users, e.g., search intention, may be leaked to the server and maliciously utilized by companies or individuals with animus. This paper presents a privacy-preserving multimedia analysis framework based on a widely-adopted structure, i.e., bipartite graph, so that multimedia analysis and retrieval in the encrypted domain is enabled. This work aims to keep the server unaware of what the user wants to retrieve, and at the same time take advantage of the server's computation power. Homomorphic encryption schemes and communication protocols in the encrypted domain are integrated to facilitate bipartite graph construction and implement the Hungarian algorithm to find the best matching. Two applications, video tag suggestion and video copy detection, are developed on top of the privacy-preserving framework, and the evaluation results demonstrate that performance obtained in the encrypted domain is comparable with that obtained in the plain text domain.
Wei-Ta Chu, Feng-Chi Chang
ICMR1
2015 Street sweeper: detecting and removing cars in street view images
Wei-Ta Chu, Ying-Chieh Chao, Yi-Sheng Chang
Multim. Tools Appl.1
2015 Optimized Comics-Based Storytelling for Temporal Image Sequences
abstract
We propose a system to transform any temporal image sequence into a comics-based presentation, as an effective and interesting storytelling manner. Three main components, including page allocation, layout selection, and speech balloon placement, are respectively formulated as optimization problems, and systematic approaches are proposed to find solutions. Page allocation is viewed as a labeling problem, and the best solution is determined by the genetic algorithm. Importance values of images and predefined layouts are both represented in vector forms, and the best layout is selected by finding the best match between vectors. Feasible solutions of speech balloons constitute a solution space, and the best solution that jointly describes the best locations of all balloons in a page is determined by the particle swarm optimization algorithm. Objective evaluation and subjective evaluation are designed from various perspectives to demonstrate effectiveness and superiority of the proposed system.
Wei-Ta Chu, Chia-Hsiang Yu, Hsin-Han Wang
IEEE Trans. Multim.1
2014 Predicting Occupation from Single Facial Images
abstract
Facial images embed age, gender, and other rich information that is implicitly related to occupation. In this work, we advocate that occupation prediction from a single facial image is a doable research direction. We first extract visual features from multiple levels of patches and describe them by locality-constrained linear coding. To avoid the curse of dimensionality and over fitting, a boost strategy called multi-feature SVM is used to integrate features. Intra-class and inter-class visual variations are jointly considered in the boosting framework to further improve performance. In the evaluation, we verify that this is a promising research topic with encouraging performance, and also discuss interesting issues from various perspectives.
Wei-Ta Chu, Chih-Hao Chiu
ISM1
2014 Fast Object Detection Using Multistage Particle Window Deformable Part Model
abstract
For object detection, evaluating all sliding windows at various scales draws a computational efficiency issue. In this paper, we propose a fast object detection framework using the multistage particle window strategy to accelerate the cascade deformable part model (DPM). Coupling this strategy with the proposed early jump scheme, adaptive particle window generation, and efficient preprocessing, we demonstrate that the proposed method runs 34.5 times faster than the conventional DPM to detect objects in images, and is able to efficiently detect vehicles and pedestrians in on-road videos.
Wei-Ta Chu, Ming-Hung Hsu
ISM1
2014 Line-Based Drawing Style Description for Manga Classification
abstract
Diversity of drawing styles of mangas can be easily perceived by humans, but are hard to be described in text. We design computational features derived from line segments to describe drawing styles, enabling style classification such as discriminating mangas targeting youth boys and youth girls, and discriminating artworks produced by different artists. With statistical analysis, we found that drawing styles can be effectively characterized by the proposed features such that explicit (e.g., density of line segments) or implicit (e.g., included angles between lines) observations can be made to facilitate various manga style classification.
Wei-Ta Chu, Ying-Chieh Chao
ACM Multimedia1
2014 Color CENTRIST: Embedding color information in scene categorization
Wei-Ta Chu, Chih-Hao Chen, Han-Nung Hsu
J. Vis. Commun. Image Represent.1
2013 ACM multimedia 2013 workshop on crowdsourcing for multimedia
abstract
The topic "Crowdsourcing for Multimedia" encompasses the full range of techniques that combine human intelligence and a large number of individual contributors to advance the state of the art in multimedia research. The ACM Multimedia 2013 Workshop on Crowdsourcing for Multimedia (CrowdMM 2013) provided a forum for presenting new crowdsourcing techniques, exchanging innovative crowdsourcing ideas, and discussing crowdsourcing best practices for multimedia. The workshop program consisted of presented papers, a keynote speech and a panel discussion. A special feature of this year's workshop was the "Crowdsourcing for Multimedia Ideas Competition", the results of which were presented at the workshop.
Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson
ACM Multimedia2
2013 Size does matter: how image size affects aesthetic perception?
abstract
There is no doubt that an image's content determines how people assess the image aesthetically. Previous works have shown that image contrast, saliency features, and the composition of objects may jointly determine whether or not an image is perceived as aesthetically pleasing. In addition to an image's content, the way the image is presented may affect how much viewers appreciate it. For example, it may be assumed that a picture will always look better when it is displayed in a larger size. Is this "the-bigger-the-better" rule always valid? If not, in what situations is it invalid?
Wei-Ta Chu, Yu-Kuang Chen, Kuan-Ta Chen
ACM Multimedia1
2013 Evaluation of Product Quantization for Image Search
Wei-Ta Chu, Chun-Chang Huang, Jen-Yu Yu
MMM (1)1
2012 Logo recognition and localization in real-world images by using visual patterns
abstract
By describing spatial relationships between feature points, we present promising logo recognition and localization, which are verified based on two state-of-the-art datasets. Given features points on the query logo, similar features on test images are efficiently found by locality sensitive hashing. After filtering out outliers, candidate regions are found by the mean-sift algorithm, and each region is compared with the logo by jointly considering visual word histogram and visual patterns. Evaluation results show that visual patterns more appropriately describe logos and provide better performance than previous approaches.
Wei-Ta Chu, Tsung-Che Lin
ICASSP1
2012 GPU-accelerated scene categorization under multiscale category-specific visual word strategy
abstract
We utilize GPU to accelerate an essential component for computer vision and multimedia information retrieval, i.e. scene categorization. To construct bag of word models, we modify calculation of Euclidean distance so that feature clustering and visual word quantization can be processed in a parallel manner. We provide details of GPU implementations and conduct comprehensive experiments to verify the efficiency of GPU on multimedia analysis.
Wei-Ta Chu, Sheng-Chun Tseng
ICASSP1
2012 Color CENTRIST: a color descriptor for scene categorization
abstract
We design a method to incorporate color information into the framework of CENsus Transform histogram (CENTRIST), a state-of-the-art visual descriptor for scene categorization. The newly proposed color CENTRIST descriptor describes global shape information by not only gradient derived from intensity values but also color variations between pixels in local image patches. Through extensive evaluations on various datasets, we demonstrate that the color CENTRIST descriptor is not only easily to be implemented, but also reliably achieves performance over that of CENTRIST.
Wei-Ta Chu, Chih-Hao Chen
ICMR1
2012 Visual pattern discovery for architecture image classification and product image search
abstract
Many objects have repetitive elements, and finding repetitive patterns facilitates object recognition and numerous applications. We devise a representation to describe configurations of repetitive elements. By modeling spatial configurations, visual patterns are more discriminative than local features, and are able to tackle with object scaling, rotation, and deformation. We transfer the pattern discovery problem into finding frequent subgraphs from a graph, and exploit a graph mining algorithm to solve this problem. Visual patterns are then exploited in architecture image classification and product image retrieval, based on the idea that visual pattern can describe elements conveying architecture styles and emblematic motifs of brands. Experimental results show that our pattern discovery approach has promising performance and is superior to the conventional bag-of-words approach.
Wei-Ta Chu, Ming-Hung Tsai
ICMR1
2012 ACM multimedia 2012 workshop on crowdsourcing for multimedia
abstract
Crowdsourcing for multimedia involves exploiting both human intelligence and the combination of a large number of individual human contributions (i.e., the 'wisdom of the crowd') to develop techniques, systems and data sets that advance the state of the art. The ACM Multimedia 2012 Workshop on Crowdsourcing for Multimedia (CrowdMM 2012) provides a forum presenting crowdsourcing techniques for multimedia, as well as innovative ideas exemplifying how multimedia research can benefit from crowdsourcing. Through presented papers, invited talks and a panel, the workshop will promote interactive discussion on the scope and research potentials of crowdsourcing. The goal is to provide information to the multimedia research community on the principles of crowdsourcing and to inspire researchers to address the limitations of current studies by innovative use of human computation and collective intelligence. The workshop views crowdsourcing in the broad sense: it encompasses both unsolicited human contributions, e.g., tags assigned by users to images, and also solicited contributions, e.g., annotations gathered by making use of crowdsourcing platforms that micro-outsource tasks to a large pool of human workers.
Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson, Wei Tsang Ooi
ACM Multimedia2
2012 Somebody helps me: Travel video scene detection using web-based context
Wei-Ta Chu, Cheng-Jung Li
Neurocomputing1
2012 Rhythm of Motion Extraction and Rhythm-Based Cross-Media Alignment for Dance Videos
abstract
We present how to extract rhythm information in dance videos and music, and accordingly correlate them based on rhythmic representation. From dancer's movement, we construct motion trajectories, detect turnings, and stops of trajectories, and then estimate rhythm of motion (ROM). For music, beats are detected to describe rhythm of music. Two modalities are therefore represented as sequences of rhythm information to facilitate finding cross-media correspondence. Two applications, i.e., background music replacement and music video generation, are developed to demonstrate the practicality of cross-media correspondence. We evaluate performance of ROM extraction, and conduct subjective/objective evaluation to show that rich browsing experience can be provided by the proposed applications.
Wei-Ta Chu, Shang-Yin Tsai
IEEE Trans. Multim.1
2011 Score Following and Retrieval Based on Chroma and Octave Representation
Wei-Ta Chu, Meng-Luen Li
MMM (1)1
2011 Travelmedia: An intelligent management system for media captured in travel
Wei-Ta Chu, Cheng-Jung Li, Sheng-Chun Tseng
J. Vis. Commun. Image Represent.1
2011 Editing by Viewing: Automatic Home Video Summarization by Viewing Behavior Analysis
abstract
In this paper, we propose the Interest Meter (IM), a system making the computer conscious of user's reactions to measure user's interest and thus use it to conduct video summarization. The IM takes account of users' spontaneous reactions when they view videos. To estimate user's viewing interest, quantitative interest measures are devised based on the perspectives of attention and emotion. For estimating attention states, variations of user's eye movement, blink, and head motion are considered. For estimating emotion states, facial expression is recognized as positive or neural emotion. By combining characteristics of attention and emotion by a fuzzy fusion scheme, we transform users' viewing behaviors into quantitative interest scores, determine interesting parts of videos, and finally concatenate them as video summaries. Experimental results show that the proposed concept “editing by viewing” works well and may provide a promising direction to consider the human factor in video summarization.
Wei-Ting Peng, Wei-Ta Chu, Chia-Han Chang, Chien-Nan Chou, Wei-Jia Huang, Wen-Yan Chang, Yi-Ping Hung
IEEE Trans. Multim.2
2010 A real-time user Interest Meter and its applications in home video summarizing
abstract
In this paper, we propose the Interest Meter (IM), a system making computer conscious of user's reactions, to measure user's interest in real time. The Interest Meter takes account of users' spontaneous reactions when users interact with computers. In this work, we analyze variations of user's eye movement, blink, head motion, and facial expression. Furthermore, we propose an algorithm to combine those signals into interest score and determine important parts of video shots when people watch raw home videos. Experimental result shows that this new type of editing mechanism can effectively generate home video summaries.
Wei-Ting Peng, Chia-Han Chang, Wei-Ta Chu, Wei-Jia Huang, Chien-Nan Chou, Wen-Yan Chang, Yi-Ping Hung
ICME3
2010 Age classification for pose variant and occluded faces
abstract
We extend the object class invariant (OCI) model to age classification, for pose variant and occluded faces. With the OCI model, we first localize faces from images captured in arbitrary views, and then determine the most distinctive features. Relationships between feature points and the invariant vector are described in terms of geometry and appearance information, in the form of a probabilistic model. In contrast to previous works on age classification/estimation, we emphasize that this method is especially useful for faces captured in real-world situations.
Wei-Ta Chu, Jen-Yu Yu
ACM Multimedia1
2010 Travel Photo and Video Summarization with Cross-Media Correlation and Mutual Influence
Wei-Ta Chu, Che-Cheng Lin, Jen-Yu Yu
MMM1
2010 Consumer photo management and browsing facilitated by near-duplicate detection with feature filtering
Wei-Ta Chu
J. Vis. Commun. Image Represent.1
2010 Modeling spatiotemporal relationships between moving objects for event tactics analysis in tennis videos
Wei-Ta Chu, Wen-Ho Tsai
Multim. Tools Appl.1
2009 Using context information and local feature points in face clustering for consumer photos
abstract
We introduce local feature points to achieve face clustering for consumer photos. After combining eigenfaces with context information like clothes, we further investigate the usage of local feature points to match face images. The relationships between face images are constructed by feature matching and then described as a graph. Outliers in the results of preliminary clustering are detected and are re-clustered according to matching characteristics. We report complete performance comparison for different datasets and show that the proposed method has superior performance than conventional approaches.
Wei-Ta Chu, Ya-Lin Lee, Jen-Yu Yu
ICASSP1
2009 Automatic summarization of travel photos using near-duplication detection and feature filtering
abstract
We try to address part of the challenge proposed by CeWe. A photo summarization method is developed to select representative photos. Based on the observation that the most important objects/views are often captured several times, we exploit near-duplicate detection techniques to represent a sequence of photo as a graph, and then the graph structure is analyzed to facilitate importance ranking. We focus on summarizing hundreds of photos taken in journeys lasting for one or two weeks. The qualitative and quantitative measurement results demonstrate the effectiveness of the proposed method.
Wei-Ta Chu
ACM Multimedia1
2009 Feature classification for representative photo selection
abstract
This paper points out that different local feature points provide different impacts to near-duplicate detection and related applications. Aiming to automatic representative photo selection, we develop three feature classification methods, i.e., point-based, region-based, and pLSA-based classification, to differentiate local feature points described by SIFT descriptors. We investigate the performance of these classification methods, and discuss how they influence near-duplicate detection and extended applications. Experiments show that, with effective feature classification, more accurate representative selection results can be achieved.
Wei-Ta Chu, Jen-Yu Yu
ACM Multimedia1
2009 Visual language model for face clustering in consumer photos
abstract
For consumer photos, this work clusters faces with large variations in lighting, pose, and expression. After matching face images by local feature points, we transform matching situations into a novel representation called visual sentences. Then, visual language models are constructed to describe the dependency of image patches on faces. With the probabilistic framework, we develop a clustering algorithm to group the same individual's face images into the same cluster. An interesting observation about evaluating face clustering performance is proposed, and we demonstrate the superiority of the proposed visual language model approach.
Wei-Ta Chu, Ya-Lin Lee, Jen-Yu Yu
ACM Multimedia1
2009 A User Experience Model for Home Video Summarization
Wei-Ting Peng, Wei-Jia Huang, Wei-Ta Chu, Chien-Nan Chou, Wen-Yan Chang, Chia-Han Chang, Yi-Ping Hung
MMM3
2009 RoleNet: Movie Analysis from the Perspective of Social Networks
abstract
With the idea of social network analysis, we propose a novel way to analyze movie videos from the perspective of social relationships rather than audiovisual features. To appropriately describe role's relationships in movies, we devise a method to quantify relations and construct role's social networks, called RoleNet. Based on RoleNet, we are able to perform semantic analysis that goes beyond conventional feature-based approaches. In this work, social relations between roles are used to be the context information of video scenes, and leading roles and the corresponding communities can be automatically determined. The results of community identification provide new alternatives in media management and browsing. Moreover, by describing video scenes with role's context, social-relation-based story segmentation method is developed to pave a new way for this widely-studied topic. Experimental results show the effectiveness of leading role determination and community identification. We also demonstrate that the social-based story segmentation approach works much better than the conventional tempo-based method. Finally, we give extensive discussions and state that the proposed ideas provide insights into context-based video analysis.
Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu
IEEE Trans. Multim.2
2008 Event detection in tennis matches based on video data mining
abstract
This paper proposes a mining-based method to achieve event detection for broadcasting tennis videos. Utilizing visual and aural information, we extract some high-level features to describe video segments. The audiovisual features are further transformed to symbolic streams and an efficient mining technique is applied to derive all frequent patterns that characterize tennis events. After mining, we categorize frequent patterns into several kinds of events and therefore achieve event detection for tennis videos by checking the correspondence between mined patterns and events. The experimental results show that the proposed approach is a promising way to detect events in broadcasting tennis video.
Min-Chun Hu 0001, Yi-Tang Wang, Chen-Wei Chou, Kuei-Yi Hsieh, Wei-Ta Chu, Ja-Ling Wu
ICME5
2008 Automatic selection of representative photo and smart thumbnailing using near-duplicate detection
abstract
This paper presents two applications about representative photo selection and smart thumbnailing using the results of near-duplicate detection. For a given photo cluster, near-duplicate photo pairs are first determined, and the relationships between them are modeled by a graph. The most typical one is then automatically selected by examining the mutual relation between them. For smart thumbnailing, we determine the region-of-interest of the selected representative photo based on locally matched feature points, which is a view different from conventional saliency-based approaches. The experiments show satisfactory performance in representative selection and promising results in ROI determination.
Wei-Ta Chu
ACM Multimedia1
2008 Aesthetics-Based Automatic Home Video Skimming System
Wei-Ting Peng, Yueh-Hsuan Chiang, Wei-Ta Chu, Wei-Jia Huang, Wei-Lun Chang, Po-Chung Huang, Yi-Ping Hung
MMM3
2008 Explicit semantic events detection and development of realistic applications for broadcasting baseball videos
Wei-Ta Chu, Ja-Ling Wu
Multim. Tools Appl.1
2007 Exploring Broadcasting Baseball Videos Based on Multimodal and Multidisciplinary Study
abstract
This demonstration presents a comprehensive work covering semantic event detection, summary/highlight generation, query answering, and efficient browsing, for broadcasting baseball videos. This system integrates the techniques of content-based analysis, key-phrase spotting, natural language processing, data mining, and human-computer interface. It presents a multimodal and multidisciplinary work to show an instance for a media-rich life.
Wei-Ta Chu, Ja-Ling Wu
ICME1
2007 Movie Analysis Based on Roles' Social Network
abstract
Roles in a movie form a small society and their interrelationship provides clues for movie understanding. Based on this observation, we present a new viewpoint to perform semantic movie analysis. Through checking the co-occurrence of roles in different scenes, we construct a roles' social network to describe their relationships. We introduce the concept of social network analysis to elaborately identify leading roles and the hidden communities. With the results of community identification, we perform storyline detection that facilitates more flexible movie browsing and higher-level movie analysis. The experimental results show that the proposed community identification method is accurate and is robust to errors.
Chung-Yi Weng, Wei-Ta Chu, Ja-Ling Wu
ICME2
2006 Extraction of Baseball Trajectory and Physics-Based Validation for Single-View Baseball Video Sequences
abstract
To enrich the viewing experience of baseball games and provide some clues for enhancing pitcher's performance, we propose a Kalman filter-based approach to track ball trajectory from single-view pitching sequences. Without setting extraordinary equipments in stadiums or other sensing instruments, this approach robustly extracts ball trajectory for pitching sequences captured from TV channels or downloaded from the Internet. To validate the detected ball trajectories, we investigate the characteristics of ball trajectories on the basis of a baseball physical model. The effectiveness of ball trajectory extraction and ball position detection are presented
Wei-Ta Chu, Chia-Wei Wang, Ja-Ling Wu
ICME1
2006 Tiling slideshow
abstract
This paper presents a new medium, called tiling slideshow, to display photos in a tile-like manner, coordinating with the pace of background music. In contrast to the conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Motivated by the concepts of technical writing, each displaying layout is composed of a larger topic photo and several small-size supportive photos. Based on this idea, the proposed tiling slideshow system consists of three major components: image clustering, music analyzer, and layout organizer. Given the limited displaying space, we consider the context and relationship between photos and model the layout organization as a constrainted optimization problem. Experiments on real consumer photograph collections show that the novel displaying method gives users more pleasant browsing experience than the methods that focus only on single photograph display.
Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu
ACM Multimedia2
2006 Audiovisual slideshow: present your journey by photos
abstract
This demonstration presents a novel way to systematically display photos and enhance the viewing experience of photo browsing. In contrast to conventional photo slideshow, multiple photos that have similar characteristics are well arranged and displayed at the same layout. Moreover, the displaying pace is coordinated with the beat of the user-selected incidental music. To automatically generate the audiovisual slideshow, we develop a system that consists of three main components: photo analysis, music analysis, and audiovisual composition. Audiovisual content analysis and cross-media synchronization issues are addressed in this work. This novel demonstration is especially suitable to present photos taken in a journey. It vigorously presents the delights of traveling and helps us recall or experience the trip.
Jun-Cheng Chen, Wei-Ta Chu, Jin-Hau Kuo, Chung-Yi Weng, Ja-Ling Wu
ACM Multimedia2
2006 Development of realistic applications based on explicit event detection in broadcasting baseball videos
abstract
This paper presents a framework that explicitly detects events in broadcasting baseball videos and facilitates the development of various extended applications. Three phases are included: reliable shot classification, explicit event detection, and elaborate applications. In the shot classification stage, color and geometric information are utilized to classify shots into several canonical views. To explicitly detect semantic events, rule-based decision and model-based decision methods are developed. We emphasize that this system efficiently and exactly identifies what happened in baseball games rather than roughly finding some interesting parts. Based on explicit event detection, many accurate and practical applications such as box score generation and game summarization can be built. The evaluation results show the effectiveness of the proposed methods and demonstrate some insights about bridging semantic gaps for sports videos.
Wei-Ta Chu, Ja-Ling Wu
MMM1
2005 Integration of rule-based and model-based decision methods for baseball event detection
abstract
To exactly detect what events occur in baseball games, a framework that integrates rule-based and model-based decision methods is proposed. The rule-based decision module infers what happened by checking the information changes in the caption. The model-based decision module further classifies the events that could not be explicitly determined by checking caption information only. Thirteen events, including hit, double, home run, and so on, are considered in this work. The promising experimental results show the effectiveness of the proposed framework and facilitate the development of advanced video applications.
Wei-Ta Chu, Ja-Ling Wu
ICME1
2005 Generative and Discriminative Modeling toward Semantic Context Detection in Audio Tracks
abstract
Semantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. We propose a hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. For action movies we focused in this work, hidden Markov models (HMMs) are used to model four representative audio events, i.e. gunshot, explosion, car-braking, and engine sounds. At the semantic context level, generative (ergodic hidden Markov model) and discriminative (support vector machine, SVM) approaches are investigated to fuse the characteristics and correlations among various audio events, which provide cues for detecting gunplay and car-chasing scenes. The experimental results demonstrate the effectiveness of the proposed approaches and draw a sketch for semantic indexing and retrieval. Moreover, the differences between two fusion schemes are discussed to be the reference for future research.
Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu
MMM1
2005 Toward better retrieval and presentation by exploring cross-media correlations
Wei-Ta Chu, Herng-Yow Chen
Multim. Syst.1
2005 Toward semantic indexing and retrieval using hierarchical audio models
Wei-Ta Chu, Wen-Huang Cheng, Yung-Jen Hsu 0001, Ja-Ling Wu
Multim. Syst.1
2004 A study of semantic context detection by using SVM and GMM approaches
abstract
Semantic-level content analysis is a crucial issue to achieve efficient content retrieval and management. In this paper, we propose an hierarchical approach that models the statistical characteristics of several audio events over a time series to accomplish semantic context detection. Two stages, including audio event and semantic context modeling/testing, are devised to bridge the semantic gap between physical audio features and semantic concepts. HMM are used to model audio events, and SVM and GMM are used to fuse the characteristics of various audio events related to some specific semantic concepts. The experimental results show that the approach is effective in detecting semantic context. The comparison between SVM- and GMM-based approaches is also studied
Wei-Ta Chu, Wen-Huang Cheng, Ja-Ling Wu, Yung-Jen Hsu 0001
ICME1
2002 Cross-media correlation: a case study of navigated hypermedia documents
abstract
The research issues on multiple media correlation have arisen with more and more integrated multimedia applications. The multimedia correlation is used to coordinate different media and facilitate cross-media access. This paper presents our work on two types of multimedia correlation: explicit and implicit relations. We develop a system to carefully capture explicit relations and devise some computed synchronization processes to discover implicit relations between media objects. The proposed computed synchronization techniques, including speech-text alignment process in temporal domain, automatic scrolling process in spatial domain, and content dependency check process in content domain, will be addressed. Experimental results show that in the speech-text alignment process 80% of forced alignment are in-sync even the speech recognition accuracy is as low as 25%. The automatic scrolling process effectively maintains a resynchronization mechanism in different displaying environments.
Wei-Ta Chu, Herng-Yow Chen
ACM Multimedia1
2002 The WSML system: web-based synchronization multimedia lecture system
abstract
This demonstration presents a web-based multimedia lecture system that perfectly integrates multimedia lecturing with static HTML pages and enables their presentation in synchronization. We develop key techniques to reproduce vivid web-teaching scenarios for on-demand access, including events capturing scheme and synchronized mechanism. To facilitate multimedia lectures creation, the easy-to-use authoring and managing tools specially designed for teachers have been also included in the development. In addition, integrated presentation and efficient access issues are also discussed for the student's aspect.
Kuo-Yu Liu, Natalius Huang, Bo-Hung Wu, Wei-Ta Chu, Herng-Yow Chen
ACM Multimedia4