VLDB 2026 Research / reviewers in the wild / expert
Tse-Yu Pan
dblp:169/3167
· DBLP profile ↗
26ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0001-8570-1575ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Investigating Generative Workflows for Immersive Prototyping: A Study on Scene- and Object-Level Creation in VRabstractThe integration of generative AI into immersive prototyping tools offers unprecedented potential to accelerate VR content creation. However, a critical gap exists in understanding the most effective interaction paradigms. We present DreamCraft, a novel VR prototyping system designed to integrate two fundamental workflows: scene-level creation, where users generate entire, navigable environments, and object-level creation, where scenes are iteratively populated with AI-generated assets. A formal user study with 12 VR designers was conducted to evaluate these workflows, evaluating both workflows through qualitative feedback and quantitative metrics, including task completion time, creative exploration, and perceived cognitive load. Our primary contribution is a set of empirically-grounded design guidelines derived from this analysis. These guidelines provide a crucial framework for the development of next-generation, user-centric generative tools for immersive virtual environments, directly informing future research and system design in the field. Cheng-Chih Tsai, Ping-Hsuan Han, Tse-Yu Pan |
VR | 3 |
| 2026 | DanceLDM: latent-based diffusion model for dance generation and editing conditioned on music and text promptabstractGenerating realistic 3D dance based on music is a unique and challenging task because real dance is a free and creative form of artistic expression. Human dance styles are diverse and ever-changing. Existing methods fail to perform well on out-of-domain music inputs or to generate movements of new dance styles. Accordingly, in this paper, we propose the Dance Latent Diffusion Model (DanceLDM), an advanced editable dance generation method for creating realistic and diverse 3D dance movements, while providing a text prompt interface for dance movements editing. Inspired by previous studies, we explore multimodal data and multi-task training for dance generation models, and we build a latent diffusion model to compare with existing approaches. Our method allows users to generate dance motions using either music or textual descriptions, while also enabling post-generation editing via text prompts. We evaluate our approach through quantitative and qualitative experiments. In quantitative analysis, DanceLDM achieves superior performance on the Beat Alignment Score (BAS), demonstrating its ability to generate rhythmically aligned dances. In qualitative studies, the participants rated our method as producing more realistic and musically coherent dances on generation and editing compared to existing baselines. Ming-Cong Su, Tse-Yu Pan |
Multim. Syst. | 3 |
| 2026 | Defect image generation using diffusion model for defect detection augmentationabstractAutomated visual inspection is critical for quality control in industrial manufacturing, yet training robust defect detection models is often hindered by the scarcity of defective samples. To address this challenge, we introduce a diffusion-based data augmentation framework capable of generating diverse and realistic synthetic defect images. Our method personalizes a text-to-image Latent Diffusion Model using DreamBooth, introducing a spatially weighted loss that forces the model to prioritize learning specific defect characteristics without attempting to reconstruct the non-defective background. To synthesize new data, we utilize Differential Diffusion to generate the learned defects onto defect-free images, eliminating boundary artifacts. We systematically evaluate the efficacy of our generated data on downstream supervised defect detection and localization tasks in two practical scenarios: as a complete substitute for real defect data and as a complement to it. Extensive experiments on the MVTec-AD dataset and a real-world industrial dataset demonstrate that our synthetic defects significantly boost downstream task performance. Furthermore, ablation studies confirm that these performance gains are robust and architecture-agnostic across various detectors. Herman Prawiro, Ching-Yeh Chiang, Nien-Yi Jan, Kai-Lin Yang, Yi-Rong Lin, Yung-Hui Li, Tse-Yu Pan, Min-Chun Hu 0001 |
Multim. Tools Appl. | 7 |
| 2025 | Movie Retrieval Systems Using Genre-Guided Multimodal Learning Techniques
Shintami Chusnul Hidayati, Tse-Yu Pan |
MMM (5) | 3 |
| 2025 | FencBuddy: Action-Aware Depth Perception Training for Fencing Attacks
Hung-Yao Peng, Zi-Heng Zhong, Cheng-Chih Tsai, Ching-Yeh Chiang, Tse-Yu Pan |
MMM (5) | 5 |
| 2025 | Fine-grained Stroke Recognition in Broadcast Table Tennis Videos with ATDTabstractThis study introduces an automated system for fine-grained stroke recognition in broadcast table tennis videos, designed to address challenges in manual annotation and tactical analysis during international competitions. The proposed framework integrates an Adaptive Temporal Difference Model with a Transformer Encoder (ATDT), leveraging a combination of Temporal Difference Networks (TDN) and Temporal Adaptive Modules (TAM) to enhance spatial and temporal feature extraction. To enhance feature discriminability, we employ supervised contrastive learning, which promotes better representation learning for fine-grained action recognition. The system is divided into two primary modules: the Action Segmentation Module (ASM) and the Action Recognition Module (ARM). ASM precisely identifies the start and end times of each stroke action by incorporating ball trajectory analysis to identify precise hit timings and placements. The precise segmentation facilitates the subsequent ARM to implement a three-stage recognition process: forehand and backhand classification, group-based classification, and intra-group action classification. This hierarchical approach improves the system’s ability to differentiate between subtle stroke variations, even under the constraints of low-resolution broadcast footage. To validate the framework, the MISTT dataset was collected, comprising 3,618 stroke action clips from 18 international matches, with professional player annotations. The proposed ATDT model outperformed existing methods, achieving a top-1 accuracy improvement of 18% for forehand strokes and 25.58% for backhand strokes compared to baseline models. Moreover, our automatic annotation system takes only 1/30 of the time compared to the manual annotation process, demonstrating its efficiency. Tang-Chen Chang, Duen-Chian Jheng, Hsuan-Ya Liang, Bill Louis Harchan, Pu Ching, Tsung-Hsun Tsai, Chih-Yi Chang, Te-Cheng Wu, Yung-Hui Li, Tse-Yu Pan, Hung-Kuo Chu, Min-Chun Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 10 |
| 2024 | Description-Driven Audiovisual Embedding Space Learning for Enhanced Movie UnderstandingabstractWith the rise of video streaming platforms, the number of videos has significantly increased, making auto movie tagging essential for better search and personalized recommendations.This paper presents a novel multimodal learning approach that not only aligns visual and auditory cues temporally and enhances their interrelation but also leverages movie descriptions and genre information to strengthen audiovisual feature extraction.Our approach has shown competitive results across nine tasks in the Long Video Understanding (LVU) benchmark, showing notable improvements in predicting directors, genres, and writers, thereby demonstrating its effectiveness in movie understanding and suitability for auto-tagging. Shao-Hung Wu, Hung-Chang Huang, Min-Chun Hu 0001, Tse-Yu Pan |
MMAsia | 5 |
| 2024 | VisionCoach: Design and Effectiveness Study on VR Vision Training for Basketball PassingabstractVision Training is important for basketball players to effectively search for teammates who has wide-open opportunities to shoot, observe the defenders around the wide-open teammates and quickly choose a proper way to pass the ball to the most suitable one. We develop an immersive virtual reality (VR) system called VisionCoach to simulate the player's viewing perspective and generate three designed systematic vision training tasks to benefit the cultivating procedure. By recording the player's eye gazing and dribbling video sequence, the proposed system can analyze the vision-related behavior to understand the training effectiveness. To demonstrate the proposed VR training system can facilitate the cultivation of vision ability, we recruited 14 experienced players to participate in a 6-week between-subject study, and conducted a study by comparing the most frequently used 2D vision training method called Vision Performance Enhancement (VPE) program with the proposed system. Qualitative experiences and quantitative training results are reported to show that the proposed immersive VR training system can effectively improve player's vision ability in terms of gaze behavior and dribbling stability. Furthermore, training in the VR-VisionCoach Condition can transfer the learned abilities to real scenario more easily than training in the 2D-VPE Condition. Pin-Xuan Liu, Tse-Yu Pan, Hsin-Shih Lin, Hung-Kuo Chu, Min-Chun Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Offensive Tactics Recognition in Broadcast Basketball Videos Based on 2D Camera View Player HeatmapsabstractIt is essential for sports teams to review their offensive and defensive tactical execution performance as well as understand their opponents’ tactics in order to identify effective counterattack strategies. This study focuses on basketball offensive tactics recognition based on 2D camera view heatmaps. Most of the current tactics recognition methods learn the spatiotemporal correlation of players based on top-view trajectory information. To obtain correct top-view player trajectories, robust camera calibration and player tracking techniques are indispensable. However, for broadcast videos having large camera movement, serious player occlusions, and similar players’ jerseys, it is quite challenging to obtain accurate camera parameters and player tracking results, resulting in poor tactical analysis performance. Instead of applying camera calibration and player tracking, this study attempts to design a tactics recognition method that directly predicts the tactics class from 2D camera-view player heatmaps in the inference phase. Our proposed method uses a recurrent convolutional neural network with coordinate embedding to directly identify the tactics. Moreover, an auxiliary top-view player trajectory reconstruction module is added in the training phase to acquire better latent codes to represent the tactics. The experimental results show that for both supervised and unsupervised settings, our proposed method achieves comparable accuracy to the current tactics classification methods that rely on perfect top-view trajectory input. subst Nico, Tse-Yu Pan, Herman Prawiro, Jain-Wei Peng, Wen-Cheng Chen, Hung-Kuo Chu, Min-Chun Hu 0001 |
ICMR | 2 |
| 2023 | SetterVision: Motion-based Tactical Training System for Volleyball Setters in Virtual RealityabstractVolleyball, a sport characterized by unpredictable factors such as ball trajectory, teammate actions, and strategic positioning, presents a challenge when it comes to modeling and training due to its high levels of complexity. Successful gameplay relies on the coordinated efforts of all team members in the receiving, setting, and attacking phases. In real-life competitions, the setter's on-ball ability and decision-making are particularly crucial to the team's offensive success: To improve the training of setters in observing player movements while running and making informed attacking decisions, we propose the design of a virtual reality (VR) system which aims to enhance players' setting skills and strategic thinking to achieve more successful offensive plays with a lower cost. Chen-Wei Fu, Ming-Cong Su, Hsin-Yu Huang, Tse-Yu Pan |
ACM Multimedia | 7 |
| 2023 | TelEmoScatter: Enabling Remote Interaction and Emotional Connections in Virtual and Physical Music PerformanceabstractTo enrich the emotional experiences of virtual reality (VR) online audiences in music performances, we developed TelEmoScatter, a system that facilitates remote interaction between music performers and onsite audiences. Our system also fosters emotional connections for online audiences through sound-visualization conversion, which is influenced by the state of the onsite audiences using computer vision techniques. In this work, we generate a 3D space using real-time sound-visualization techniques by converting MIDI signals from musical instruments into dynamic animations. Additionally, we employ video analysis to predict the emotions of the onsite audience, allowing seamless integration of emotional visual cues into the virtual scene. With our system, users can effortlessly immerse themselves in the emotional expressions of performers through music and experience the unique atmosphere of a live performance venue simply by wearing a VR headset. Chen-Wei Fu, Pin-Xuan Liu, Ming-Cong Su, Ping-Hsuan Han, Tse-Yu Pan |
MMAsia | 8 |
| 2023 | Efficient Hand Gesture Recognition using Multi-Task Multi-Modal Learning and Self-DistillationabstractIn this paper, we propose a lightweight model for hand gesture recognition using an RGB camera. The proposed model enables recognition of first-person hand gestures using a single camera and achieves near-real-time computational performance on both high-end and low-end computing devices. The proposed framework utilizes multi-task multi-modal learning and self-distillation to deal with the challenges in hand gesture recognition. We integrate additional modalities (depth) and a future prediction mechanism to enhance the model’s ability to learn spatio-temporal information. Furthermore, we employ self-distillation to compress the model, achieving a balance between accuracy and computational efficiency. We compared the proposed hand gesture recognition model with the state-of-the-art method, and our model outperforms the SOTA by 0.88% and 3.52% on the EgoGesture and NVGesture datasets, respectively. In terms of computational efficiency, our model takes only 161ms in average to recognize a gesture on a device with low-end GPUs (NVIDIA Jetson TX2), which is acceptable for interaction in XR applications. Jie-Ying Li, Herman Prawiro, Chia-Chen Chiang, Hsin-Yu Chang, Tse-Yu Pan, Chih-Tsun Huang, Min-Chun Hu 0001 |
MMAsia | 5 |
| 2022 | ScoreActuary: Hoop-Centric Trajectory-Aware Network for Fine-Grained Basketball Shot AnalysisabstractWe propose a fine-grained basketball shot analysis system called ScoreActuary to analyze the players' shot events, which can be applied to game analysis, player training, and highlight generation. Given a basketball video as input, our system first detects/segments shot candidates and then analyzes "Shot Type", "Shot Result", and "Ball Status" of each shot candidate in real-time. Our approach is composed of a customized object detector and a trajectory-aware network to learn the information of ball trajectory. Compared to the existing methods that analyze basketball shots, our algorithm can better handle videos with arbitrary camera movements while improving the accuracy. To the best of our knowledge, this work is the first system that can analyze fine-grained shot events accurately in real basketball games with arbitrary camera movements. Ting-Yang Kao, Tse-Yu Pan, Chen-Ni Chen, Tsung-Hsun Tsai, Hung-Kuo Chu, Min-Chun Hu 0001 |
ACM Multimedia | 2 |
| 2022 | BetterSight: Immersive Vision Training for Basketball PlayersabstractVision training is important for athletes and is key to win a sports game. Traditional vision training methods are suitable for sports that focus only on the ball. For basketball, however, players need to observe multiple moving objects (i.e., ball and players) concurrently on a large court. We propose BetterSight, an immersive vision training system for basketball that not only trains the vision of the player but also requires the player to dribble the ball stably, mimicking the situation in a real basketball game. BetterSight is composed of an Interaction Module (IM), a Training Content Generation Module (TCGM), and an Analysis Module (AM). IM allows the trainee to interact with the system more intuitively based on gesture and speech rather than the controller. TCGM simulates the training scenarios based on the training configurations selected by the trainee. AM collects the video sequences capturing the trainee and the trainee's eye movements during the training phase, and then analyzes the trainee's gaze, dribbling movements, and number of dribbles. The analyzed data can be used to evaluate the training effectiveness of using the proposed BetterSight. Pin-Xuan Liu, Tse-Yu Pan, Hsin-Shih Lin, Hung-Kuo Chu, Min-Chun Hu 0001 |
ACM Multimedia | 2 |
| 2022 | StimulusLoop: Game-Actuated Mutuality Artwork for Evoking Affective StateabstractAs transmission technology has advanced, large-scale media data is delivered to users. However, machines may collect user behavior while users are receiving messages and analyze how to stimulate users' senses to grab more attention. To demonstrate the relationship between machines and users, we propose a game-actuated mutuality artwork, StimulusLoop. We design the visualization to present the users' behavior and affective state and design the game mechanics that one participant throws the dart to change the video watched by the other participant. The interactions between two participants form a game loop, and different kinds of messages passing between the two participants are visualized as mutuality artwork. Tai-Chen Tsai, Tse-Yu Pan, Min-Chun Hu 0001, Ya-Lun Tao |
ACM Multimedia | 2 |
| 2022 | GetWild: A VR Editing System with AI-Generated 3D Object and Terrainabstract3D environment artists typically use 2D screens and 3D modeling software to achieve their creation. However, creating 3D content using 2D tools is counterintuitive. Moreover, the process would be inefficient for junior artists in the absence of a reference. We develop a system called GetWild, which employs artificial intelligence (AI) models to generate the prototype of 3D objects/terrain and allows users to further edit the generated content in the virtual space. With the aid of AI, the user can capture an image to obtain a rough 3D object model, or start with drawing simple sketches representing the river, the mountain peak and the mountain ridge to create a 3D terrain prototype. Further, the virtual reality (VR) technique is used to provide an immersive design environment and intuitive interaction (such as painting, sculpturing, coloring, and transformation) for users to edit the generated prototypes. Compared with the existing 3D modeling software and systems, the proposed VR editing system with AI-generated 3D objects/terrain provides a more efficient way for the user to create virtual artwork. Shing Ming Wong, Chien-Wen Chen, Tse-Yu Pan, Hung-Kuo Chu, Min-Chun Hu 0001 |
ACM Multimedia | 3 |
| 2022 | A Hierarchical Hand Gesture Recognition Framework for Sports Referee Training-Based EMG and Accelerometer SensorsabstractTo cultivate professional sports referees, we develop a sports referee training system, which can recognize whether a trainee wearing the Myo armband makes correct judging signals while watching a prerecorded professional game. The system has to correctly recognize a set of gestures related to official referee's signals (ORSs) and another set of gestures used to intuitively interact with the system. These two gesture sets involve both large motion and subtle motion gestures, and the existing sensor-based methods using handcrafted features do not work well on recognizing all kinds of these gestures. In this work, deep belief networks (DBNs) are utilized to learn more representative features for hand gesture recognition, and selective handcrafted features are combined with the DBN features to achieve more robust recognition results. Moreover, a hierarchical recognition scheme is designed to first recognize the input gesture as a large or subtle motion gesture, and the corresponding classifiers for large motion gestures and subtle motion gestures are further used to obtain the final recognition result. Moreover, the Myo armband consists of eight-channel surface electromyography (sEMG) sensors and an inertial measurement unit (IMU), and these heterogeneous signals can be fused to achieve better recognition accuracy. We take basketball as an example to validate the proposed training system, and the experimental results show that the proposed hierarchical scheme considering DBN features of multimodality data outperforms other methods. Tse-Yu Pan, Wan-Lun Tsai, Chen-Yuan Chang, Chung-Wei Yeh, Min-Chun Hu 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Feasibility Study on Virtual Reality Based Basketball Tactic TrainingabstractIn this article, a VR-based basketball training system comprising a standalone VR device and a tablet is proposed. The system is intended to improve the ability of players to understand offensive tactics and practice these tactics correctly. We compare the training effectiveness of various degrees of immersion, including a conventional basketball tactic board, a 2D monitor, and virtual reality. A multi-camera-based human tracking system was designed and built around a real-world basketball court to record and analyze the running trajectory of each player during tactical execution. The accuracy of the running path and hesitation time at each tactical step were evaluated for each participant. Furthermore, we assessed several subjective measurements, including simulator sickness, presence, and sport imagery ability, to conduct a more comprehensive exploration of the feasibility of the proposed VR framework for basketball tactics training. The results indicate that the proposed system is useful for learning complex tactics. Furthermore, high VR immersion training improves athletes' abilities with regards to strategic imagery. Wan-Lun Tsai, Tse-Yu Pan, Min-Chun Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Does Virtual Odor Representation Influence the Perception of Olfactory Intensity and Directionality in VR?abstractIntroducing olfactory display in the virtual reality (VR) system brings the immersive experience to new heights. However, it is intractable to simulate olfactory features (such as the intensity and the direction) with multiple levels. Visual stimuli have been proved to dominate human perception among multiple sensors in virtual environments. If visual stimuli can be used to guide the olfactory sense in VR, the design of the olfactory display can be simpler but still able to provide olfactory experience with more diversity. To understand the visual-olfactory effect on different olfactory characteristics, a portable olfactory display that can control the intensity and direction of odors was developed. An experimental study was conducted to investigate cross-modal human perception, i.e. how the visually virtual odor representation in VR influences human perception of real odor produced by the proposed olfactory display. The results showed that the perception of odor intensity and directionality can be modulated by visually virtual odor representation. Shou-En Tsai, Wan-Lun Tsai, Tse-Yu Pan, Chia-Ming Kuo, Min-Chun Hu 0001 |
VR | 3 |
| 2020 | Emotion Recognition from Galvanic Skin Response Signal Based on Deep Hybrid Neural NetworksabstractEmotion reacts human beings' physiological and psychological status. Galvanic Skin Response (GSR) can reveal the electrical characteristics of human skin and is widely used to recognize the presence of emotion. In this work, we propose an emotion recognition frame-work based on deep hybrid neural networks, in which 1D CNN and Residual Bidirectional GRU are employed for time series data analysis. The experimental results show that the proposed method can outperform other state-of-the-art methods. In addition, we port the proposed emotion recognition model on Raspberry Pi and design a real-time emotion interaction robot to verify the efficiency of this work. Imam Yogie Susanto, Tse-Yu Pan, Chien-Wen Chen, Min-Chun Hu 0001, Wen-Huang Cheng |
ICMR | 2 |
| 2020 | An Empirical Study of Emotion Recognition from Thermal Video Based on Deep Neural NetworksabstractEmotion recognition is a crucial problem in affective computing. Most of previous works utilized facial expression from visible spectrum data to solve emotion recognition task. Thermal videos provide temperature measurement of human body over time, which can be used to recognize affective states by learning its temporal pattern. In this paper, we conduct comparative experiments to study the effectiveness of the existing deep neural networks when applied to emotion recognition task from thermal video. We analyze the effect of various approaches for frame sampling in video, temporal aggregation between frames, and different convolutional neural network architectures. To the best of our knowledge, we are the first w ork t o c onduct s tudy on emotion recognition from thermal video based on deep neural networks. Our work can provide preliminary study to design new methods for emotion recognition in thermal domain. Herman Prawiro, Tse-Yu Pan, Min-Chun Hu 0001 |
VCIP | 2 |
| 2019 | Robust Basketball Player Tracking Based on a Hybrid Detection Grouping Framework for Overlapping CamerasabstractWe propose a robust basketball player tracking framework for multi-cameras which have high portion of overlapping with each other and are set at human height. A novel detection grouping method is proposed to more correctly merge the projected detection results. Instead of using linear motion assumption to predict the human motion, we applied a regional consistency assumption to calculate the motion affinity. Further-more, we design a one-to-one clustering method to associate the most matching tracklets together using correlation values between tracklets and generate final trajectory results. Since there is no public labeled overlapping cross-cameras basketball dataset, we collected our own dataset, MISBasketball, and labeled the ground truth to evaluate the proposed tracking framework. Kuan-Hsien Wu, Wan-Lun Tsai, Tse-Yu Pan, Min-Chun Hu 0001 |
IEEE BigData | 3 |
| 2019 | Furniture style compatibility recommendation with cross-class triplet loss
Tse-Yu Pan, Yi-Zhu Dai, Min-Chun Hu 0001, Wen-Huang Cheng |
Multim. Tools Appl. | 1 |
| 2017 | Deep model style: Cross-class style compatibility for 3D furniture within a sceneabstractHarmonizing the style of all the furniture placed within a constrained space/scene has been regarded as one of the most important tasks in interior design. Most previous style analysis works measure the style similarity or compatibility of the objects based on predefined geometric features extracted from 3D models. However, “style” is a high-level semantic concept, which is difficult to be described explicitly by handcrafted geometric features. Deep neural network has been claimed to have more powerful ability to mimic the perception of human visual cortex. Therefore, in this work we utilize Triplet Convolutional Neural Network (Triplet CNN) to analyze style compatibility between 3D furniture models of different classes (e.g., a table and a lamp). It should be noted that analyzing the style compatibility between two or more furniture of different classes is quite difficult, as the given furniture may have distinctive structures or geometric elements. We conducted experiments based on a collected dataset containing 420 textured 3D furniture models. A group of raters were recruited from Amazon Mechanical Turk (AMT) to evaluate the comparative suitability of paired models within the dataset. The experimental results reveal that the proposed furniture style compatibility method based on deep learning is better than the state-of-the-art method and can be used for furniture recommendation. Tse-Yu Pan, Yi-Zhu Dai, Wan-Lun Tsai, Min-Chun Hu 0001 |
IEEE BigData | 1 |
| 2017 | A Sensor-Based Official Basketball Referee Signals Recognition System Using Deep Belief Networks
Chung-Wei Yeh, Tse-Yu Pan, Min-Chun Hu 0001 |
MMM (1) | 2 |
| 2015 | B-box Mixer: An Interactive UI for Generating B-box MusicabstractB-box is a form of vocal percussion that imitates rhythms in various types of sound, especially musical instruments. As b-box becoming popular, more and more people want to learn b-box and make their own b-box music. However, not everyone has the talent for generating harmonic b-box music. In this work, we develop an interactive system which helps the user easily compose b-box music given two inputs: an unaccompanied vocal song and a piece of b-box rhythm. The audio signals of the two inputs are analyzed and adaptively matched on the basis of their beats. The state of the art beat detection technique does not perform well on vocal songs. Hence, we propose to partition the song into short segments and estimate the average tempo for each segment so that the adjustment of tempo will not be affected too much by the wrongly detected beats. With the proposed system, people who love b-box or are not familiar with b-box can enjoy producing their own b-box music. Yi-Zhu Dai, Ting-Chia Lee, Xin-Yu Kuo, Tse-Yu Pan, Min-Chun Hu 0001 |
ACM Multimedia | 4 |