EDBT 2026 Demo / reviewers in the wild / expert
Tse-Yu Pan
dblp:169/3167
· DBLP profile ↗
7ranked-venue papers in the field
1as first author
4since 2021 · last 2024
0000-0001-8570-1575ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 3Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Description-Driven Audiovisual Embedding Space Learning for Enhanced Movie UnderstandingabstractWith the rise of video streaming platforms, the number of videos has significantly increased, making auto movie tagging essential for better search and personalized recommendations.This paper presents a novel multimodal learning approach that not only aligns visual and auditory cues temporally and enhances their interrelation but also leverages movie descriptions and genre information to strengthen audiovisual feature extraction.Our approach has shown competitive results across nine tasks in the Long Video Understanding (LVU) benchmark, showing notable improvements in predicting directors, genres, and writers, thereby demonstrating its effectiveness in movie understanding and suitability for auto-tagging. Shao-Hung Wu, Hung-Chang Huang, Min-Chun Hu 0001, Tse-Yu Pan |
MMAsia | 5 |
| 2023 | Offensive Tactics Recognition in Broadcast Basketball Videos Based on 2D Camera View Player HeatmapsabstractIt is essential for sports teams to review their offensive and defensive tactical execution performance as well as understand their opponents’ tactics in order to identify effective counterattack strategies. This study focuses on basketball offensive tactics recognition based on 2D camera view heatmaps. Most of the current tactics recognition methods learn the spatiotemporal correlation of players based on top-view trajectory information. To obtain correct top-view player trajectories, robust camera calibration and player tracking techniques are indispensable. However, for broadcast videos having large camera movement, serious player occlusions, and similar players’ jerseys, it is quite challenging to obtain accurate camera parameters and player tracking results, resulting in poor tactical analysis performance. Instead of applying camera calibration and player tracking, this study attempts to design a tactics recognition method that directly predicts the tactics class from 2D camera-view player heatmaps in the inference phase. Our proposed method uses a recurrent convolutional neural network with coordinate embedding to directly identify the tactics. Moreover, an auxiliary top-view player trajectory reconstruction module is added in the training phase to acquire better latent codes to represent the tactics. The experimental results show that for both supervised and unsupervised settings, our proposed method achieves comparable accuracy to the current tactics classification methods that rely on perfect top-view trajectory input. subst Nico, Tse-Yu Pan, Herman Prawiro, Jain-Wei Peng, Wen-Cheng Chen, Hung-Kuo Chu, Min-Chun Hu 0001 |
ICMR | 2 |
| 2023 | TelEmoScatter: Enabling Remote Interaction and Emotional Connections in Virtual and Physical Music PerformanceabstractTo enrich the emotional experiences of virtual reality (VR) online audiences in music performances, we developed TelEmoScatter, a system that facilitates remote interaction between music performers and onsite audiences. Our system also fosters emotional connections for online audiences through sound-visualization conversion, which is influenced by the state of the onsite audiences using computer vision techniques. In this work, we generate a 3D space using real-time sound-visualization techniques by converting MIDI signals from musical instruments into dynamic animations. Additionally, we employ video analysis to predict the emotions of the onsite audience, allowing seamless integration of emotional visual cues into the virtual scene. With our system, users can effortlessly immerse themselves in the emotional expressions of performers through music and experience the unique atmosphere of a live performance venue simply by wearing a VR headset. Chen-Wei Fu, Pin-Xuan Liu, Ming-Cong Su, Ping-Hsuan Han, Tse-Yu Pan |
MMAsia | 8 |
| 2023 | Efficient Hand Gesture Recognition using Multi-Task Multi-Modal Learning and Self-DistillationabstractIn this paper, we propose a lightweight model for hand gesture recognition using an RGB camera. The proposed model enables recognition of first-person hand gestures using a single camera and achieves near-real-time computational performance on both high-end and low-end computing devices. The proposed framework utilizes multi-task multi-modal learning and self-distillation to deal with the challenges in hand gesture recognition. We integrate additional modalities (depth) and a future prediction mechanism to enhance the model’s ability to learn spatio-temporal information. Furthermore, we employ self-distillation to compress the model, achieving a balance between accuracy and computational efficiency. We compared the proposed hand gesture recognition model with the state-of-the-art method, and our model outperforms the SOTA by 0.88% and 3.52% on the EgoGesture and NVGesture datasets, respectively. In terms of computational efficiency, our model takes only 161ms in average to recognize a gesture on a device with low-end GPUs (NVIDIA Jetson TX2), which is acceptable for interaction in XR applications. Jie-Ying Li, Herman Prawiro, Chia-Chen Chiang, Hsin-Yu Chang, Tse-Yu Pan, Chih-Tsun Huang, Min-Chun Hu 0001 |
MMAsia | 5 |
| 2020 | Emotion Recognition from Galvanic Skin Response Signal Based on Deep Hybrid Neural NetworksabstractEmotion reacts human beings' physiological and psychological status. Galvanic Skin Response (GSR) can reveal the electrical characteristics of human skin and is widely used to recognize the presence of emotion. In this work, we propose an emotion recognition frame-work based on deep hybrid neural networks, in which 1D CNN and Residual Bidirectional GRU are employed for time series data analysis. The experimental results show that the proposed method can outperform other state-of-the-art methods. In addition, we port the proposed emotion recognition model on Raspberry Pi and design a real-time emotion interaction robot to verify the efficiency of this work. Imam Yogie Susanto, Tse-Yu Pan, Chien-Wen Chen, Min-Chun Hu 0001, Wen-Huang Cheng |
ICMR | 2 |
| 2019 | Robust Basketball Player Tracking Based on a Hybrid Detection Grouping Framework for Overlapping CamerasabstractWe propose a robust basketball player tracking framework for multi-cameras which have high portion of overlapping with each other and are set at human height. A novel detection grouping method is proposed to more correctly merge the projected detection results. Instead of using linear motion assumption to predict the human motion, we applied a regional consistency assumption to calculate the motion affinity. Further-more, we design a one-to-one clustering method to associate the most matching tracklets together using correlation values between tracklets and generate final trajectory results. Since there is no public labeled overlapping cross-cameras basketball dataset, we collected our own dataset, MISBasketball, and labeled the ground truth to evaluate the proposed tracking framework. Kuan-Hsien Wu, Wan-Lun Tsai, Tse-Yu Pan, Min-Chun Hu 0001 |
IEEE BigData | 3 |
| 2017 | Deep model style: Cross-class style compatibility for 3D furniture within a sceneabstractHarmonizing the style of all the furniture placed within a constrained space/scene has been regarded as one of the most important tasks in interior design. Most previous style analysis works measure the style similarity or compatibility of the objects based on predefined geometric features extracted from 3D models. However, “style” is a high-level semantic concept, which is difficult to be described explicitly by handcrafted geometric features. Deep neural network has been claimed to have more powerful ability to mimic the perception of human visual cortex. Therefore, in this work we utilize Triplet Convolutional Neural Network (Triplet CNN) to analyze style compatibility between 3D furniture models of different classes (e.g., a table and a lamp). It should be noted that analyzing the style compatibility between two or more furniture of different classes is quite difficult, as the given furniture may have distinctive structures or geometric elements. We conducted experiments based on a collected dataset containing 420 textured 3D furniture models. A group of raters were recruited from Amazon Mechanical Turk (AMT) to evaluate the comparative suitability of paired models within the dataset. The experimental results reveal that the proposed furniture style compatibility method based on deep learning is better than the state-of-the-art method and can be used for furniture recommendation. Tse-Yu Pan, Yi-Zhu Dai, Wan-Lun Tsai, Min-Chun Hu 0001 |
IEEE BigData | 1 |