VLDB 2026 Research / reviewers in the wild / expert
Shih-Yao Lin 0001
dblp:62/1050-1
· DBLP profile ↗
24ranked-venue papers
9as first author
5since 2021 · last 2026
0000-0003-3160-669XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoReframe: Context-Aware Horizontal-to-Vertical Video Transformation with Temporal Smoothness
Minmin Shen, Yarong Feng, Haiyun Jin, Ganesh Samarth, Chamarahalli Arunkumar, Shih-Yao Lin 0001, Solale Tabarestani, Caren Chen |
ICPR (1) | 8 |
| 2022 | AQT: Adversarial Query Transformers for Domain Adaptive Object DetectionabstractAdversarial feature alignment is widely used in domain adaptive object detection. Despite the effectiveness on CNN-based detectors, its applicability to transformer-based detectors is less studied. In this paper, we present AQT (adversarial query transformers) to integrate adversarial feature alignment into detection transformers. The generator is a detection transformer which yields a sequence of feature tokens, and the discriminator consists of a novel adversarial token and a stack of cross-attention layers. The cross-attention layers take the adversarial token as the query and the feature tokens from the generator as the key-value pairs. Through adversarial learning, the adversarial token in the discriminator attends to the domain-specific feature tokens, while the generator produces domain-invariant features, especially on the attended tokens, hence realizing adversarial feature alignment on transformers. Thorough experiments over several domain adaptive object detection benchmarks demonstrate that our approach performs favorably against the state-of-the-art methods. Source code is available at https://github.com/weii41392/AQT. Wei-Jie Huang, Yu-Lin Lu, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin |
IJCAI | 3 |
| 2021 | TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
Liangjian Chen, Deying Kong, Xingwei Liu, Hao Tang 0010, Xiangyi Yan, Yusheng Xie, Shih-Yao Lin 0001, Xiaohui Xie |
BMVC | 9 |
| 2021 | MVHM: A Large-Scale Multi-View Hand Mesh Benchmark for Accurate 3D Hand Pose EstimationabstractEstimating 3D hand poses from a single RGB image is challenging because depth ambiguity leads the problem ill-posed. Training hand pose estimators with 3D hand mesh annotations and multi-view images often results in significant performance gains. However, existing multi-view datasets are relatively small with hand joints annotated by off-the-shelf trackers or automated through model predictions, both of which may be inaccurate and can introduce biases. Collecting a large-scale multi-view 3D hand pose images with accurate mesh and joint annotations is valuable but strenuous. In this paper, we design a spin match algorithm that enables a rigid mesh model matching with any target mesh ground truth. Based on the match algorithm, we propose an efficient pipeline to generate a large-scale multi-view hand mesh (MVHM) dataset with accurate 3D hand mesh and joint labels. We further present a multi-view hand pose estimation approach to verify that training a hand pose estimator with our generated dataset greatly enhances the performance. Experimental results show that our approach achieves the performance of 0.990 in AUC20-50 on the MHP dataset compared to the previous state-of-the-art of 0.939 on this dataset. Our datasset is available at https://github.com/Kuzphi/MVHM. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Xiaohui Xie |
WACV | 2 |
| 2021 | Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in VideosabstractEstimating 3D hand pose directly from RGB images is challenging but has gained steady progress recently by training deep models with annotated 3D poses. However annotating 3D poses is difficult and as such only a few 3D hand pose datasets are available, all with limited sample sizes. In this study, we propose a new framework of training 3D pose estimation models from RGB images without using explicit 3D annotations, i.e., trained with only 2D information. Our framework is motivated by two observations: 1) Videos provide richer information for estimating 3D poses as opposed to static images; 2) Estimated 3D poses ought to be consistent whether the videos are viewed in the forward order or reverse order. We leverage these two observations to develop a self-supervised learning model called temporal-aware self-supervised network (TASSN). By enforcing temporal consistency constraints, TASSN learns 3D hand poses and meshes from videos with only 2D keypoint position annotations. Experiments show that our model achieves surprisingly good results, with 3D estimation accuracy on par with the state-of-the-art models trained with 3D annotations, highlighting the benefit of the temporal consistency in constraining 3D prediction models. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Xiaohui Xie |
WACV | 2 |
| 2020 | MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose SynthesisabstractEstimating the 3D hand pose from a monocular RGB image is important but challenging. A solution is training on large-scale RGB hand images with accurate 3D hand keypoint annotations. However, it is too expensive in practice. Instead, we develop a learning-based approach to synthesize realistic, diverse, and 3D pose-preserving hand images under the guidance of 3D pose information. We propose a 3D-aware multi-modal guided hand generative network (MM-Hand), together with a novel geometry-based curriculum learning strategy. Our extensive experimental results demonstrate that the 3D-annotated images generated by MM-Hand qualitatively and quantitatively outperform existing options. Moreover, the augmented data can consistently improve the quantitative performance of the state-of-the-art 3D hand pose estimators on two benchmark datasets. The code will be available at https://github.com/ScottHoang/mm-hand. Zhenyu Wu 0002, Duc Hoang, Shih-Yao Lin 0001, Yusheng Xie, Liangjian Chen, Yen-Yu Lin, Zhangyang Wang, Wei Fan 0001 |
ACM Multimedia | 3 |
| 2020 | DGGAN: Depth-image Guided Generative Adversarial Networks for Disentangling RGB and Depth Images in 3D Hand Pose EstimationabstractEstimating 3D hand poses from RGB images is essential to a wide range of potential applications, but is challenging owing to substantial ambiguity in the inference of depth information from RGB images. State-of-the-art estimators address this problem by regularizing 3D hand pose estimation models during training to enforce the consistency between the predicted 3D poses and the ground-truth depth maps. However, these estimators rely on both RGB images and the paired depth maps during training. In this study, we propose a conditional generative adversarial network (GAN) model, called Depth-image Guided GAN (DGGAN), to generate realistic depth maps conditioned on the input RGB image, and use the synthesized depth maps to regularize the 3D hand pose estimation model, therefore eliminating the need for ground-truth depth maps. Experimental results on multiple benchmark datasets show that the synthesized depth maps produced by DGGAN are quite effective in regularizing the pose estimation model, yielding new state-of-the-art results in estimation accuracy, notably reducing the mean 3D endpoint errors (EPE) by 4.7%, 16.5%, and 6.8% on the RHD, STB and MHP datasets, respectively. Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yen-Yu Lin, Wei Fan 0001, Xiaohui Xie |
WACV | 2 |
| 2019 | TAGAN: Tonality Aligned Generative Adversarial Networks for Realistic Hand Pose Synthesis
Liangjian Chen, Shih-Yao Lin 0001, Yusheng Xie, Yufan Xue, Yen-Yu Lin, Xiaohui Xie, Wei Fan 0001 |
BMVC | 2 |
| 2018 | Action Recognition with the Augmented MoCap Data using Neural Data Translation
Shih-Yao Lin 0001, Yen-Yu Lin |
BMVC | 1 |
| 2018 | IR drop prediction of ECO-revised circuits using machine learningabstractExcessive power supply noise (PSN), such as IR drop, can cause timing violation in VLSI chips. However, simulation PSN takes a very long time, especially when multiple iterations are needed in IR drop signoff. In this work, we propose a machine learning technique to build an IR drop prediction model based on circuits before ECO (engineer change order) revision. After revision, we can re-use this model to predict the IR drop of the revised circuit. Because the previous circuit(s) and the revised circuit are very similar, the model can be applied with small error. We proposed seven feature extractions, which are simple and scalable for large designs. Our experiment results show that prediction accuracy (average error 3.7mV) and correlation (0.55) are very high for a three million-gate real design. The run time speedup is up to 30X. The proposed method is very useful for designers to save the simulation time when fixing the IR drop problem. Shih-Yao Lin 0001, Yen-Chun Fang, Yu-Ching Li, Tsung-Shan Yang, Shang-Chien Lin, Chien-Mo James Li, Eric Jia-Wei Fang |
VTS | 1 |
| 2017 | Learning and inferring human actions with temporal pyramid features based on conditional random fieldsabstractFinding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle motions. Many existing approaches model actions by directly learning the transitional relationships between those fine-grained features. Yet, an action data may have many similar observations with occasional and irregular changes, which make commonly used fine-grained features less reliable. This paper presents a set of temporal pyramid features that enriches action representation with various levels of semantic granularities. For learning and inferring the proposed pyramid features, we adopt a discriminative model with latent variables to capture the hidden dynamics in each layer of the pyramid. Our method is evaluated on a Tai-Chi Chun dataset and a daily activities dataset. Both of them are collected by us. Experimental results demonstrate that our approach achieves more favorable performance than existing methods. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 1 |
| 2017 | Recognizing Human Actions with Outlier Frames by Observation Filtering and CompletionabstractThis article addresses the problem of recognizing partially observed human actions. Videos of actions acquired in the real world often contain corrupt frames caused by various factors. These frames may appear irregularly, and make the actions only partially observed. They change the appearance of actions and degrade the performance of pretrained recognition systems. In this article, we propose an approach to address the corrupt-frame problem without knowing their locations and durations in advance. The proposed approach includes two key components: outlier filtering and observation completion . The former identifies and filters out unobserved frames, and the latter fills up the filtered parts by retrieving coherent alternatives from training data. Hidden Conditional Random Fields (HCRFs) are then used to recognize the filtered and completed actions. Our approach has been evaluated on three datasets, which contain both fully observed actions and partially observed actions with either real or synthetic corrupt frames. The experimental results show that our approach performs favorably against the other state-of-the-art methods, especially when corrupt frames are present. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2014 | Opportunities for Persuasive Technology to Motivate Heavy Computer Users for Stretching Exercise
Yong-Xiang Chen, Siek-Siang Chiang, Shu-Yun Chih, Wen-Ching Liao, Shih-Yao Lin 0001, Shang-Hua Yang, Shun-Wen Cheng, Shih-Sung Lin, Yu-Shan Lin, Ming-Sui Lee, Jau-Yih Tsauo, Cheng-Min Jen, Chia-Shiang Shih, King-Jen Chang, Yi-Ping Hung |
PERSUASIVE | 5 |
| 2013 | Target-driven video summarization in a camera networkabstractNowadays, ever expanding camera network makes it difficult to find the suspect from lengthy video records. This paper proposes a target-driven video summarization framework which provides two-step Filtered Summarized Video (FSV) for tracing suspects. Before the target is identified, users can find the target efficiently using the firststep FSV of any arbitrary camera. The first-step FSV filters all the attributes of the target including the time information and the target's categories. After identifying the target, the second-step FSV with additional spatio-temporal & appearance cues are triggered in the neighbor cameras. To enhance the accuracy of the object classification for FSV, we propose a Perspective Dependent Model (PDM) which consists of many grid-based models. Finally, the experimental results show that grid-based model is more robust than general detectors and the user study demonstrates better performance for target finding and tracking in camera network for surveillance. Shen-Chi Chen, Shih-Yao Lin 0001, Kuan-Wen Chen, Chih-Wei Lin 0004, Chu-Song Chen, Yi-Ping Hung |
ICIP | 3 |
| 2013 | Real-time camera tampering detection using two-stage scene matchingabstractWe propose a tampering detection method using two-stage scene matching for real application with high efficiency and low false alarm rate. In the first stage, we use the intensity of edges as the main cue to detect the camera tampering events. Instead of using the entire edge points of the images, we sample the most significant edge points to represent the scene. Analyzing the edge variation with only the sample points, we discover that the events of camera tampering can be detected with low computation cost. Whenever the first stage detects the tampering event, the second stage is triggered to reduce false alarms. In the second stage, we propose an illumination change detector which can check the consistency of the scene structure using cell-based matching method. The experimental results demonstrate that our system can detect the camera tampering precisely and minimize false alarm even when the illumination changes dramatically or large crowds passing through the scene. Chao-Ching Shih, Shen-Chi Chen, Cheng-Feng Hung, Kuan-Wen Chen, Shih-Yao Lin 0001, Chih-Wei Lin 0004, Yi-Ping Hung |
ICME | 5 |
| 2013 | AirTouch panel: a re-anchorable virtual touch panelabstractTo achieve maximum mobility, device-less approaches for home appliance remote control have received increasing attention in recent years. In this paper, we propose a screen-less virtual touch panel, called AirTouch Panel, which can be positioned at any place with various orientations around users. The proposed virtual touch panel provides a potential ability to remotely control the home appliances, such as television, air conditioner, and so on. The proposed system allows users to anchor the panel at the place with comfortable poses. If the users want to change panel's position or orientation, they only need to re-anchor it, and then the panel will be reset. In this paper, our main contribution is to design a re-anchorable virtual panel for digital home remote control. Most importantly, we explore the design of such imaginary interface through two user studies. In our user studies, we analyze task completion time, satisfaction rate, and the number of miss-clicks. We are interested in the feasibility issues, for example, proper click gesture, panel size and button size, etc. Moreover, based on the AirTouch Panel, we also developed an intelligent TV to demonstrate the usability for controlling home appliance. Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Yi-Ping Hung |
ACM Multimedia | 1 |
| 2012 | 2D Face Alignment and Pose Estimation Based on 3D Facial ModelsabstractFace alignment and head pose estimation has become a thriving research field with various applications for the past decade. Several approaches process on 2D texture image but most of them perform decently only with small pose variation. Recently, many approaches apply depth information to align objects. However, applications are restricted because depth cameras are more expensive than common cameras, and many original image resources contain no depth information. Therefore, we propose a 3D face alignment algorithm in 2D image based on Active Shape Model, and use Speeded-Up Robust Features (SURF) descriptors as local texture model. We train a 3D shape model with different view-based local texture models from a 3D database, and then fit a face in a 2D image by these models. We also improve the performance by two-stage search strategy. Furthermore, the head pose can be estimated by the alignment result of the proposed 3D model. Finally, we demonstrate some applications applied by our method. Shen-Chi Chen, Chia-Hsiang Wu, Shih-Yao Lin 0001, Yi-Ping Hung |
ICME | 3 |
| 2012 | Human action recognition using Action Trait Code
Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Ming-Sui Lee, Yi-Ping Hung |
ICPR | 1 |
| 2012 | Action recognition for human-marionette interactionabstractIn this paper, we propose a human-marionette interaction system based on a human action recognition approach for applications to interactive artistic puppetry and a mimicking-marionette game. We developed an intelligent marionette called "i-marionette" that is controlled by a sophisticated control device to achieve various human actions. Moreover, we utilized an action recognition approach to enable the i-marionette to learn and recognize complex dance movements. The idea of artistic puppetry is to present a conflict scenario between two different cultural worlds: the performer is active and represents the culture of modern technology based in the real world. In contrast, the i-marionette represents traditional culture and is passive and based in a virtual world. The active performer guides the passive i-marionette to form a space-time connection between the real world and the virtual world. The i-marionette mimics the performer's action, while the performer also mimics the i-marionette's action. The performance represents an artistic conception in which humans invent technology and the i-marionette is manipulated by human control. However, in this interactive circle, the human is implicitly affected by the i-marionette. In our mimicking-marionette game, a player mimics the i-marionette's action. Subsequently, our human action recognition system measures the action similarity between the player and the i-marionette, and our system provides a similarity score. Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Yi-Ping Hung |
ACM Multimedia | 1 |
| 2011 | Novel projector calibration approaches of multi-resolution displayabstractThis paper proposes convenient and useful approaches to automatically calibrate the projectors of a multi-resolution display. The proposed approaches estimate both the keystone effect and misalignment of the projections with an assistance of a color camera. Structured light patterns are employed to construct the geometric relationship between projectors and the projection surface, and then pre-warp the images so that they appear undistorted as a result. Experimental results demonstrate that the proposed approaches successfully reduce the human-effort and lower the calibration time of multi-resolution display calibration task. Po-Hsun Chiu, Shih-Yao Lin 0001, Li-Wei Chan 0001, Neng-Hao Yu, Yi-Ping Hung |
ICME | 2 |
| 2010 | Real-Time 3D Model-Based Gesture Tracking for Multimedia ControlabstractThis paper presents a new 3D model-based gesture tracking system for controlling multimedia player in an intuitive way. The motivation of this paper is to make home appliance aware of user's intention. This 3D model-based gesture tracking system adopts a Bayesian framework to track the user's 3D hand position and to recognize meaning of these postures for controlling 3D player interactively. To avoid the high dimensionality of the whole 3D upper body model, which may complicate the gesture tracking problem, our system applies a novel hierarchical tracking algorithm to improve the system performance. Moreover, this system applies multiple cues for improving the accuracy of tracking results. Based on the above idea, we have implemented a 3D hand gesture interface for controlling multimedia players. Experimental results have shown that the proposed system robustly tracks the 3D position of the hand and has high potential for controlling the multimedia player. Shih-Yao Lin 0001, Yun-Chien Lai, Li-Wei Chan 0001, Yi-Ping Hung |
ICPR | 1 |
| 2010 | i-m-Space: interactive multimedia-enhanced space for rehabilitation of breast cancer patientsabstractThis paper presents i-m-Space, an interactive multimedia rehabilitation space that helps the post-surgery recovery of breast cancer patients. Our goal is to improve patients' physical therapy and psychological relaxation experience through careful applications of multimedia technology. i-m-Space consists of three types of breathing-based relaxation and three types for interactive exercise-based rehabilitation. Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Jin-Yao Lin, Szu-Wei Wu, Yi-Yu Chung, I-Ling Hu, Wei-Ting Peng, Shih-Yao Lin 0001, Chia-Han Chang, Pei-Hsuan Chou, King-Jen Chang, Mei-Lan Chang, Sue-Huei Chen, Jin-Shing Chen, Ming-Sui Lee, Mike Y. Chen, Yi-Ping Hung |
ACM Multimedia | 10 |
| 2010 | 3D human motion tracking based on a progressive particle filter
I-Cheng Chang, Shih-Yao Lin 0001 |
Pattern Recognit. | 2 |
| 2009 | Dynamic Kernel-Based Progressive Particle Filter for 3D Human Motion Tracking
Shih-Yao Lin 0001, I-Cheng Chang |
ACCV (2) | 1 |