Wen-Jiin Tsai

dblp:87/5580 · DBLP profile ↗
← Back
47ranked-venue papers
15as first author
10since 2021 · last 2025
0000-0001-9030-5139ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 13 first-author · 10 since 2021Systems, architecture and hardware · 2 · 1 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 React: Reference-Based Anime Colorization Transformer
abstract
The animation industry is expanding, yet coloring frames, particularly sketch coloring, remains labor-intensive. Automating anime sketch colorization is challenging due to sparse details and frame variations. Existing methods struggle with color bleeding and misapplication. We present the Reference-based Anime Colorization Transformer (ReACT), which uses two reference images to transfer color to sketches. ReACT incorporates a density-aware distance field map to improve sketch semantics and preserve details, while a Transformer leverages multi-source attention for accurate color alignment. Experimental results demonstrate that ReACT delivers high-quality, continuous colorization, outperforming existing methods.
Li-Ting Wang, Wen-Jiin Tsai
ICIP2
2025 Automatic Natural Image Matting via Dual Encoder Aggregation
abstract
Traditional image matting requires manual trimap input, making it complex. Automatic image matting, however, extracts an alpha mask from natural images without needing trimaps or other inputs. While recent CNN-based methods perform well on images with clear subjects, they struggle with transparent or unclear objects. This is because CNNs capture local features, making it hard to handle scattered elements like fire or water droplets. On the other hand, Transformer-based encoders, which capture global features, perform better on such images but are less effective on images with distinct objects. To address this, we propose a hybrid matting network combining both CNN and Transformer encoders. Our model automatically predicts a trimap and detail map, combining them to produce the final alpha mask without auxiliary inputs. Trained on a composite dataset, it performs well on both clear and transparent subjects across four mainstream real-world datasets.
Meng-Lun Yu, Wen-Jiin Tsai
ICME2
2024 Shadow-Aware Makeup Transfer with Lighting Adaptation
abstract
In recent years, makeup transfer methods have shown satisfactory performance on reference images without shadows but struggle with those containing shadows. These methods often mistakenly include shadows as part of the makeup style, leading to poor results. To address this, we propose a shadow-aware makeup transfer approach. This method utilizes facial symmetry by flipping the face features and replacing shadowed areas with shadow-free makeup features. Additionally, we introduce a module to predict shadow maps, guiding makeup transfer from non-shadowed areas in the reference images. We also present a lighting-aware Pseudo Ground Truth (PGT) generator that evaluates shadow presence and lighting quality on the reference face, ensuring high-quality PGT. Experimental results demonstrate our method’s effectiveness in shadow removal and its ability to produce visually pleasing results.
Hao-Yun Chang, Wen-Jiin Tsai
ICIP2
2024 An Anchor-Free Contour-Based Method For Instance Segmentation
abstract
Instance segmentation methods can be broadly categorized into two types: mask-based and contour-based. Mask-based methods treat it as a pixel-level classification task, while Contour-based methods consider it a regression problem, predicting object boundary polygons. Regardless of the method, many approaches rely on object detectors to identify candidate bounding boxes, limiting performance to detector capabilities. In this paper, we propose a contour-based, anchor-free instance segmentation approach that eliminates the need to use an anchor box as assistance. Our approach leverages learnable initial contours and employs dynamic convolution to achieve this goal. The dynamic convolution utilizes a generated filter to predict an offset map. This map is then utilized by our deformation module to iteratively increase the number of predicted vertices and gradually refine the initial contours, aligning them with object boundaries and ultimately achieving the desired instance segmentation.
Tzu-Han Huang, Wen-Jiin Tsai
ICIP2
2024 Talking Head Generation Based on 3D Morphable Facial Model
abstract
This paper presents a framework for one-shot talking-head video generation which takes a single person image and audio clips as input and synthesizes photo-realistic videos with natural head-poses and lip motion synced to the driving audio. The main idea behind this framework is to use 3D Morphable Model (3DMM) parameters as intermediate representation in generating the videos. We design an Expression Predictor and a Head Pose Predictor to predict facial expression and head-pose parameters from audio, respectively, and adopt a 3DMM model to extract identity and texture parameters from the reference image. With these parameters, facial images are rendered as an auxiliary to guide video generation. Compared to widely used facial landmarks, 3DMM parameters are more powerful in representing facial details. Experimental results show that our method can generate realistic talking-head videos and outperform many state-of-the-art methods.
Hsin-Yu Shen, Wen-Jiin Tsai
PCS2
2023 Attention-based Video Virtual Try-On
abstract
This paper presents a video virtual try-on model which is based on appearance flow warping and is parsing-free. In this model, we utilized attention methods from Transformer [15] and proposed three attention-based modules: a Person-Cloth Transformer, a Self-Attention Generator, and a Cloth Refinement Transformer. The Person-Cloth Transformer enables clothing features to refer to person information, which is beneficial for style vector calculation and also improves the style warping process to estimate better appearance flows. The Self-Attention Generator utilizes a self-attention mechanism at the deepest feature layer, which enables the feature map to learn global context from all the other pixels, helping it synthesize more realistic results. The Cloth Refinement Transformer utilizes two cross-attention modules: one enables the current warped clothes to refer to previously warped clothes to ensure it is temporally consistent, and the other enables the current warped clothes to refer to person information to ensure it is spatially aligned. Our ablation study shows that each proposed module contributes to the improvement of the results. Experiment results show that our model can generate realistic try-on videos with high quality and perform better than existing methods.
Wen-Jiin Tsai, Yi-Cheng Tien
ICMR1
2023 Transformer-based spatial-temporal feature lifting for 3D hand mesh reconstruction
abstract
This paper presents a novel model for reconstructing hand meshes in video sequences. The model extends the MobRecon [1] pipeline and incorporates a variant of the Transformer architecture which effectively models both spatial and temporal relationships using distinct positional encodings. The Transformer encoder enhances the feature representation by modeling joint relationships and learning hidden depth information. Leveraging temporal information from consecutive frames, the Transformer decoder further enhances the feature representation for the mesh decoder’s final prediction. Additionally, we incorporate techniques such as Twice-LN, confidence-based attention, scaling in place of Softmax, and learnable encodings to improve the feature representation. Experimental results demonstrate the superiority of the proposed method over existing approaches.
Meng-Xue Lin, Wen-Jiin Tsai
VCIP2
2022 Component-Based Transformation for Person Image Generation
abstract
Person image generation has attracted more and more attention because it has many applications. This paper focuses on pose transfer and virtual try-on. We propose a flow-based method called attribute-decomposited spatial transformation network. It warps different body parts separately using segmentation and then fuses different body parts and generates images. The flow-based technique enables the model to generate high-quality person images for pose transfer and the attribute-decomposited technique enables the model to support virtual try-on. The proposed method was evaluated on Deepfashion dataset. Quantitative and qualitative experimental results show the superiority of the method.
Wen-Jiin Tsai, Po-Hsiang Chen
ICIP1
2021 Hierarchical Embedding Guided Network for Video Object Segmentation
abstract
Semi-supervised video object segmentation is to segment the target objects given the ground truth annotation of the first frame. Previous successful methods mostly rely on online learning or static image pre-train to improve accuracy. However, online learning methods require huge time costs at inference time, thus restrict their practical use. Methods with static image pre-train require heavy data augmentation that is complicated and time-consuming. This paper presents a fast Hierarchical Embedding Guided Network (HEGNet) which is only trained on Video Object Segmentation (VOS) datasets and does not utilize online learning. Our HEGNet integrates propagation-based and matching-based methods. It propagates the predicted mask of the previous frame as a soft cue and extracts hierarchical embedding at both deep and shallow layers to do feature matching. The produced label map of the deep layer is also used to guide the matching of the shallow layer. We evaluated our method on the DAVIS-2016 and DAVIS-2017 validation sets and achieved overall scores of 84.9% and 71.9% respectively. Our method surpasses the methods without online learning and static image pre-train and runs at 0.08 seconds per frame.
Chin-Hsuan Shih, Wen-Jiin Tsai
ICIP2
2021 Data Transformer for Anomalous Trajectory Detection
abstract
Anomaly detection is an important task in many traffic applications. Methods based on deep learning networks reach high accuracy; however, they typically rely on supervised training with large annotated data. Considering that anomalous data are not easy to obtain, we present data transformation methods which convert the data obtained from one intersection to other intersections to mitigate the effort of collecting training data. The proposed methods are demonstrated on the task of anomalous trajectory detection. A General model and a Universal model are proposed. The former focuses on saving data collection effort; the latter further reduces the network training effort. We evaluated the methods on the dataset with trajectories from four intersections in GTA V virtual world. The experimental results show that with significant reduction in data collecting and network training efforts, the proposed anomalous trajectory detection still achieves state-of-the-art accuracy.
Hsuan-Jen Psan, Wen-Jiin Tsai
VCIP2
2020 Joint Detection, Re-Identification, And Lstm In Multi-Object Tracking
abstract
Using Convolutional Neural Networks (CNN) in object tracking typically utilizes spatial features, while ignores the temporal correlation of frames in the whole film, causing that it is easy to lose the target when it is occluded by other objects. To cope with the problem, a robust system combining CNN and long short-term memory (LSTM) is proposed for multi-object tracking. The system consists of three modules: object detection, data association, and LSTM tracking. With the proposed approach, the tracking accuracy can be greatly improved especially when the tracking targets suffer from occlusion. Experimental results showed that the proposed system exhibits outstanding tracking accuracy and stability.
Wen-Jiin Tsai, Zih-Jie Huang, Chen-En Chung
ICME1
2020 DeepPear: Deep Pose Estimation and Action Recognition
abstract
Human action recognition has been a popular issue recently because it can be applied in many applications such as intelligent surveillance systems, human-robot interaction, and autonomous vehicle control. Human action recognition using RGB video is a challenging task because the learning of actions is easily affected by the cluttered background. To cope with this problem, the proposed method estimates 3D human poses first which can help remove the cluttered background and focus on the human body. In addition to the human poses, the proposed method also utilizes appearance features nearby the predicted joints to make our action prediction context-aware. Instead of using 3D convolutional neural networks as many action recognition approaches did, the proposed method uses a two-stream architecture that aggregates the results from skeleton-based and appearance-based approaches to do action recognition. Experimental results show that the proposed method achieved state-of-the-art performance on NTU RGB+D which is a large-scale dataset for human action recognition.
You-Ying Jhuang, Wen-Jiin Tsai
ICPR2
2018 Asymmetry dual-LFSR reseeding for low power BIST
Jen-Cheng Ying, Wang-Dauh Tseng, Wen-Jiin Tsai
Integr.3
2018 Understanding and Removal of False Contour in HEVC Compressed Images
abstract
A contour-like artifact called false contour is often observed in large smooth areas of decoded images and video. Without loss of generality, we focus on detection and removal of false contours resulting from the state-of-the-art High Efficiency Video Coding codec. First, we identify the cause of false contours by explaining the human perceptual experiences on them with specific experiments. Next, we propose a precise pixel-based false contour detection method based on the evolution of a false contour candidate (FCC) map. The number of points in the FCC map becomes fewer by imposing more constraints step by step. Special attention is paid to separating false contours from real contours such as edges and textures in the video source. Then, a decontour method is designed to remove false contours in the exact contour position while preserving edge/texture details. Extensive experimental results are provided to demonstrate the superior performance of the proposed false contour detection and removal method in both compressed images and videos.
Qin Huang 0006, Hui Yong Kim, Wen-Jiin Tsai, Seyoon Jeong, Jin Soo Choi, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.3
2017 Fast multi-view people localization using a torso-high reference plane
abstract
People locations bring rich information for a wide spectrum of applications in intelligent video surveillance systems. In addition to localization accuracy, computational efficiency is another significant issue to be highly concerned in people localization. As an essential early stage, people localization has to be accomplished in a very short time, enabling further semantic analysis. However, most state-of-the-art people localization methods pay little attention to computational efficiency. Hence, in this paper we propose an effective and efficient multi-view people localization scheme with several acceleration mechanisms. First, a torso-high reference plane is introduced since in general the torso part (after foreground segmentation) is more intact and stable than the other parts of a human body, and thus can predict potential people locations more reliably. Then, a novel and computationally efficient bitwise-operation scheme is proposed to predict people locations at the intersection regions of foreground line samples from multiple views. After rule-based validation, people locations can be accurately obtained and visualized on a real world plane. Experiments on multi-view surveillance videos not only validate the high accuracy of the proposed method in locating people under crowded scenes with serious occlusions, but also demonstrate an outstanding computational speed.
Chun-Chieh Hsu, Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee
VCIP3
2016 Adaptive accordion transformation based video compression method on HEVC
abstract
Fast processing speed and high compression ratio are always expected for video coding. But generally it is hard to compromise between each other since high compression ratio usually requires high coding complexities. Ouni et al. proposed a novel method called accordion-based (ACC) coding in which multiple video frames are combined into an accordion-like representation for compression and it can achieve high compression ratio with low coding complexity. For low-motion videos at high bit-rates, its coding efficiency is close to H.264/AVC. However, it may no longer have benefit when applying the ACC to HEVC because HEVC has much better coding efficiency than H.264/AVC. Besides, it is not likely that all the frames in a video sequence are low-motion and thus suitable for ACC coding. To cope with these problems, in this paper we propose an adaptive method which dynamically switches between accordion-based coding and traditional HEVC coding according to the characteristics of video frames. The experiment result shows that the proposed method can achieve up to 11.5% of total bit-rate saving and 56.1% of encoding time saving, compared to traditional HEVC coding.
Chia-Hsin Chan, Hong-Syun Ji, Wen-Jiin Tsai
VCIP3
2016 Incorporating frequent pattern analysis into multimodal HMM event classification for baseball videos
Hsuan-Sheng Chen, Wen-Jiin Tsai
Multim. Tools Appl.2
2015 Hybrid image retargeting
abstract
An image retargeting method is presented in this paper, which uses multiple operators to improve the performance. In the method, cropping is adopted first in order to remove non-important content to the desired aspect ratio while keeping significant content intact. If the desired aspect ratio cannot be met by cropping, an aspect ratio adjusting method is then adopted to fit the image to the desired aspect ratio. Uniform scaling is used in the last stage to adjust the image to the desired size. Since aspect ratio needs no change, the scaling can retain all the information without causing any distortion. Since different operators are applied in different stages, their benefits are fully utilized and drawbacks can be avoided. Experimental result shows that this approach can improve the retargeting performance at a low computation cost.
Wen-Jiin Tsai, Chun-Fu Chen 0003
VCIP1
2014 Foveation-based image quality assessment
abstract
Since human vision has much greater resolutions at the center of our visual field than elsewhere, different criteria of quality assessment should be applied on the image areas with different visual resolutions. This paper proposed a foveation-based image quality assessment method which adopted different sizes of windows in quality assessment for a single image. Visual salience models which estimate visual attention regions are used to determine the foveation center and foveation resolution models are used to guide the selection of window sizes for the areas over spatial extent of the image. Finally, the quality scores obtained from different window sizes are pooled together to get a single value for the image. The proposed method has been applied to IQA metrics, SSIM, PSNR, and UQI. The result shows that both Spearman and Kendall correlation coefficients can be improved significantly by our foveation-based method.
Wen-Jiin Tsai, Yi-Shih Liu
VCIP1
2014 A framework for video event classification by modeling temporal context of multimodal features using HMM
Hsuan-Sheng Chen, Wen-Jiin Tsai
J. Vis. Commun. Image Represent.2
2014 Rate-distortion optimized mode selection method for multiple description video coding
Wen-Jiin Tsai
Multim. Tools Appl.2
2014 Robust Video Coding Based on Hybrid Hierarchical B Pictures
abstract
Compressed video streams transmitted over error-prone environments are usually corrupted by transmission errors. Error-concealment techniques can be used to recover the lost information. In this paper, a hybrid model is proposed to improve the error-concealment performance. The model combines two hierarchical B-picture coding structures such that key-frames, reference B frames, or even nonreference B frames have buddy frames to serve as their data recovery frames when they are lost. With buddy frames, the distance between a lost frame and its recovering frame can be substantially reduced. In addition, an improved estimation method is also proposed to further increase the accuracy of recovering motion. Error-concealment performance can thus be significantly improved with little bit-rate redundancy. We have conducted experiments to compare its performance with other methods, and the results show that the proposed hybrid model outperforms these competed methods. The advantages of the proposed hybrid model are demonstrated in error-free and packet-loss environments.
Wen-Jiin Tsai, Po-Jui Chiu
IEEE Trans. Circuits Syst. Video Technol.1
2013 Error-resilient video coding using multiple reference frames
abstract
An error-resilient scheme based on multiple reference frames is proposed, which employs error-resilient frames (ER-frame) as part of reference frames and adopts a rate distortion optimized (RDO) technique to choose the blocks referring to ER-frames. To reduce the computational cost, two techniques are further proposed. Experimental results show that adopting ER-frames enhances error-resilient performance; using RDO makes the scheme adaptive to various packet loss rates and video sequences; and applying time reduction techniques reduces computational cost significantly with neglectable performance loss.
Wen-Jiin Tsai
ICIP1
2012 Error resilient video coding using hybrid hierarchical B pictures
abstract
In this paper, a hybrid model based on hierarchical B picture structure is proposed to improve error concealment effects when there is a whole-frame loss. The model combines two hierarchical B-picture coding structures such that key-frames, reference B frames, or even non-reference B frames have buddy frames to serve as their data recovery frames when they are lost. With buddy frames, the distance between a lost frame and its recovering frame can be substantially reduced and thus error concealment performance can be improved with little bit-rate redundancy.
Wen-Jiin Tsai, Wan-Han Liu
ICIP1
2012 Technicolor challenge: an event classification framework by probabilistic context modeling of multimodal features
abstract
Semantic high-level event recognition of videos is one of most interesting issues for multimedia searching and indexing. Since high-level events are usually domain-specific, a generic framework which can adapt itself to new domains without or with a few modifications is needed. To this end, this paper presents a generic framework for video event classification using temporal context of interval-based multimodal features. In the framework, a co-occurrence symbol transformation method is proposed to explore full temporal relations among multiple modalities in probabilistic HMM event classification. The results of our experiments on baseball video event classification demonstrate the superiority of the proposed approach.
Hsuan-Sheng Chen, Wen-Jiin Tsai
ACM Multimedia2
2012 Ball tracking and 3D trajectory approximation with applications to tactics analysis from single-camera volleyball sequences
Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee, Jen-Yu Yu
Multim. Tools Appl.2
2012 Multiple Description Video Coding Based on Hierarchical B Pictures Using Unequal Redundancy
abstract
Multiple description video coding (MDC) is one of the approaches for reducing the detrimental effects caused by transmission over error-prone networks. In this paper, a MDC model based on hierarchical B pictures is proposed to optimize the tradeoff between coding efficiency and error resilience. The model produces two descriptors by applying different MDC techniques such as duplication, spatial splitting and temporal splitting on the different frames of video sequences, taking into account unequal importance of frames at different hierarchical levels. Duplication (high redundancy) is for key frames: spatial splitting (medium redundancy) for reference B frames, and temporal splitting (low redundancy) for nonreference B frames. For one descriptor loss, the model applies different estimation methods, but for the two descriptor loss case, the same temporal estimation is employed. As a consequence, better error resilience can be achieved at high coding efficiency. The advantages of the proposed model are demonstrated in error-free and packet loss networks.
Wen-Jiin Tsai, Hao-Yu You
IEEE Trans. Circuits Syst. Video Technol.1
2011 Extraction and representation of human body for pitching style recognition in broadcast baseball video
abstract
In baseball games, different release points of pitchers form several kinds of pitching styles. Different pitching styles possess individual advantages. This paper presents a novel pitching style recognition approach for automatic generation of game information and video annotation. First, an effective object segmentation algorithm is designed to compute the body contour and extract the pitcher's body. Then, star skeleton is used as the representative descriptor of the pitcher posture for pitching style recognition. The proposed approach has been tested on broadcast baseball video and the promising experimental results validate the robustness and practicability.
Hua-Tsung Chen, Chien-Li Chou, Wen-Jiin Tsai, Suh-Yin Lee, Jen-Yu Yu
ICME3
2011 3D ball trajectory reconstruction from single-camera sports video for free viewpoint virtual replay
abstract
Free viewpoint video presentation is a new challenge in multimedia analysis. This paper presents an innovative physics-based scheme to reconstruct the 3D ball trajectory from single-camera volleyball video sequences for free viewpoint virtual replay. The problem of 2D-to-3D inference is arduous due to the loss of 3D information in projection to 2D images. The proposed scheme incorporates the domain knowledge of court specification and the physical characteristics of ball motion to accomplish the 2D-to-3D inference. Motion equations with the parameters are set up to define the 3D trajectories based on physical characteristics. Utilizing the geometric transformation of camera calibration, the 2D ball coordinates extracted over frames are used to approximate the parameters of the 3D motion equations, and finally the 3D ball trajectory can be reconstructed from single-camera sequences. The experiments show promising results. The reconstructed 3D trajectory enables the free viewpoint virtual replay and enriched visual presentation, making game watching a whole new experience.
Hua-Tsung Chen, Chien-Li Chou, Wen-Jiin Tsai, Suh-Yin Lee
VCIP3
2011 Screen-strategy analysis in broadcast basketball video using player tracking
abstract
In basketball games, screen is a blocking move performed by an offensive player, who stands beside or behind a defender, in order to free a teammate to shoot, to receive a pass, or to drive in to score. Screen is the fundamental essence that most offensive tactics are executed with. In this paper, a screen- strategy analysis system is designed, and through combining the identified screens, what tactics are executed in basketball games can be speculated. The proposed system is capable of court region detection, camera calibration and player extraction. Player trajectories are computed by a Kalman filter-based tracking method and mapped to the real-world court coordinates. The player position/trajectory information greatly assists professional-oriented applications such as screen-strategy analysis and tactic inference. The experiments on broadcast basketball videos show encouraging results.
Tsung-Sheng Fu, Hua-Tsung Chen, Chien-Li Chou, Wen-Jiin Tsai, Suh-Yin Lee
VCIP4
2010 Joint temporal and spatial multiple description coding for H.264 video
abstract
Multiple description video coding (MDC) is one of the techniques used to reduce the detrimental effects caused by transmission over error-prone networks. This paper presents a hybrid MDC method which segments the video along spatial and temporal dimensions, and provides efficient estimation methods for missing description reconstruction. The experimental results confirm the improved error resilience capability achieved by the proposed hybrid MDC in lossy networks.
Wen-Jiin Tsai
ICIP2
2010 Contour-based strike zone shaping and visualization in broadcast baseball video: providing reference for pitch location positioning and strike/ball judgment
Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee
Multim. Tools Appl.2
2010 A new unequal error protection scheme based on FMO
Jhong-Yu Shih, Wen-Jiin Tsai
Multim. Tools Appl.2
2010 Hybrid Multiple Description Coding Based on H.264
abstract
Multiple description (MD) video coding is one of the approaches that can be used to reduce the detrimental effects caused by transmission over error-prone networks. A number of approaches have been proposed for MD coding, where each provides a different tradeoff between compression efficiency and error resilience. This paper first presents two basic MD coding methods; one segments the video in the spatial domain, while the other in the frequency domain. Then a hybrid MD coding method is proposed. The hybrid MD encoder segments the video in both the spatial and frequency domains. In the case of data loss, the hybrid MD decoder takes advantage of the residual-pixel correlations in the spatial domain, and the coefficient correlations in the frequency domain, for error concealment. As a result, better error resilience can be achieved at high compression efficiency. The advantages of the proposed hybrid MD method are demonstrated in the contexts of descriptor loss in ideal channels and in packet-loss networks.
Chia-Wei Hsiao, Wen-Jiin Tsai
IEEE Trans. Circuits Syst. Video Technol.2
2010 Joint Temporal and Spatial Error Concealment for Multiple Description Video Coding
abstract
Transmission of compressed video signals over error-prone networks exposes the information to losses and errors. To reduce the effects of these losses and errors, this paper presents a joint spatial-temporal estimation method which takes advantages of data correlation in these two domains for better recovery of the lost information. The method is designed for the hybrid multiple description coding which splits video signals along spatial and temporal dimensions. In particular, the proposed method includes fixed and content-adaptive approaches for estimation method selection. The fixed approach selects the estimation method based on description loss cases, while the adaptive approach selects the method according to pixel gradients. The experimental results demonstrate that improved error resilience can be accomplished by the proposed estimation method.
Wen-Jiin Tsai
IEEE Trans. Circuits Syst. Video Technol.1
2010 Scene Change Aware Intra-Frame Rate Control for H.264/AVC
abstract
Most of rate-control research focuses on inter-coded frames, instead of intra-coded frames which are more possible to cause the problem of buffer overflow. This letter presents a rate control algorithm for intra-frame coding. We propose a Taylor series-based rate-QS model and a scene-change aware rate-QS model to determine quantization parameters for general intra frames and scene-change frames, respectively. Simulation results show that compared to competed approaches, the proposed method achieves better and stable quality with low buffer fullness.
Wen-Jiin Tsai, Ting-Li Chou
IEEE Trans. Circuits Syst. Video Technol.1
2009 Stance-based strike zone shaping and visualization in broadcast baseball video: Providing reference for pitch location positioning
abstract
In baseball, the strike zone plays an important role in each pitch. Pitches which pass through the strike zone count as strikes, three of which strike out the batter. Thus, pitchers should acquire mastery of the strike zone. Moreover, the strike zone also provides the reference for positioning the pitch locations, about which sports fans and professionals have an intense interest in compile statistics. This paper presents a stance-based approach for strike zone shaping and visualization in broadcast baseball video. We design efficient and effective algorithms to detect the home plate, contour the batter, locate the features points on the batter contour, and finally outline the strike zone. The experiments show that the proposed framework is able to shape the strike zone fairly well in various broadcast baseball video sequences, needless of manual operation and additional camera setting.
Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee
ICME2
2009 Physics-based ball tracking and 3D trajectory reconstruction with applications to shooting location estimation in basketball video
Hua-Tsung Chen, Min-Chun Hu 0001, Yi-Wen Chen, Wen-Jiin Tsai, Suh-Yin Lee
J. Vis. Commun. Image Represent.4
2008 A new unequal error protection scheme based on FMO
abstract
In this paper we present a slice group based unequal error protection (UEP) scheme for video transmission over error-prone networks. We propose a method to assign macroblock to slice groups based on a variation of H.264/AVC dispersed FMO mode and k-means clustering algorithm. In addition, a Converged Motion Estimation (CME) is proposed to further improve our UEP scheme. The idea behind the CME is to make the macroblocks be referenced in a skewed manner, such that highly important macroblocks are converged on only a few and the use of redundancy for error protection is efficient. Our experiments on many video sequences show promising results.
Jhong-Yu Shih, Wen-Jiin Tsai
ICIP2
2008 Human action recognition based on layered-HMM
abstract
We address the problem of human action understanding of the upper human body from video sequences. Time-sequential images expressing human actions are transformed to sequences of feature vectors containing the configuration of the human body. A human is modeled as a collection of body parts, linked in a kinematic structure. The relation of the joints is used to estimate the human pose. A proposed layered HMM framework decomposes the human action recognition problem into two layers. The first layer models the actions of two arms individually from low-level features. The second layer models the interrelationship of two arms as an action. Experiments with a set of six types of human actions demonstrate the effectiveness of our proposed scheme, and the comparisons with other HMM systems show the robustness.
Yen-Chieh Wu, Hsuan-Sheng Chen, Wen-Jiin Tsai, Suh-Yin Lee, Jen-Yu Yu
ICME3
2007 Pitch-by-Pitch Extraction from Single View Baseball Video Sequences
abstract
This paper presents a novel method for reducing a baseball video segment from one batter to next batter into a more compact pitch-by-pitch video by pitching ball trajectory detection. The pitch-by-pitch video shows the complete pitching and batting process and largely reduces the source video data, making pitching analysis of broadcast baseball sequences an easier task. The proposed method has been tested for several long sequences, and promising results are reported.
Hsuan-Sheng Chen, Hua-Tsung Chen, Wen-Jiin Tsai, Suh-Yin Lee, Jen-Yu Yu
ICME3
2007 A Tempo Analysis System for Automatic Music Accompaniment
abstract
This paper proposes a music accompaniment system capable of catching the tempo of music. The original signal is first reduced to a detection function revealing the pulses of music. To induce the tempo and locate beats, a Fourier analysis-based algorithm with high practicality and generality is designed that it is robust to beat strength and not restricted to specific music genres. The regularity and periodicity of music beats are extracted and even the missing beats are recovered. The performance is validated using a comprehensive testing data set, and results in both formal objective experiments and subjective listening evaluations show convincible performance. Moreover, an interactive interface is designed that users are allowed to select an instrument to accompany the music based on the obtained tempo.
Hua-Tsung Chen, Ming-Ho Hsiao, Wen-Jiin Tsai, Suh-Yin Lee, Jen-Yu Yu
ICME3
2007 Searching the Video: An Efficient Indexing Method for Video Retrieval in Peer to Peer Network
Ming-Ho Hsiao, Wen-Jiin Tsai, Suh-Yin Lee
MMM (2)2
1999 Buffer-Sharing Techniques in Service-Guaranteed Video Servers
Wen-Jiin Tsai, Suh-Yin Lee
Multim. Tools Appl.1
1998 Dynamic Buffer Management for Near Video-On-Demand Systems
Wen-Jiin Tsai, Suh-Yin Lee
Multim. Tools Appl.1
1997 Multi-Partition RAID: A New Method for Improving Performance of Disk Arrays under Failure
abstract
Disk arrays have been proposed as a way of improving I/O performance by using parallelism among multiple disks. This paper focuses, however, on improving the performance of disk array systems in the presence of disk failures, which are significant for applications where continuous operation is of concern. Although several approaches have been explored, the goals of achieving high performance and storage efficiency often conflict. In this paper, we propose a new variation of RAID organization, multi-partition RAID (mP-RAID), to improve storage efficiency and reduce performance degradation when disk failures occur. The idea is to recognize that frequently demanded data dominate the degree of the performance degradation when disk failures occur. mP-RAID subdivides a disk array into several partitions associated with different block organizations. Based upon data popularity, we assign data to appropriate partitions so that high performance and better storage efficiency can be achieved simultaneously.
Wen-Jiin Tsai, Suh-Yin Lee
Comput. J.1
1994 Storage Design and Retrieval of Continuous Multimedia Data Using Multi-Disks
abstract
In the domain of multimedia applications, continuous display is an important issue. In this paper, we present a practical method to allocate disk storage for multimedia data so that continuous requirement can be met. This technique explores data-transfer parallelism on a multidisk system. Moreover, in order to ensure that continuous retrieval can be achieved in a multiuser environment, we propose the dynamic scheduling mechanism for real-time object retrieval. It can be seen that, a good scheduling can explore higher access concurrency in display of multimedia applications. Several approaches with different trade-off based upon this mechanism are proposed in this paper, which include delay initiation, read-ahead, migration, segmentation and integration.
Wen-Jiin Tsai, Suh-Yin Lee
ICPADS1