EDBT 2026 Demo / reviewers in the wild / expert
Chiou-Ting Hsu
dblp:70/6821
· DBLP profile ↗
79ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0001-8857-2481ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 12 first-author · 16 since 2021Artificial intelligence and machine learning · 26 · 15 since 2021Security and privacy · 3 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OCC-FAS: A New Benchmark and Feature-Disentangled Mixture-of-Experts Framework for Occlusion-Aware Face Anti-spoofing
Jun-Ren Chen, Cheng-Hsiang Su, Yi-Chen Ou, Kai-Heng Chien, Pei-Kai Huang, Chiou-Ting Hsu |
ICPR (11) | 7 |
| 2026 | BASIL-rPPG: Basis Learning with Predictive rPPG Reconstruction for Heart Rate Estimation from Ultra-Short Facial Videos
Jhih-Wei Jhao, Wen-Pin Chen, Jun-Ren Chen, Yen-Chun Chou, Shih-Yu Yang, Pei-Kai Huang, Chiou-Ting Hsu |
ICPR (8) | 7 |
| 2026 | Multi-modal face anti-spoofing via cross-modal feature transitions
Jun-Xiong Chong, Fang-Yu Hsu, Ming-Tsung Hsu, Kai-Heng Chien, Chiou-Ting Hsu, Pei-Kai Huang |
Expert Syst. Appl. | 6 |
| 2026 | UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake DetectionabstractIn deepfake detection, the varying degrees of compression employed by social media platforms pose significant challenges for model generalization and reliability. Although existing methods have progressed from single-modal to multimodal approaches, they face critical limitations: single-modal methods struggle with feature degradation under data compression in social media streaming, while multimodal approaches require expensive data collection and labeling and suffer from inconsistent modal quality or accessibility in real-world scenarios. To address these challenges, we propose a novel Unimodal-generated Multimodal Contrastive Learning (UMCL) framework for robust cross-compression-rate (CCR) deepfake detection. In the training stage, our approach transforms a single visual modality into three complementary features: compression-robust rPPG signals, temporal landmark dynamics, and semantic embeddings from pre-trained vision-language models. These features are explicitly aligned through an affinity-driven semantic alignment (ASA) strategy, which models inter-modal relationships through affinity matrices and optimizes their consistency through contrastive learning. Subsequently, our cross-quality similarity learning (CQSL) strategy enhances feature robustness across compression rates. Extensive experiments demonstrate that our method achieves superior performance across various compression rates and manipulation types, establishing a new benchmark for robust deepfake detection. Notably, our approach maintains high detection accuracy even when individual features degrade, while providing interpretable insights into feature relationships through explicit alignment. Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin |
Int. J. Comput. Vis. | 6 |
| 2026 | Fully test-time rPPG estimation via synthetic signal-guided feature learning
Pei-Kai Huang, Tzu-Hsien Chen, Ya-Ting Chan, Kuan-Wen Chen, Shih-Yu Yang, Yen-Chun Chou, Chiou-Ting Hsu |
Pattern Recognit. | 7 |
| 2026 | Ultra-short rPPG estimation via periodicity guidance and signal reconstruction
Pei-Kai Huang, Ya-Ting Chan, Kuan-Wen Chen, Chiou-Ting Hsu, Xiaoding Wang, Mohammad Jalil Piran |
Pattern Recognit. | 4 |
| 2026 | Enhancing learnable descriptive convolutional vision transformer for face anti-spoofing
Pei-Kai Huang, Jun-Xiong Chong, Ming-Tsung Hsu, Fang-Yu Hsu, Kai-Heng Chien, Chiou-Ting Hsu |
Pattern Recognit. | 7 |
| 2025 | SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image PretrainingabstractFace anti-spoofing (FAS) plays a pivotal role in ensuring the security and reliability of face recognition systems. With advancements in vision-language pretrained (VLP) models, recent two-class FAS techniques have leveraged the advantages of using VLP guidance, while this potential remains unexplored in one-class FAS methods. The one-class FAS focuses on learning intrinsic liveness features solely from live training images to differentiate between live and spoof faces. However, the lack of spoof training data can lead one-class FAS models to inadvertently incorporate domain information irrelevant to the live/spoof distinction (\eg, facial content), causing performance degradation when tested with a new application domain. To address this issue, we propose a novel framework called Spoof-aware one-class face anti-spoofing with Language Image Pretraining (SLIP). Given that live faces should ideally not be obscured by any spoof-attack-related objects (\eg, paper, or masks) and are assumed to yield zero spoof cue maps, we first propose an effective language-guided spoof cue map estimation to enhance one-class FAS models by simulating whether the underlying faces are covered by attack-related objects and generating corresponding nonzero spoof cue maps. Next, we introduce a novel prompt-driven liveness feature disentanglement to alleviate live/spoof-irrelative domain variations by disentangling live/spoof-relevant and domain-dependent information. Finally, we design an effective augmentation strategy by fusing latent features from live images and spoof prompts to generate spoof-like image features and thus diversify latent spoof features to facilitate the learning of one-class FAS. Our extensive experiments and ablation studies support that SLIP consistently outperforms previous one-class FAS methods. Pei-Kai Huang, Jun-Xiong Chong, Cheng-Hsuan Chiang, Tzu-Hsien Chen, Tyng-Luh Liu, Chiou-Ting Hsu |
AAAI | 6 |
| 2025 | Channel difference transformer for face anti-spoofing
Pei-Kai Huang, Jun-Xiong Chong, Ming-Tsung Hsu, Fang-Yu Hsu, Chiou-Ting Hsu |
Inf. Sci. | 5 |
| 2025 | DD-rPPGNet: De-Interfering and Descriptive Feature Learning for Unsupervised rPPG Estimation
Pei-Kai Huang, Tzu-Hsien Chen, Ya-Ting Chan, Kuan-Wen Chen, Chiou-Ting Hsu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Prompt-guided Multi-modal contrastive learning for Cross-compression-rate Deepfake Detection
Ching-Yi Lai, Chiou-Ting Hsu, Chih-Chung Hsu, Chia-Wen Lin |
BMVC | 2 |
| 2024 | One-Class Face Anti-Spoofing via Spoof Cue Map-Guided Feature LearningabstractMany face anti-spoofing (FAS) methods have focused on learning discriminative features from both live and spoof training data to strengthen the security of face recognition systems. However, since not every possible attack type is available in the training stage, these FAS methods usually fail to detect unseen attacks in the inference stage. In comparison, one-class FAS, where training data comprise only live faces, aims to detect whether a test face image belongs to the live class or not. In this paper, we propose a novel One-Class Spoof Cue Map estimation Network (OC-SCMNet) to address the one-class FAS detection problem. Our first goal is to learn to extract latent spoof features from live images so that their estimated Spoof Cue Maps (SCMs) should have zero responses. To avoid trapping to a trivial solution, we devise a novel SCM-guided feature learning by combining many SCMs as pseudo ground-truths to guide a conditional generator to create latent spoof features for spoof data. Our second goal is to simulate the potential out-of-distribution spoof attacks approximately. To this end, we propose using a memory bank to dynamically preserve a set of sufficiently “independent” latent spoof features to encourage the generator to probe the latent spoof feature space. Extensive experiments conducted on eight FAS benchmark datasets demonstrate that the proposed OC-SCMNet not only outperforms previous one-class FAS approaches but also achieves performance comparable to the state-of-the-art two-class FAS methods. The code is available at https://github.com/Pei-KaiHuang/CVPR24_OC_SCMNet. Pei-Kai Huang, Cheng-Hsuan Chiang, Tzu-Hsien Chen, Jun-Xiong Chong, Tyng-Luh Liu, Chiou-Ting Hsu |
CVPR | 6 |
| 2023 | Test-Time Adaptation for Robust Face Anti-Spoofing
Pei-Kai Huang, Chen-Yu Lu, Shu-Jung Chang, Jun-Xiong Chong, Chiou-Ting Hsu |
BMVC | 5 |
| 2023 | Single-Domain Generalization for Semantic Segmentation Via Dual-Level Domain AugmentationabstractThe goal of single-domain generalization is to learn a domain- generalized model from only one single source domain. To avoid overfitting to the source domain, recent research focused on domain augmentation for learning domain generalized features. Therefore, domain diversity is indeed crucial to the generalization ability of the model. In this paper, we propose a novel dual-level domain augmentation framework to enrich the domain diversity for single-domain generalized semantic segmentation. We specifically devise an Image-Level and a Class-Level Augmentation Module (IAM and CAM) to enlarge the diversity of augmented images and per-class features, respectively. From the original and augmented data, we then design a Domain-Generalized Feature Learning to learn representative features regularized by a large-scale pretrained model. Experimental results on semantic segmentation benchmarks demonstrate the effectiveness and outperformance of the proposed method over previous work. Shu-Jung Chang, Chen-Yu Lu, Pei-Kai Huang, Chiou-Ting Hsu |
ICIP | 4 |
| 2023 | LDCformer: Incorporating Learnable Descriptive Convolution to Vision Transformer for Face Anti-SpoofingabstractFace anti-spoofing (FAS) aims to counter facial presentation attacks and heavily relies on identifying live/spoof discriminative features. While vision transformer (ViT) has shown promising potential in recent FAS methods, there remains a lack of studies examining the values of incorporating local descriptive feature learning with ViT. In this paper, we propose a novel LDCformer by incorporating Learnable Descriptive Convolution (LDC) with ViT and aim to learn distinguishing characteristics of FAS through modeling long-range dependency of locally descriptive features. In addition, we propose to extend LDC to a Decoupled Learnable Descriptive Convolution (Decoupled-LDC) for improving the optimization efficiency. With the new Decoupled-LDC, we further develop an extended model LDCformerDfor FAS. Extensive experiments on FAS benchmarks show that LDCformerDoutperforms previous methods on most of the protocols in both intra-domain and cross-domain testings. The codes are available at https://github.com/Pei-KaiHuang/ICIP23_D-LDCformer. Pei-Kai Huang, Cheng-Hsuan Chiang, Jun-Xiong Chong, Tzu-Hsien Chen, Hui-Yu Ni, Chiou-Ting Hsu |
ICIP | 6 |
| 2023 | Towards Diverse Liveness Feature Representation and Domain Expansion for Cross-Domain Face Anti-SpoofingabstractFace anti-spoofing (FAS) aims to strengthen security of facial identity authentication by distinguishing live faces from spoof ones. Although disentangled feature learning has achieved much success in FAS, the representation capacity of disentangled feature space remains limited and does not extend beyond the training domains. In this paper, we propose to further augment the disentangled liveness and domain features with a two-fold goal. Our first goal is to enrich the diversity of liveness features so as to encompass a wide range of facial representation attacks. The second goal is to expand the domain features toward well-generalized and unseen domains. To reach the two goals, we develop a Disentangled Feature Augmentation Network (DFANet) with two feature augmentation strategies, including Affine Feature Transformation (AFT) and Adversarial Domain Learning (ADL). Extensive experiments on four FAS benchmark datasets show that the proposed DFANet outperforms previous methods on most of the protocols under cross-domain testings. The codes are available at https://github.com/Jxchong1999/DFANet. Pei-Kai Huang, Jun-Xiong Chong, Hui-Yu Ni, Tzu-Hsien Chen, Chiou-Ting Hsu |
ICME | 5 |
| 2022 | Domain Generalized RPPG Network: Disentangled Feature Learning with Domain Permutation and Domain Augmentation
Wei-Hao Chung, Cheng-Ju Hsieh, Sheng-Hung Liu, Chiou-Ting Hsu |
ACCV (2) | 4 |
| 2022 | Learnable Descriptive Convolutional Network for Face Anti-Spoofing
Pei-Kai Huang, Hui-Yu Ni, Yanqin Ni, Chiou-Ting Hsu |
BMVC | 4 |
| 2022 | Towards Robust In-domain and Out-of-Domain Generalization: Contrastive Learning with Prototype Alignment and Collaborative Attention
Yuan-Jhe Kuo, Cheng-Yu Yang, Chiou-Ting Hsu |
BMVC | 3 |
| 2022 | Augmentation of rPPG Benchmark Datasets: Learning to Remove and Embed rPPG Signals via Double Cycle Consistent Learning from Unpaired Facial Videos
Cheng-Ju Hsieh, Wei-Hao Chung, Chiou-Ting Hsu |
ECCV (16) | 3 |
| 2022 | Learning to Augment Face Presentation Attack Dataset via Disentangled Feature Learning from Limited Spoof DataabstractFace presentation attack detection methods have been de-veloped to counter presentation attacks and achieved consid-erable success, thanks to large training data and newly developed deep-learning technology. However, when encountering the attacks provided with few training examples, the learning-based detection methods tend to overfit to the small dataset and lead to poor generalization. In this paper, we study this scenario and propose to augment the limited data via disentangled feature learning. We include the live/spoof classifi-cation task and the person identification task in a multi-task learning framework to disentangle the liveness and identity features. To enlarge the number of training samples, we de-sign two remixing strategies on the disentangled features under the identity preservation constraint and the reconstruction constraint, and also adopt the idea of contrastive learning to ensure the discriminability of the augmented samples. Exper-imental results on several benchmark datasets show that the proposed augmentation method significantly improves many detection methods under the limited data scenario. Pei-Kai Huang, Chu-Ling Chang, Hui-Yu Ni, Chiou-Ting Hsu |
ICME | 4 |
| 2022 | Source Free Domain Adaptation for Semantic Segmentation via Distribution Transfer and Adaptive Class-Balanced Self-TrainingabstractUnsupervised Domain Adaptation (UDA) for semantic seg-mentation aims to transfer the knowledge learned from the source domain to the target domain. Unlike the source-available UDA setting, Source-Free Domain Adaptation (SFDA) has no access to the source data and rely solely on the well-trained source model for adaptation. Without the source data for reference, SDFA often leads to unstable adaptation and mostly focuses on common semantic classes. In this pa-per, we propose a Distribution Transfer and Adaptive Class-balanced self-training (DTAC) framework to tackle the issues of SFDA for semantic segmentation. First, in the distribution transfer stage, we propose to narrow the domain gap by aligning the implicit feature characteristics of source model with the feature statistics of the target data. Next, in the self-training stage, we propose a multi-class negative learning method with adaptive thresholding to dynamically select per-class pseudo labels for self-supervision. Experimental re-sults on urban scene benchmarks show that DTAC outper-forms other SFDA baselines and even achieves competitive results with source-available UDA methods. Cheng-Yu Yang, Yuan-Jhe Kuo, Chiou-Ting Hsu |
ICME | 3 |
| 2021 | Self-Guided Adversarial Learning For Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation has been introduced to generalize semantic segmentation models from labeled synthetic images to unlabeled real-world images. Although much effort was devoted to minimize the cross-domain gap, the segmentation results on real-world data remain highly unstable. In this paper, we discuss two main issues which hinder previous methods from achieving satisfactory results and propose a novel self-guided adversarial learning to leverage the capability of domain adaptation. Firstly, to deal with the unpredictable data variation in the real-world domain, we develop a self-guided adversarial learning method by selecting reliable target pixels as guidance to lead the adaptation of the other pixels. Secondly, to address the class-imbalanced issue, we devise the selection strategy in each class independently and incorporate this idea with a class-level adversarial learning in a unified framework. Experimental results show that the proposed method significantly improves the previous methods on several benchmark datasets. Yu-Ting Pang, Jui Chang, Chiou-Ting Hsu |
ICIP | 3 |
| 2020 | Multi-task Learning for Simultaneous Video Generation and Remote Photoplethysmography Estimation
Yun-Yun Tsou, Yi-An Lee, Chiou-Ting Hsu |
ACCV (5) | 3 |
| 2019 | Vision-Based Heart Rate Estimation Via A Two-Stream CNNabstractRemote photoplethysmography (rPPG) is a non-contact method for heart rate (HR) estimation from facial videos. In this paper, we propose a novel two-stream convolutional neural network for remote HR estimation. We introduce a feature extraction stream by adopting a low-rank constraint to guide the network to learn a robust feature representation. We also develop a complementary stream, the rPPG extraction stream, to extract reliable rPPG signals from facial regions. After fusing the two streams, we develop a unified neural network to learn the feature extraction and to estimate HR simultaneously. Experimental results on COHFACE dataset demonstrate that our proposed method achieves state-of-the-art performance for HR estimation. Zhi-Kuan Wang, Ying Kao, Chiou-Ting Hsu |
ICIP | 3 |
| 2018 | Generative Adversarial Guided Learning for Domain Adaptation
Kai-Ya Wei, Chiou-Ting Hsu |
BMVC | 2 |
| 2017 | AST-Net: An Attribute-based Siamese Temporal Network for Real-Time Emotion Recognition
Shu-hui Wang, Chiou-Ting Hsu |
BMVC | 2 |
| 2016 | Towards Deep Style Transfer: A Content-Aware Perspective
Yi-Lei Chen, Chiou-Ting Hsu |
BMVC | 2 |
| 2016 | Semantic Segmentation for Real-World Data by Jointly Exploiting Supervised and Transferrable Knowledge
Li-Hsien Lu, Chiou-Ting Hsu |
BMVC | 2 |
| 2015 | Revealing Smooth Structure of Visual Data by Permutation on ManifoldsabstractIn this paper, we address the issue of visual data organization by recovering an intrinsic order from an unorganized dataset. The proposed method exploits the inherent nature of manifold. This new perspective, posing no hypothesis on local topology of observed data, is simply built on the smoothness prior of manifold geometry. Under the observation that strong relation exists among visual content, we assume a visual dataset lies on a manifold and thus changes smoothly from point to point. By exploiting the linearity within nearby data points, our goal becomes to visit all of the data points along a manifold-guided order and to characterize the specific manifold’s shape. Yi-Lei Chen, Chiou-Ting Hsu |
BMVC | 2 |
| 2015 | Nonparametric scene parsing with deep convolutional features and dense alignmentabstractThis paper addresses two key issues which concern the performance of nonparametric scene parsing: (1) the semantic quality of image retrieval; and (2) the accuracy in label transfer. First, because nonparametric methods annotate a query image through transferring labels from retrieved images, the task of image retrieval should find a set of “semantically similar” images to the query. Second, with the retrieval set, a good strategy should be developed to transfer semantic labels in pixel-level accuracy. In this paper, we focus on improving scene parsing accuracy in these two issues. We propose using the state-of-the-art deep convolutional features as image descriptors to improve the semantic quality of retrieved images. In addition, we include dense alignment into the Markov Random Field inference framework to transfer labels at pixel-level accuracy. Our experiments on the SIFT Flow dataset shows the improvement of the proposed approach over other nonparametric methods. Chih-Hao Ma, Chiou-Ting Hsu, Benoit Huet |
ICIP | 2 |
| 2015 | Single-Image Dehazing via Optimal Transmission Map Under Scene PriorsabstractThe challenge of single-image dehazing mainly comes from double uncertainty of scene radiance and scene transmission. Most existing methods focus on restoring the visibility of hazy images and tend to derive a rough estimate of scene transmission. Unlike previous work, in this paper we advocate the significance of accurate transmission estimation and recast our problem as deriving the optimal transmission map directly from the haze model under two scene priors. We introduce theoretic and heuristic bounds of scene transmission to guide the optimum and show that the proposed theoretic bound happens to justify the well-known dark channel prior of haze-free images. With the constraints on the solution space, we then incorporate two scene priors, including locally consistent scene radiance and context-aware scene transmission, to formulate a constrained minimization problem and solve it by quadratic programming. The global optimality is guaranteed. Simulations on synthetic data set quantitatively verify the accuracy and show that the transmission map successfully captures fine-grained depth boundaries. Experimental results on color/gray-level images demonstrate that our method outperforms most state of the arts in terms of both accurate transmission maps and realistic haze-free images. Yi-Hsuan Lai, Yi-Lei Chen, Chuan-Ju Chiou, Chiou-Ting Hsu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Dual Subspace Nonnegative Graph Embedding for Identity-Independent Expression RecognitionabstractFacial expression is one of the intricate biometric traits, where different persons exhibit various appearance changes when posing the same expression. Because facial cues involved in the recognition of facial expression are not fully separate from that of facial identity, this identity-dependent behavior often complicates automatic facial expression recognition. In this paper, to address the identity-independent expression recognition problem, we propose a dual subspace nonnegative graph embedding (DSNGE) to represent expressive images using two subspaces: 1) identity subspace and 2) expression subspace. The identity subspace characterizes identity-dependent appearance variations; whereas the expression subspace characterizes identity-independent expression variations. With DSNGE, we propose to decompose each facial image into an identity part and an expression part represented by their corresponding nonnegative bases. We also address the intra-class variation issue in the expression recognition problem, and further devise a graph-embedding constraint on the expression subspace to tackle this problem. Our experimental results show that the proposed DSNGE outperforms other graph-based nonnegative factorization methods and existing expression recognition methods on CK+, JAFFE, and TFEID databases. Hsin-Wen Kung, Yi-Han Tu, Chiou-Ting Hsu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Alignment-free exposure fusion of image pairsabstractThis paper presents an effective fusion technique for an exposure bracketed pair of images, which may contain motion blur from moving objects or hand trembling. Exposure fusion is a technique to represent a high dynamic scene by fusing differently exposed images. Existing exposure fusion methods either assume the input images are perfectly aligned or conduct an additional alignment step before/during the fusion process. To have a high quality result without resorting to image alignment, we propose to fuse a histogram-transformed image with the short-exposed image to preserve desired properties from the input image pair. We model the fuse problem as determining the fusing map via Markov Random Field in terms of spatial continuity and two fusion criteria. Our experiments demonstrate that our results are comparable with existing methods which include additional alignment steps. Wei-Rong Sie, Chiou-Ting Hsu |
ICIP | 2 |
| 2014 | Implicit Rank-Sparsity Decomposition: Applications to Saliency/Co-saliency DetectionabstractModern techniques rely on convex relaxation to derive tractable approximations for rank-sparsity decomposition. However, the resultant precision loss usually deteriorates the performance in real-world applications. In this paper, we focus on the topic of visual saliency detection and consider the inherent uncertainty existing in observations, which may originate from both low-rank and sparse components. We formulate the rank-sparsity model with an implicit weighting factor and show that this weighting factor characterizes the nature of visual saliency. The proposed model is generalized to solve saliency and co-saliency detection in a unified way. In addition, this model can easily incorporate center-prior or other top-down priors and can extend to multi-task learning to explore the interrelation between multiple features. Experimental results demonstrate that our method improves existing rank-sparsity decomposition, and also outperforms most state of the arts on two salient object databases. Yi-Lei Chen, Chiou-Ting Hsu |
ICPR | 2 |
| 2014 | Simultaneous Tensor Decomposition and Completion Using Factor PriorsabstractThe success of research on matrix completion is evident in a variety of real-world applications. Tensor completion, which is a high-order extension of matrix completion, has also generated a great deal of research interest in recent years. Given a tensor with incomplete entries, existing methods use either factorization or completion schemes to recover the missing parts. However, as the number of missing entries increases, factorization schemes may overfit the model because of incorrectly predefined ranks, while completion schemes may fail to interpret the model factors. In this paper, we introduce a novel concept: complete the missing entries and simultaneously capture the underlying model structure. To this end, we propose a method called simultaneous tensor decomposition and completion (STDC) that combines a rank minimization technique with Tucker model decomposition. Moreover, as the model structure is implicitly included in the Tucker model, we use factor priors, which are usually known a priori in real-world tensor objects, to characterize the underlying joint-manifold drawn from the model factors. By exploiting this auxiliary information, our method leverages two classic schemes and accurately estimates the model factors and missing entries. We conducted experiments to empirically verify the convergence of our algorithm on synthetic data and evaluate its effectiveness on various kinds of real-world data. The results demonstrate the efficacy of the proposed method and its potential usage in tensor-based applications. It also outperforms state-of-the-art methods on multilinear model analysis and visual data completion tasks. Yi-Lei Chen, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Multilinear Graph Embedding: Representation and Regularization for ImagesabstractGiven a set of images, finding a compact and discriminative representation is still a big challenge especially when multiple latent factors are hidden in the way of data generation. To represent multifactor images, although multilinear models are widely used to parameterize the data, most methods are based on high-order singular value decomposition (HOSVD), which preserves global statistics but interprets local variations inadequately. To this end, we propose a novel method, called multilinear graph embedding (MGE), as well as its kernelization MKGE to leverage the manifold learning techniques into multilinear models. Our method theoretically links the linear, nonlinear, and multilinear dimensionality reduction. We also show that the supervised MGE encodes informative image priors for image regularization, provided that an image is represented as a high-order tensor. From our experiments on face and gait recognition, the superior performance demonstrates that MGE better represents multifactor images than classic methods, including HOSVD and its variants. In addition, the significant improvement in image (or tensor) completion validates the potential of MGE for image regularization. Yi-Lei Chen, Chiou-Ting Hsu |
IEEE Trans. Image Process. | 2 |
| 2014 | Example-Based Human Motion Extrapolation and Motion Repairing Using Contour ManifoldabstractWe propose a human motion extrapolation algorithm that synthesizes new motions of a human object in a still image from a given reference motion sequence. The algorithm is implemented in two major steps: contour manifold construction and object motion synthesis. Contour manifold construction searches for low-dimensional manifolds that represent the temporal-domain deformation of the reference motion sequence. Since the derived manifolds capture the motion information of the reference sequence, the representation is more robust to variations in shape and size. With this compact representation, we can easily modify and manipulate human motions through interpolation or extrapolation in the contour manifold space. In the object motion synthesis step, the proposed algorithm generates a sequence of new shapes of the input human object in the contour manifold space and then renders the textures of those shapes to synthesize a new motion sequence. We demonstrate the efficacy of the algorithm on different types of practical applications, namely, motion extrapolation and motion repair. Nick C. Tang, Chiou-Ting Hsu, Ming-Fang Weng, Tsung-Yi Lin, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 2 |
| 2013 | A Generalized Low-Rank Appearance Model for Spatio-temporally Correlated Rain StreaksabstractIn this paper, we propose a novel low-rank appearance model for removing rain streaks. Different from previous work, our method needs neither rain pixel detection nor time-consuming dictionary learning stage. Instead, as rain streaks usually reveal similar and repeated patterns on imaging scene, we propose and generalize a low-rank model from matrix to tensor structure in order to capture the spatio-temporally correlated rain streaks. With the appearance model, we thus remove rain streaks from image/video (and also other high-order image structure) in a unified way. Our experimental results demonstrate competitive (or even better) visual quality and efficient run-time in comparison with state of the art. Yi-Lei Chen, Chiou-Ting Hsu |
ICCV | 2 |
| 2013 | What has been tampered? From a sparse manipulation perspectiveabstractExisting forensic fingerprints mostly rely on robust statistical estimates, which usually hinder accurate image tampering detection at fine-grained level. To date, people still put a big question mark behind “what has been tampered?” In this paper, we try to answer this question from a counterfeiter's perspective, devil in the details, that image tampering is usually sparsely and delicately manipulated. Thanks to recently well-established rank-sparsity incoherence, we formulate the fine-grained tampering detection as a constrained minimization problem in order to discriminate the authentic areas (sharing similar feature behaviours) from the tampered areas (inconsistently and sparsely distributed) in a forensic feature space. Our formulation could incorporate with any applicable forensic features and, unlike existing methods, needs neither statistical analysis nor model factor estimation. Our experimental results show that the proposed method successfully locates various kinds of image tampering, including copy-move forgery, resampling and recompression, at fine-grained level. Yi-Lei Chen, Chiou-Ting Hsu |
MMSP | 2 |
| 2013 | Subspace Learning for Facial Age Estimation Via Pairwise Age RankingabstractAge is one of the important biometric traits for reinforcing the identity authentication. The challenge of facial age estimation mainly comes from two difficulties: (1) the wide diversity of visual appearance existing even within the same age group and (2) the limited number of labeled face images in real cases. Motivated by previous research on human cognition, human beings can confidently rank the relative ages of facial images, we postulate that the age rank plays a more important role in the age estimation than visual appearance attributes. In this paper, we assume that the age ranks can be characterized by a set of ranking features lying on a low-dimensional space. We propose a simple and flexible subspace learning method by solving a sequence of constrained optimization problems. With our formulation, both the aging manifold, which relies on exact age labels, and the implicit age ranks are jointly embedded in the proposed subspace. In addition to supervised age estimation, our method also extends to semi-supervised age estimation via automatically approximating the age ranks of unlabeled data. Therefore, we can successfully include more available data to improve the feature discriminability. In the experiments, we adopt the support vector regression on the proposed ranking features to learn our age estimators. The results on the age estimation demonstrate that our method outperforms classic subspace learning approaches, and the semi-supervised learning successfully incorporates the age ranks from unlabeled data under different scales and sources of data set. Yi-Lei Chen, Chiou-Ting Hsu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Single image dehazing with optimal transmission map
Yi-Shuan Lai, Yi-Lei Chen, Chiou-Ting Hsu |
ICPR | 3 |
| 2012 | Dual subspace nonnegative matrix factorization for person-invariant facial expression recognition
Yi-Han Tu, Chiou-Ting Hsu |
ICPR | 2 |
| 2012 | Examplar-based object posture super-resolution using manifold learningabstractThis paper proposes a learning-based approach to increase the temporal resolutions of human motion sequences. Given a set of high resolution motion sequences, our idea is first to learn the motion tendency from this learning dataset and then synthesize new postures for the low-resolution sequence according to the learned motion tendency. We summarize the proposed framework in the following steps: (1) Each motion sequence is first projected into a low-dimension manifold space, where the local distance between postures could be better preserved. We then represent each of the projected motion sequences as a motion trajectory. (2) Next, motion priors learned from the HR training sequences are used to reconstruct the motion trajectory for the input sequence. (3) Finally, we use the reconstructed motion trajectory combined with object inpainting technique to generate the final result. Our experimental results demonstrate the effectiveness of the proposed method, and also show its outperformance over existing approaches. Chih-Hung Ling, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
MMSP | 3 |
| 2012 | Ridge Network Detection in Crumpled Paper via Graph Density MaximizationabstractCrumpled sheets of paper tend to exhibit a specific and complex structure, which is described by physicists as ridge networks. Existing literature shows that the automation of ridge network detection in crumpled paper is very challenging because of its complex structure and measuring distortion. In this paper, we propose to model the ridge network as a weighted graph and formulate the ridge network detection as an optimization problem in terms of the graph density. First, we detect a set of graph nodes and then determine the edge weight between each pair of nodes to construct a complete graph. Next, we define a graph density criterion and formulate the detection problem to determine a subgraph with maximal graph density. Further, we also propose to refine the graph density by including a pairwise connectivity into the criterion to improve the connectivity of the detected ridge network. Our experimental results show that, with the density criterion, our proposed method effectively automates the ridge network detection. Chiou-Ting Hsu, Marvin Huang |
IEEE Trans. Image Process. | 1 |
| 2011 | Single-frame-based rain removal via image decompositionabstractRain removal from a video is a challenging problem and has been recently investigated extensively. Nevertheless, the problem of rain removal from a single image has been rarely studied in the literature, where no temporal information among successive images can be exploited, making it more challenging. In this paper, to the best of our knowledge, we are among the first to propose a single-frame-based rain removal framework via properly formulating rain removal as an image decomposition problem based on morphological component analysis (MCA). Instead of directly applying conventional image decomposition technique, we first decompose an image into the low-frequency and high frequency parts using a bilateral filter. The high-frequency part is then decomposed into "rain component" and "non rain component" via performing dictionary learning and sparse coding. As a result, the rain component can be successfully removed from the image while preserving most original image details. Experimental results demonstrate the efficacy of the proposed algorithm. Yu-Hsiang Fu, Li-Wei Kang, Chia-Wen Lin, Chiou-Ting Hsu |
ICASSP | 4 |
| 2011 | Time-variant modeling for general surface appearanceabstractDescribing time-variant appearance of object surface is still an open problem. With intricate environmental factors and different material characteristics over time, no researcher did tackle the principal problem: how to formulate time-variant change on general surface appearance? In this paper, we attempt to solve this challenging issue. Using multilinear algebra representation, we propose a novel appearance model and characterize the surface-specific and time-variant properties. When given an unknown sample, we propose a robust method to estimate its aging degree. In addition, we also propose an approach to synthesize its realistic appearance changes even when the material of this given sample does not exist in our database. Experimental results demonstrate the feasibility and effectiveness of our proposed approach. To the best of our knowledge, this challenging issue is first explored in image processing applications. Yi-Lei Chen, Chiou-Ting Hsu |
ICIP | 2 |
| 2011 | Example-based human motion extrapolation based on manifold learningabstractIn this paper, we propose a new framework to synthesize human motions based on only one single posture given in the input image. To generate visually pleasing motion sequences, the proposed framework consists of two key techniques. One is motion retrieval, which retrieves reference motions from a human motion database on a low-dimensional motion manifold. Another one is human motion extrapolation, which first generates new postures by deforming the shape of the input posture according to the retrieved motions and then synthesizes the corresponding motion sequence. To demonstrate the efficacy of the proposed method, we generate several human motion sequences using input images with different postures and show that the results are indeed visually pleasing. Nick C. Tang, Chiou-Ting Hsu, Tsung-Yi Lin, Hong-Yuan Mark Liao |
ACM Multimedia | 2 |
| 2011 | Narrative Generation by Repurposing Digital Videos
Nick C. Tang, Hsiao-Rong Tyan, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
MMM (1) | 3 |
| 2011 | Automatic ridge network detection in crumpled paper based on graph densityabstractCrumpled sheets of paper tend to exhibit specific and complex structure, which is usually described as ridge network by physicists. Existing literature has showed that it is difficult to automate ridge network detection in crumpled paper because of its complex structure. In this paper, we attempt to develop an automatic detection process in terms of our proposed density criterion. We model the ridge network as a weighted graph, where the nodes indicate the intersections of ridges and the edges are the straightened ridges detected in crumpled paper. We construct the weighted graph by first detecting the nodes and then determining the edge weight using the ridge responses. Next, we formulate a graph density criterion to evaluate the detected ridge network. Finally, we propose an edge linking method to construct the graph by maximizing the proposed density criterion. Our experimental results show that, with the density criterion, our proposed node detection together with the edge line linking method could effectively automate the ridge network detection. Marvin Huang, Chiou-Ting Hsu, Kazuyuki Tanaka |
MMSP | 2 |
| 2011 | Video Inpainting on Digitized Vintage Films via Maintaining Spatiotemporal ContinuityabstractVideo inpainting is an important video enhancement technique used to facilitate the repair or editing of digital videos. It has been employed worldwide to transform cultural artifacts such as vintage videos/films into digital formats. However, the quality of such videos is usually very poor and often contain unstable luminance and damaged content. In this paper, we propose a video inpainting algorithm for repairing damaged content in digitized vintage films, focusing on maintaining good spatiotemporal continuity. The proposed algorithm utilizes two key techniques. Motion completion recovers missing motion information in damaged areas to maintain good temporal continuity. Frame completion repairs damaged frames to produce a visually pleasing video with good spatial continuity and stabilized luminance. We demonstrate the efficacy of the algorithm on different types of video clips. Nick C. Tang, Chiou-Ting Hsu, Chih-Wen Su, Timothy K. Shih, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 2 |
| 2010 | Color transfer for complex content images based on intrinsic componentabstractThis paper proposes an automatic color transfer method for processing images with complex content based on intrinsic component. Although several automatic color transfer methods has been proposed by including region information and/or using multiple references, these methods tend to become ineffective when processing images with complex content and lighting variation. In this paper, our goal is to incorporate the idea of intrinsic component to better characterize the local organization within an image and to reduce the color-bleeding artifact across complex regions. Using intrinsic information, we first represent each image in region level and determine the best-matched reference region for each target region. Next, we conduct color transfer between the best-matched region pairs and perform weighted color transfer for pixels across complex regions in a de-correlated color space. Both subjective and objective evaluation of our experiments demonstrates that the proposed method outperforms the existing methods. Wan-Chien Chiou, Yi-Lei Chen, Chiou-Ting Hsu |
MMSP | 3 |
| 2010 | Face hallucination using Bayesian global estimation and local basis selectionabstractThis paper proposes a two-step prototype-face-based scheme of hallucinating the high-resolution detail of a low-resolution input face image. The proposed scheme is mainly composed of two steps: the global estimation step and the local facial-parts refinement step. In the global estimation step, the initial high-resolution face image is hallucinated via a linear combination of the global prototype faces with a coefficient vector. Instead of estimating coefficient vector in the high-dimensional raw image domain, we propose a maximum a posteriori (MAP) estimator to estimate the optimum set of coefficients in the low-dimensional coefficient domain. In the local refinement step, the facial parts (i.e., eyes, nose and mouth) are further refined using a basis selection method based on overcomplete nonnegative matrix factorization (ONMF). Experimental results demonstrate that the proposed method can achieve significant subjective and objective improvement over state-of-the-art face hallucination methods, especially when an input face does not belong to a person in the training data set. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao, Jen-Yu Yu |
MMSP | 3 |
| 2009 | Region-based color transfer from multi-reference with graph-theoretic region correspondence estimationabstractThis paper proposes an automatic color transfer method based on multi-reference and graph-theoretic region correspondence estimation. When multiple high-quality reference images are available, our goal is to determine a set of best reference colors for transferring the color characteristics of the target image. Given a target image, we first employ content-based image retrieval technique to obtain a small number of relevant images as its multi-reference. Next, we represent each image in region level and determine the best-matched reference region for each target region. We propose to incorporate both region attribute and spatially adjacent relationships between regions into the region mapping criterion. Finally, we conduct color transfer between the best-matched region pairs in a de-correlated color space. Both subjective and objective measures of our experiments demonstrate that the proposed method outperforms the existing methods. Wan-Chien Chiou, Chiou-Ting Hsu |
ICIP | 2 |
| 2009 | Cooperative face hallucination using multiple referencesabstractThis paper proposes a cooperative example-based face hallucination method using multiple references. The proposed method first uses clustering and residual prototype faces construction to improve the performance of hallucinating a single low-resolution (LR) face to obtain a high-resolution (HR) counterpart. In the case that multiple LR face images for a person are available, a unique feature of the proposed method is to cooperatively enhance the qualities of hallucinated HR images by taking into account the multiple input face images jointly as prior models. Experimental results demonstrate that the proposed cooperative method achieve significant subjective and objective improvement over single-prior schemes. Chih-Chung Hsu, Chia-Wen Lin, Chiou-Ting Hsu, Hong-Yuan Mark Liao |
ICME | 3 |
| 2009 | A novel color-context descriptor and its applicationsabstractThis paper presents a new descriptor for object categorization and pedestrian identification applications. One of the main drawbacks of shape-context descriptor is its vulnerability and distinctness to color images. We propose a spherical descriptor that simultaneously adopts the spatial and color information as a discriminative representation. Based on the descriptor, this paper also contributes a bag-of-features framework to pedestrian identification for video surveillance. In contrast to the previous works, the proposed scheme does not require background subtraction stage. Thus the potential problems, such as the susceptibility to shadows and highlights from the background subtraction procedure, are avoided. Experiments validate the discriminant power of the proposed descriptor in object categorization on COIL- 100 database and pedestrian identification in surveillance videos. Chia-Te Liao, Yu-Lin Wang, Shang-Hong Lai, Chiou-Ting Hsu |
ICME | 4 |
| 2009 | Bayesian age estimation on face imagesabstractThis paper proposes to formulate the age estimation on face images as a Bayesian estimation problem. The proposed framework incorporates a probabilistic model on face aging distribution into the formulation. Individual-specific prior information as well as individual-specific aging models, if available, could then be easily included into the unified age estimation framework. We conduct experiments on the publicly available FG-NET database and compare the estimation results with existing methods. The experimental results demonstrate that, by combining the probabilistic modeling of facial feature distribution, the proposed method indeed outperforms the existing methods. Chung-Chun Wang, Yi-Chueh Su, Chiou-Ting Hsu, Chia-Wen Lin, Hong-Yuan Mark Liao |
ICME | 3 |
| 2009 | Detecting doubly compressed images based on quantization noise model and image restorationabstractSince JPEG has been a popularly used image compression standard, forgery detection in JPEG images now plays an important role. Forgeries on compressed images often involve recompression and tend to erase those forgery traces existed in uncompressed images. We could, however, try to discover new traces caused by recompression and use these traces to detect the recompression forgeries. Quantization is the critical step in lossy compression which maps the DCT coefficients in an irreversible way under the quantization constraint set (QCS) theorem. In this paper, we first derive that a doubly compressed image no longer follows the QCS theorem and then propose a novel quantization noise model to characterize single and doubly compressed images. In order to detect double compression forgery, we further propose to approximate the uncompressed ground truth image using image restoration techniques. We conduct a series of experiments to demonstrate the validity of the proposed quantization noise model and also the effectiveness of the forgery detection method with the proposed image restoration techniques. Yi-Lei Chen, Chiou-Ting Hsu |
MMSP | 2 |
| 2008 | Image tampering detection by blocking periodicity analysis in JPEG compressed imagesabstractSince JPEG image format has been a popularly used image compression standard, tampering detection in JPEG images now plays an important role. The artifacts introduced by lossy JPEG compression can be seen as an inherent signature for compressed images. In this paper, we propose a new approach to analyse the blocking periodicity by, 1) developing a linearly dependency model of pixel differences, 2) constructing a probability map of each pixel’s belonging to this model, and 3) finally extracting a peak window from the Fourier spectrum of the probability map. We will show that, for single and double compressed images, their peaks’ energy distribution behave very differently. We exploit this property and derive statistic features from peak windows to classify whether an image has been tampered by cropping and recompression. Experimental results demonstrate the validity of the proposed approach. Yi-Lei Chen, Chiou-Ting Hsu |
MMSP | 2 |
| 2008 | Video forgery detection using correlation of noise residueabstractWe propose a new approach for locating forged regions in a video using correlation of noise residue. In our method, block-level correlation values of noise residual are extracted as a feature for classification. We model the distribution of correlation of temporal noise residue in a forged video as a Gaussian mixture model (GMM). We propose a two-step scheme to estimate the model parameters. Consequently, a Bayesian classifier is used to find the optimal threshold value based on the estimated parameters. Two video inpainting schemes are used to simulate two different types of forgery processes for performance evaluation. Simulation results show that our method achieves promising accuracy in video forgery detection. Chih-Chung Hsu, Tzu-Yi Hung, Chia-Wen Lin, Chiou-Ting Hsu |
MMSP | 4 |
| 2008 | Image Retrieval With Relevance Feedback Based on Graph-Theoretic Region Correspondence EstimationabstractThis paper presents a graph-theoretic approach for interactive region-based image retrieval. When dealing with image matching problems, we use graphs to represent images, transform the region correspondence estimation problem into an inexact graph matching problem, and propose an optimization technique to derive the solution. We then define the image distance in terms of the estimated region correspondence. In the relevance feedback steps, with the estimated region correspondence, we propose to use a maximum likelihood method to re-estimate the ideal query and the image distance measurement. Experimental results show that the proposed graph-theoretic image matching criterion outperforms the other methods incorporating no spatially adjacent relationship within images. Furthermore, our maximum likelihood method combined with the estimated region correspondence improves the retrieval performance in feedback steps. Chueh-Yu Li, Chiou-Ting Hsu |
IEEE Trans. Multim. | 2 |
| 2007 | Online Selection of Tracking Features using AdaBoostabstractThis paper, a novel feature selection algorithm for object tracking is proposed. This algorithm performs more robust than the previous works by taking the correlation between features into consideration. Pixels of object/background regions are first treated as training samples. The feature selection problem is then modeled as finding a good subset of features and constructing a compound likelihood image with better discriminability for the tracking process. By adopting the AdaBoost algorithm, we iteratively select one best feature which compensate the previous selected features and linearly combine the set of corresponding likelihood images to obtain the compound likelihood image. We include the proposed algorithm into the mean shift based tracking system. Experimental results demonstrate that the proposed algorithm achieve very promising results. Ying-Jia Yeh, Chiou-Ting Hsu |
ICCCN | 2 |
| 2007 | Source Camera Identification Based on Camera Gain HistogramabstractIn this paper, we propose a novel approach for source camera identification based on camera gain histogram. By using the photon transfer curve (PTC) as camera noise model, we construct camera gain histogram from the occurrences of different camera gain constants. With the distribution of camera gain histogram for each camera, we extract four features to characterize the camera. In our experiments, 400 photos acquired from two high-end digital cameras at two different exposure levels are used to evaluate the effectiveness of the proposed approach. A two-class support vector machine (SVM) is employed as a classifier. Our experimental results demonstrate that the distinction rate in identifying different cameras achieves promising performance. Sz-Han Chen, Chiou-Ting Hsu |
ICIP (4) | 2 |
| 2006 | Background Modeling from GMM Likelihood Combined with Spatial and Color CoherencyabstractThis paper proposes to combine spatial and color coherency with the pixel-wise GMM to determine the background model. We first represent each pixel with a hybrid feature vector, which includes its GMM likelihood, color and spatial features, and estimate the density for each video frame by a non-parametric method. Next, we apply a clustering process to segment the video frame into clusters with similar hybrid features. Finally, we replace the background likelihood for each cluster with the GMM likelihood in the cluster mode. Hence, the resulting background model becomes a smoothed GMM in terms of spatial and color coherency. For accurate object detection, we develop an adaptive thresholding scheme using Markov random field. Moreover, in order to reduce the computational load, we also propose a filtering step to skip pixels from the time-consuming clustering process. Our experimental results and comparisons demonstrate that the proposed background model indeed achieves better detection results with accurate object contours even in dynamic scenes. Sheng-Yan Yang, Chiou-Ting Hsu |
ICIP | 2 |
| 2006 | Region tracking for non-rigid video objects in a non-parametric MAP framework
Chiou-Ting Hsu, Ming-Shen Hsieh |
Signal Process. Image Commun. | 1 |
| 2006 | Fusion of audio and motion information on HMM-based highlight extraction for baseball gamesabstractThis paper aims to extract baseball game highlights based on audio-motion integrated cues. In order to better describe different audio and motion characteristics in baseball game highlights, we propose a novel representation method based on likelihood models. The proposed likelihood models measure the "likeliness" of low-level audio features and motion features to a set of predefined audio types and motion categories, respectively. Our experiments show that using the proposed likelihood representation is more robust than using low-level audio/motion features to extract the highlight. With the proposed likelihood models, we then construct an integrated feature representation by symmetrically fusing the audio and motion likelihood models. Finally, we employ a hidden Markov model (HMM) to model and detect the transition of the integrated representation for highlight segments. A series of experiments have been conducted on a 12-h video database to demonstrate the effectiveness of our proposed method and show that the proposed framework achieves promising results over a variety of baseball game sequences. Chih-Chieh Cheng, Chiou-Ting Hsu |
IEEE Trans. Multim. | 2 |
| 2005 | Learning hidden semantic cues using support vector clusteringabstractThis paper presents a method to infer hidden semantic cues by accumulating the knowledge learned from relevance feedback sessions. We propose to explicitly represent a semantic space using a probabilistic model. In short-term learning, we apply the general 2-class SVM classification to initialize the semantic space. Once the accumulated semantic space becomes impractically large, we propose using support vector clustering (SVC) to construct a more compact and still meaningful semantic space with lower dimensionality. Given a dimension-reduced semantic space, we then perform the image query in terms of the semantic attributes instead of merely the visual features. Our experimental results and comparisons demonstrate that the proposed semantic representation as well as the SVC-based technique indeed achieves promising results. Jia-Wen Tung, Chiou-Ting Hsu |
ICIP (1) | 2 |
| 2005 | Soft Region Correspondence Estimation for Graph-Theoretic Image Retrieval Using Quadratic Programming ApproachabstractThis paper proposes employing a graph-theoretic approach to estimate the region correspondence between two images. We represent each image as an attributed undirected graph and transform the image matching problem into an inexact graph matching problem. We formulate the estimation of the soft matching matrix between two graphs as a quadratic programming problem, and apply KKT (Karush-Kuhn-Tucker) conditions and the modified simplex algorithm to solve the constrained optimization problem. With the soft matching matrix, we are capable to integrate both the region correspondence and low-level visual features into an effective matching measurement for image matching. Experiments have been conducted on image retrieval to show the effectiveness of the proposed estimation algorithm. Chueh-Yu Li, Chiou-Ting Hsu |
ICME | 2 |
| 2004 | Region correspondence for image retrieval using graph-theoretic approach and maximum likelihood estimationabstractThis paper proposes employing an efficient graph-theoretic approach to estimate the region correspondence between two images. We represent each image as an attributed graph and transform the image matching problem into a graph matching problem. During the image retrieval process, we formulate the matching problem as a maximum likelihood estimation and propose an optimization technique to derive its closed-form solution. Hence, we are capable of measuring the image distance in terms of both the estimated region correspondence and the low-level features. This paper has two main contributions. First, our proposed matching technique is efficient and applicable to the interactive process of image retrieval. Second, with the estimated region correspondence, the proposed matching criterion, which is defined in terms of matched regions and penalized with unmatched regions, achieves satisfactory performance for retrieval applications. Chueh-Yu Li, Chiou-Ting Hsu |
ICIP | 2 |
| 2004 | Statistical motion characterization for video content classificationabstractThis work proposed using a unified model to characterize the motion variations along both the spatial and temporal domains. To this end, we estimate the motion quantities from the pixelwise normal flow and represent the motion distribution using two Gibbs models: temporal and spatial Gibbs models. We measure the potential values of the two Gibbs models by the maximum likelihood criterion. To demonstrate the effectiveness of the proposed model, we have applied the motion model for the application of video content classification. Experimental results show that using the proposed model indeed improves the classification performance. Chiou-Ting Hsu, Ching-Wei Lee |
ICME | 1 |
| 2004 | Mosaics of video sequences with moving objects
Chiou-Ting Hsu, Yu-Chun Tsan |
Signal Process. Image Commun. | 1 |
| 2003 | Segmentation of nonrigid object in a nonparametric MAP frameworkabstractThis paper presents an efficient segmentation approach for nonrigid video object. We propose to formulate the video object segmentation problem as the maximum a posteriori probability (MAP) problem and define the probabilistic models in terms of the object's density function. Furthermore, in order to accurately represent the density function for video object with arbitrary shape and complex texture, we employ a nonparametric method to estimate the density function. Our proposed density estimation mostly relies on the object's color features and requires no time-consuming motion estimation. In addition, we further employ an efficient mean-shift procedure in the MAP optimization step to largely reduce the computational cost. Our experiments demonstrate that the segmentation results are very promising even when the video objects are severely deformed or occluded. Chiou-Ting Hsu, Ming-Shen Hsieh |
ICIP (1) | 1 |
| 2002 | Motion trajectory based video indexing and retrievalabstractThis paper presents a technique to efficiently index and retrieve video clips in terms of motion-trajectory-based similarity. We describe the motion trajectory in three representations: the horizontal and vertical movements of the trajectory, and the motion trail that indicates shape of the trajectory. Each representation is approximated by a polynomial function. We index the polynomial coefficients and combine different spatio-temporal characteristics to provide flexible retrieval processes. A unified framework is also proposed to deal with various query types: query-by-example, query-by-sketch, and query-by-specification. Chiou-Ting Hsu, Shang-Ju Teng |
ICIP (1) | 1 |
| 2001 | Mosaics of video sequences with moving objectsabstractWe propose an efficient method for creating a mosaic of a video sequence in the presence of moving objects. This method includes two principal processes. The first one removes the moving objects from the background, and as a side effect, obtains the global motion. This global motion provides a good initial estimation to the next stage. Second, we employ a feature-based technique to derive the precise global motion with eight projective parameters. The performance of our work is demonstrated by experiments. Chiou-Ting Hsu, Yu-Chun Tsan |
ICIP (2) | 1 |
| 2000 | Feature-Based Video MosaicabstractA complete method to create a panoramic video mosaic is proposed in this paper. This method includes three principal processes: real-time frame selection, feature based estimation of an eight-parameter projective coordinate transformation, and image mosaic composing. This proposed method works well on many video sequences captured from a PC camera to generate both 360 degree and planar panoramas. Experimental results demonstrate the validity and superior quality of this proposed method. Chiou-Ting Hsu, Tzu-Hung Cheng, Rob A. Beuker, Jyh-Kuen Horng |
ICIP | 1 |
| 2000 | Multiresolution feature-based image registration
Chiou-Ting Hsu, Rob A. Beuker |
VCIP | 1 |
| 1999 | Hidden digital watermarks in imagesabstractIn this paper, an image authentication technique by embedding digital "watermarks" into images is proposed. Watermarking is a technique for labeling digital pictures by hiding secret information into the images. Sophisticated watermark embedding is a potential method to discourage unauthorized copying or attest the origin of the images. In our approach, we embed the watermarks with visually recognizable patterns into the images by selectively modifying the middle-frequency parts of the image. Several variations of the proposed method are addressed. The experimental results show that the proposed technique successfully survives image processing operations, image cropping, and the Joint Photographic Experts Group (JPEG) lossy compression. Chiou-Ting Hsu, Ja-Ling Wu |
IEEE Trans. Image Process. | 1 |
| 1996 | Hidden signatures in imagesabstractAn image authentication technique by embedding each image with a signature so as to discourage unauthorized copying is proposed. The proposed technique could actually survive several kinds of image processing and the JPEG lossy compression. Chiou-Ting Hsu, Ja-Ling Wu |
ICIP (3) | 1 |
| 1996 | Multiresolution mosaicabstractMosaic techniques have been used to combine two or more signals into a new one with an invisible seam, and with as little distortion of each signal as possible. Multiresolution representation is an effective method for analyzing the information content of signals, and it also fits a wide spectrum of visual signal processing and visual communication applications. The wavelet transform is one kind of multiresolution representations, and has found a wide variety of application in many aspects, including signal analysis, image coding, image processing, computer vision and etc. Due to its characteristic of multiresolution signal decomposition, the wavelet transform is used for the image mosaic by choosing the width of the mosaic transition zone proportional to the frequency represented in the band. Both 1-D and 2-D signal mosaics are described, and some factors which affect the mosaics are discussed. Chiou-Ting Hsu, Ja-Ling Wu |
ICIP (3) | 1 |