Shang-Hong Lai

dblp:27/679 · DBLP profile ↗
← Back
193ranked-venue papers
14as first author
39since 2021 · last 2026
0000-0002-5092-993XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 161 · 3 first-author · 35 since 2021Artificial intelligence and machine learning · 75 · 12 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 STEN-FAWA: Spatial-Temporal Expert Network with Forgery-Aware Weighted-Adaptive Aggregator for Deepfake Detection
Min-Xuan Qiu, Kai-Wei Chuang, Shang-Hong Lai
FG3
2026 Beyond Texture: Advanced Facial Privacy Protection via Hierarchical Diffusion Autoencoder
Ting-Yi Lu, Che-Tsung Lin, Christopher Zach, Shang-Hong Lai
ICPR (7)4
2026 FaceML-MOE: Face Multi-task Learning via Attribute-Specific Expert Routing
Lu-Yan Wang, Shang-Hong Lai
ICPR (9)2
2026 VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
abstract
Video anomaly understanding (VAU) aims to provide detailed interpretation and semantic comprehension of anomalous events within videos, addressing limitations of traditional methods that focus solely on detecting and localizing anomalies. However, existing approaches often neglect the deeper causal relationships and interactions between objects, which are critical for understanding anomalous behaviors. In this paper, we propose VADER, an LLM-driven framework for Video Anomaly unDErstanding, which integrates keyframe object Relation features with visual cues to enhance anomaly comprehension from video. Specifically, VADER first applies an Anomaly Scorer to assign per-frame anomaly scores, followed by a Context-AwarE Sampling (CAES) strategy to capture the causal context of each anomalous event. A Relation Feature Extractor and a COntrastive Relation Encoder (CORE) jointly model dynamic object interactions, producing compact relational representations for downstream reasoning. These visual and relational cues are integrated with LLMs to generate detailed, causally grounded descriptions and support robust anomaly-related question answering. Experiments on multiple real-world VAU benchmarks demonstrate that VADER achieves strong results across anomaly description, explanation, and causal reasoning tasks, advancing the frontier of explainable video anomaly analysis. Project page is available at https://vader-vau.github.io/.
Yu-Ho Lin, Min-Hung Chen, Fu-En Yang, Shang-Hong Lai
WACV5
2026 2S-CEDiff: A Two-Stage Diffusion Framework for Generating High-Fidelity Contrast-Enhanced CT Images from Non-Contrast Scans
abstract
Contrast-enhanced Computed Tomography (CT) plays a vital role in modern medical diagnostics, particularly in cardiovascular assessment. However, the use of intravenous contrast agents can pose potential health risks for vulnerable patient populations. To address this limitation, we present 2S-CEDiff, a clinically-inspired image translation framework that synthesizes high-fidelity contrast-enhanced 3D CT volumes from non-contrast inputs. Our method adopts a two-stage architecture. The first stage employs a 2.5D diffusion model to generate anatomically accurate slice-wise predictions, guided by cross-attention-based positional conditioning and structural priors derived from segmentation masks produced by a pre-trained TotalSegmentator model. The second stage involves applying a 3D UNet model trained with a multi-objective loss that jointly optimizes pixel-level fidelity and volumetric coherence. We evaluate 2S-CEDiff on an in-house paired 3D cardiac CT dataset, where it achieves state-of-the-art performance in PSNR (29.87), SSIM (0.89), and inter-slice coherence. Moreover, the synthesized contrast-enhanced images significantly enhance downstream anatomical segmentation accuracy, improving the overall Dice score to 0.9447 and yielding a substantial +0.0798 increase in myocardium segmentation performance with TotalSegmentator — underscoring their clinical utility and translational potential.
Yibang Wu, Tzung-Dau Wang, Shang-Hong Lai
WACV3
2025 KeyGS: A Keyframe-Centric Gaussian Splatting Method for Monocular Image Sequences
abstract
Reconstructing high-quality 3D models from sparse 2D images has garnered significant attention in computer vision. Recently, 3D Gaussian Splatting (3DGS) has gained prominence due to its explicit representation with efficient training speed and real-time rendering capabilities. However, existing methods still heavily depend on accurate camera poses for reconstruction. Although some recent approaches attempt to train 3DGS models without the Structure-from-Motion (SfM) preprocessing from monocular video datasets, these methods suffer from prolonged training times, making them impractical for many applications. In this paper, we present an efficient framework that operates without any depth or matching model. Our approach initially uses SfM to quickly obtain rough camera poses within seconds, and then refines these poses by leveraging the dense representation in 3DGS. This framework effectively addresses the issue of long training times. Additionally, we integrate the densification process with joint refinement and propose a coarse-to-fine frequency-aware densification to reconstruct different levels of details. This approach prevents camera pose estimation from being trapped in local minima or drifting due to high-frequency signals. Our method significantly reduces training time from hours to minutes while achieving more accurate novel view synthesis and camera pose estimation compared to previous methods.
Keng Wei Chang, Shang-Hong Lai
AAAI3
2025 MovieCORE: COgnitive REasoning in Movies
abstract
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Ying Cheng, Hung-Ting Su, Yung-Hao Tang, Shang-Hong Lai, Winston H. Hsu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Hung-Ting Su, Yung-Hao Tang, Shang-Hong Lai, Winston H. Hsu
EMNLP7
2025 InstAD: Instance-aware Segmentation Framework for Zero-shot Multi-instance Anomaly Detection
abstract
Multi-instance anomaly detection and segmentation play a crucial role in automated industrial inspection. Previous works mainly focus on single-instance detection tasks that require well-aligned input and extensive training sets. In this work, we introduce InstAD, a zero-shot multi-instance anomaly detection framework that achieves high accuracy with unaligned multi-instances. Combining segmentation results from the Segment Anything Model [1] and Grounded-SAM [2], we further refine the segmentation results with the proposed Adaptive Bandwidth Segmentation Refinement scheme to achieve accurate multi-instance segmentation. By using our instance-aware anomaly detection strategy, we achieve 93.0% and 99.1% for image-level and pixel-level AUROC, respectively, with the four-shot setting on the multi-instance classes of the VisA and MPDD datasets. Under the zero-shot scenario, we reach 87.2% and 98.4% for image-level and pixel-level AUROC, respectively. Our experiments on public datasets show that the proposed InstAD method significantly outperforms SOTA methods for multi-instance anomaly detection and segmentation.
Cheng-Yu Ho, Shang-Hong Lai
ICASSP2
2025 HERMES: Temporal-Coherent Long-form Understanding with Episodes and Semantics
abstract
Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processing redundant information efficiently, and extracting high-level semantic concepts. To address these challenges, we propose a novel approach that more accurately reflects human cognition. This paper introduces HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics, featuring two versatile modules that can enhance existing video-language models or operate as a standalone system. Our Episodic COmpressor (ECO) efficiently aggregates representations from micro to semi-macro levels, reducing computational overhead while preserving temporal dependencies. Our Semantics ReTRiever (SeTR) enriches these representations with semantic information by focusing on broader context, dramatically reducing feature dimensionality while preserving relevant macro-level information. We demonstrate that these modules can be seamlessly integrated into existing SOTA models, consistently improving their performance while reducing inference latency by up to 43% and memory usage by 46%. As a standalone system, HERMES achieves state-of-the-art performance across multiple long-video understanding benchmarks in both zero-shot and fully-supervised settings.
Gueter Josmy Faure, Jia-Fong Yeh, Min-Hung Chen, Hung-Ting Su, Shang-Hong Lai, Winston H. Hsu
ICCV5
2025 CLIP-FSQAE: Clip-Guided Finite Scalar Quantized Autoencoder for Few-Shot Anomaly Detection
abstract
Industrial anomaly detection is hindered by data inefficiency and dependence on large-scale training sets. We introduce CLIP-FSQAE, a novel framework for few-shot anomaly detection that integrates Finite Scalar Quantization (FSQ) into a CLIP-guided autoencoder. By leveraging CLIP’s vision-language alignment and image-to-text tuning, the model mitigates shortcut learning and reconstructs defective inputs using only a few normal samples. A diffusion-based synthetic anomaly generator further improves robustness and diversity. Evaluations on MVTecAD and VisA show that CLIP-FSQAE achieves state-of-the-art image- and pixel-level performance. It matches or exceeds full-data baselines with fewer than four training samples, establishing a new benchmark for few-shot industrial anomaly detection.
Shih-Chih Lin, Shang-Hong Lai
ICIP2
2025 CACE: Sim-to-Real Indoor 3D Semantic Segmentation via Context-Aware Augmentation and Consistency Enforcement
abstract
Indoor 3D domain adaptation for semantic segmentation is an understudied task. The first unsupervised sim-to-real benchmark was only proposed recently. Existing methods try to modify the source domain data by simulating the occlusion and noise pattern of the target domain. However, this methodology unrealistically demands a clear definition of the real-world data patterns, and is highly dependent on the simulation quality. In this paper, we propose a novel adaptation framework via Context-aware Augmentation and Consistency Enforcement (CACE). Our CACE framework consists of two modules, a space and context-aware augmentation module that is invariant of target data pattern and domain gaps, and a carefully designed self-supervision module that maximizes the utility of the augmented data. Our CACE surpasses the state-of-the-art method by over 6% on the indoor 3D sim-to-real benchmark$3D-FRONT\rightarrow ScanNet$.
Tsung-Yu Chen, Luyu Yang, Tzu-Yu Chuang, Shang-Hong Lai
WACV4
2025 Spatio-Temporal Context Prompting for Zero-Shot Action Detection
abstract
Spatio-temporal action detection encompasses the tasks of localizing and classifying individual actions within a video. Recent works aim to enhance this process by incorporating interaction modeling, which captures the relationship between people and their surrounding context. However, these approaches have primarily focused on fully-supervised learning, and the current limitation lies in the lack of generalization capability to recognize unseen action categories. In this paper, we aim to adapt the pretrained image-language models to detect unseen actions. To this end, we propose a method which can effectively leverage the rich knowledge of visual-language models to perform Person-Context Interaction. Meanwhile, our Context Prompting module will utilize contextual information to prompt labels, thereby enhancing the generation of more representative text features. Moreover, to address the challenge of recognizing distinct actions by multiple people at the same timestamp, we design the Interest Token Spotting mechanism which employs pretrained visual knowledge to find each person's interest context tokens, and then these tokens will be used for prompting to generate text features tailored to each individual. To evaluate the ability to detect unseen actions, we propose a comprehensive benchmark on J-HMDB, UCF101-24, and AVA datasets. The experiments show that our method achieves superior results compared to previous approaches and can be further extended to multi-action videos, bringing it closer to real-world applications. The code and data can be found in ST-CLIP.
Wei-Jhe Huang, Min-Hung Chen, Shang-Hong Lai
WACV3
2025 Text in the dark: Extremely low-light text image enhancement
Che-Tsung Lin, Chun Chet Ng, Zhi Qin Tan, Wan Jun Nah, Xinyu Wang 0010, Jie-Long Kew, Po-Hao Hsu, Shang-Hong Lai, Chee Seng Chan, Christopher Zach
Signal Process. Image Commun.8
2024 CSAD: Unsupervised Component Segmentation for Logical Anomaly Detection
Yu Hsuan Hsieh, Shang-Hong Lai
BMVC2
2024 APTPose: Anatomy-aware Pre-Training for 3D Human Pose Estimation
Qing-Wen Yang, Kai-Wen Duan, Ting-Yi Lu, Cheng-Yen Yang, Jenq-Neng Hwang, Shang-Hong Lai
BMVC8
2024 Domain Adaptation for Machinery Fault Diagnosis Based on Critic Classifier GAN
Tso-Sung Hung, Shang-Hong Lai
ICPR (6)2
2024 Few-Shot Deep Structure-Based Camera Localization with Pose Augmentation
Cheng-Yu Tsai, Shang-Hong Lai
ICPR (18)2
2023 MixFairFace: Towards Ultimate Fairness via MixFair Adapter in Face Recognition
abstract
Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than the others. In this paper, we propose MixFairFace framework to improve the fairness in face recognition models. First of all, we argue that the commonly used attribute-based fairness metric is not appropriate for face recognition. A face recognition system can only be considered fair while every person has a close performance. Hence, we propose a new evaluation protocol to fairly evaluate the fairness performance of different approaches. Different from previous approaches that require sensitive attribute labels such as race and gender for reducing the demographic bias, we aim at addressing the identity bias in face representation, i.e., the performance inconsistency between different identities, without the need for sensitive attribute labels. To this end, we propose MixFair Adapter to determine and reduce the identity bias of training samples. Our extensive experiments demonstrate that our MixFairFace approach achieves state-of-the-art fairness performance on all benchmark datasets.
Fu-En Wang, Chien-Yi Wang, Min Sun 0001, Shang-Hong Lai
AAAI4
2023 Masked Attention ConvNeXt Unet with Multi-Synthesis Dynamic Weighting for Anomaly Detection and Localization
Shih-Chih Lin, Ho-Weng Lee, Yu-Shuan Hsieh, Cheng Yu Ho, Shang-Hong Lai
BMVC5
2023 KFC: Kinship Verification with Fair Contrastive loss and Multi-Task Learning
Jia Luo Peng, Keng Wei Chang, Shang-Hong Lai
BMVC3
2023 Generalized Face Anti-Spoofing via Multi-Task Learning and One-Side Meta Triplet Loss
abstract
With the increasing variations of face presentation attacks, model generalization becomes an essential challenge for a practical face anti-spoofing system. This paper presents a generalized face anti-spoofing framework that consists of three tasks: depth estimation, face parsing, and live/spoof classification. With the pixel-wise supervision from the face parsing and depth estimation tasks, the regularized features can better distinguish spoof faces. While simulating domain shift with meta-learning techniques, the proposed one-side triplet loss can further improve the generalization capability by a large margin. Extensive experiments on four public datasets demonstrate that the proposed framework and training strategies are more effective than previous works for model generalization to unseen domains.
Chu-Chun Chuang, Chien-Yi Wang, Shang-Hong Lai
FG3
2023 Robust Multi-Object Tracking With Spatial Uncertainty
abstract
Most methods address the multi-object tracking (MOT) problem by tracking-by-detection paradigm, which tracks objects from the detected windows by associating detection boxes whose scores are higher than a given threshold. As such, the confidence score becomes the only indicator of bounding boxes when handling complicated cases, such as occlusions. However, a high confidence score cannot guarantee that the bounding box does not overlap with nearby objects, especially in crowded scenarios. In this paper, spatial uncertainty is proposed for MOT. Firstly, the statistical analysis indicates that spatial uncertainty is highly correlated to the occlusion ratio, which can better represent the level of occlusion of the detection boxes. It is measured by the proposed Sparse Tracker with Spatial Uncertainty (SSUTracker). Then, it is adopted to learn robust tracklet representation. The experimental results demonstrate that it improves overall performance. As a result, our approach achieves very competitive results on popular MOT17 and MOT20 benchmarks compared to state-of-the-art methods.
Pin-Jie Liao, Yu-Cheng Huang, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP4
2023 ReST: A Reconfigurable Spatial-Temporal Graph Model for Multi-Camera Multi-Object Tracking
abstract
Multi-Camera Multi-Object Tracking (MC-MOT) utilizes information from multiple views to better handle problems with occlusion and crowded scenes. Recently, the use of graph-based approaches to solve tracking problems has become very popular. However, many current graph-based methods do not effectively utilize information regarding spatial and temporal consistency. Instead, they rely on single-camera trackers as input, which are prone to fragmentation and ID switch errors. In this paper, we propose a novel reconfigurable graph model that first associates all detected objects across cameras spatially before reconfiguring it into a temporal graph for Temporal Association. This two-stage association approach enables us to extract robust spatial and temporal-aware features and address the problem with fragmented tracklets. Furthermore, our model is designed for online tracking, making it suitable for real-world applications. Experimental results show that the proposed graph model is able to extract more discriminating features for object tracking, and our model achieves state-of-the-art performance on several public datasets. Code is available at https://github.com/chengche6230/ReST.
Cheng-Che Cheng, Min-Xuan Qiu, Chen-Kuo Chiang, Shang-Hong Lai
ICCV4
2023 Rethinking Long-Tailed Visual Recognition with Dynamic Probability Smoothing and Frequency Weighted Focusing
abstract
Deep learning models trained on long-tailed (LT) datasets often exhibit bias towards head classes with high frequency. This paper highlights the limitations of existing solutions that combine class- and instance-level re-weighting loss in a naive manner. Specifically, we demonstrate that such solutions result in overfitting the training set, significantly impacting the rare classes. To address this issue, we propose a novel loss function that dynamically reduces the influence of outliers and assigns class-dependent focusing parameters. We also introduce a new long-tailed dataset, ICText-LT, featuring various image qualities and greater realism than artificially sampled datasets. Our method has proven effective, outperforming existing methods through superior quantitative results on CIFAR-LT, Tiny ImageNet-LT, and our new ICText-LT datasets. The source code and new dataset are available at https://github.com/nwjun/FFDS-Loss.
Wan Jun Nah, Chun Chet Ng, Che-Tsung Lin, Yeong Khang Lee, Jie-Long Kew, Zhi Qin Tan, Chee Seng Chan, Christopher Zach, Shang-Hong Lai
ICIP9
2023 Holistic Interaction Transformer Network for Action Detection
abstract
Actions are about how we interact with the environment, including other people, objects, and ourselves. In this paper, we propose a novel multi-modal Holistic Interaction Transformer Network (HIT) that leverages the largely ignored, but critical hand and pose information essential to most human actions. The proposed HIT network is a comprehensive bi-modal framework that comprises an RGB stream and a pose stream. Each of them separately models person, object, and hand interactions. Within each sub-network, an Intra-Modality Aggregation module (IMA) is introduced that selectively merges individual interaction units. The resulting features from each modality are then glued using an Attentive Fusion Mechanism (AFM). Finally, we extract cues from the temporal context to better classify the occurring actions using cached memory. Our method significantly outperforms previous approaches on the J-HMDB, UCF101-24, and MultiSports datasets. We also achieve competitive results on AVA. The code will be available at https://github.com/joslefaure/HIT.
Gueter Josmy Faure, Min-Hung Chen, Shang-Hong Lai
WACV3
2023 Cycle-object consistency for image-to-image domain adaptation
Che-Tsung Lin, Jie-Long Kew, Chee Seng Chan, Shang-Hong Lai, Christopher Zach
Pattern Recognit.4
2022 FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition
abstract
Current state-of-the-art deep learning based face recognition (FR) models require a large number of face identities for central training. However, due to the growing privacy awareness, it is prohibited to access the face images on user devices to continually improve face recognition models. Federated Learning (FL) is a technique to address the privacy issue, which can collaboratively optimize the model without sharing the data between clients. In this work, we propose a FL based framework called FedFR to improve the generic face representation in a privacy-aware manner. Besides, the framework jointly optimizes personalized models for the corresponding clients via the proposed Decoupled Feature Customization module. The client-specific personalized model can serve the need of optimized face recognition experience for registered identities at the local device. To the best of our knowledge, we are the first to explore the personalized face recognition in FL setup. The proposed framework is validated to be superior to previous approaches on several generic and personalized face recognition benchmarks with diverse FL scenarios. The source codes and our proposed personalized FR benchmark under FL setup are available at https://github.com/jackie840129/FedFR.
Chih-Ting Liu, Chien-Yi Wang, Shao-Yi Chien, Shang-Hong Lai
AAAI4
2022 Siamese U-Net for Image Anomaly Detection and Segmentation with Contrastive Learning
Chia Ying Lin, Shang-Hong Lai
BMVC2
2022 PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition
abstract
Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof types. However, the local characteristics of image captures, i.e., capturing devices and presenting materials, are ignored in existing works and we argue that such information is required for networks to discriminate between live and spoof images. In this work, we propose PatchNet which reformulates face anti-spoofing as a fine-grained patch-type recognition problem. To be specific, our framework recognizes the combination of capturing devices and presenting materials based on the patches cropped from non-distorted face images. This reformulation can largely improve the data variation and enforce the network to learn discriminative feature from local capture patterns. In addition, to further improve the generalization ability of the spoof feature, we propose the novel Asymmetric Margin-based Classification Loss and Self-supervised Similarity Loss to regularize the patch embedding space. Our experimental results verify our assumption and show that the model is capable of recognizing unseen spoof types robustly by only looking at local regions. Moreover, the fine-grained and patch-level reformulation of FAS outperforms the existing approaches on intra-dataset, cross-dataset, and domain generalization benchmarks. Furthermore, our PatchNet framework can enable practical applications like FewShot Reference-based FAS and facilitate future exploration of spoof-related intrinsic cues.
Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang, Shang-Hong Lai
CVPR4
2022 Local-Adaptive Face Recognition via Graph-based Meta-Clustering and Regularized Adaptation
abstract
Due to the rising concern of data privacy, it's reasonable to assume the local client data can't be transferred to a centralized server, nor their associated identity label is provided. To support continuous learning and fill the last-mile quality gap, we introduce a new problem setup called “local-adaptive face recognition (LaFR)”. Leveraging the environment-specific local data after the deployment of the initial global model, LaFR aims at getting optimal performance by training local-adapted models automatically and un-supervisely, as opposed to fixing their initial global model. We achieve this by a newly proposed embedding cluster model based on Graph Convolution Network (GCN), which is trained via meta-optimization procedure. Compared with previous works, our meta-clustering model can generalize well in unseen local environments. With the pseudo identity labels from the clustering results, we further introduce novel regularization techniques to improve the model adaptation performance. Extensive experiments on racial and internal sensor adaptation demonstrate that our proposed solution is more effective for adapting face recognition models in each specific environment. Meanwhile, we show that LaFR can further improve the global model by a simple federated aggregation over the updated local models.
Chien-Yi Wang, Kuan-Lun Tseng, Shang-Hong Lai, Baoyuan Wang
CVPR4
2022 Learning Monocular 3D Human Pose Estimation With Skeletal Interpolation
abstract
Deep learning has achieved unprecedented accuracy for monocular 3D human pose estimation. However, current learning-based 3D human pose estimation still suffers from poor generalization. Inspired by skeletal animation, which is popular in game development and animation production, we put forward an simple, intuitive yet effective interpolation-based data augmentation approach to synthesize continuous and diverse 3D human body sequences to enhance model generalization. The Transformer-based lifting network, trained with the augmented data, utilizes the self-attention mechanism to perform 2D-to-3D lifting and successfully infer high-quality predictions in the qualitative experiment. The quantitative result of cross-dataset experiment demonstrates that our resulting model achieves superior generalization accuracy on the publicly available dataset.
Akihiro Sugimoto, Shang-Hong Lai
ICASSP3
2022 CyEDA: Cycle-Object Edge Consistency Domain Adaptation
abstract
A difficulty of global-level translation is to preserve instance-level details in an image. Although some instance level translation methods can retain the details, most of them require either pre-trained object detection/segmentation network or annotation labels. In this work, we propose a novel method namely CyEDA to perform global level domain adaptation that can preserve image contents without any pre-trained networks integration or annotation labels. Specifically, we introduce blending masks and cycle-object edge consistency loss which exploit the preservation of image objects. We show that our approach can outperform other SOTAs in terms of image quality and FID score in both BDD100K and GTA datasets. The code and pre-trained models are publicly available at https://github.com/bjc1999/CyEDA.
Jing Chong Beh, Kam Woh Ng, Jie-Long Kew, Che-Tsung Lin, Chee Seng Chan, Shang-Hong Lai, Christopher Zach
ICIP6
2022 Extremely Low-Light Image Enhancement with Scene Text Restoration
abstract
Deep learning-based methods have made impressive progress in enhancing extremely low-light images - the image quality of the reconstructed images has generally improved. However, we found out that most of these methods could not sufficiently recover the image details, for instance, the texts in the scene. In this paper, a novel image enhancement framework is proposed to precisely restore the scene texts, as well as the overall quality of the image simultaneously under extremely low-light conditions. Mainly, we employed a self-regularised attention map, an edge map, and a novel text detection loss. In addition, leveraging the synthetic low-light images is beneficial for image enhancement on the genuine ones in terms of text detection. The quantitative and qualitative experimental results have shown that the proposed model outperforms state-of-the-art methods in image restoration, text detection, and text spotting on See In the Dark and ICDAR15 datasets.
Po-Hao Hsu, Che-Tsung Lin, Chun Chet Ng, Jie-Long Kew, Mei Yih Tan, Shang-Hong Lai, Chee Seng Chan, Christopher Zach
ICPR6
2022 Multi-Scale Patch-Based Representation Learning for Image Anomaly Detection and Segmentation
abstract
Unsupervised representation learning has been proven to be effective for the challenging anomaly detection and segmentation tasks. In this paper, we propose a multi-scale patch-based representation learning method to extract critical and representative information from normal images. By taking the relative feature similarity between patches of different local distances into account, we can achieve better representation learning. Moreover, we propose a refined way to improve the self-supervised learning strategy, thus allowing our model to learn better geometric relationship between neighboring patches. Through sliding patches of different scales all over an image, our model extracts representative features from each patch and compares them with those in the training set of normal images to detect the anomalous regions. Our experimental results on MVTec AD dataset and BTAD dataset demonstrate the proposed method achieves the state-of-the-art accuracy for both anomaly detection and segmentation.
Chin-Chia Tsai, Tsung-Hsuan Wu, Shang-Hong Lai
WACV3
2022 Disentangled Representation with Dual-stage Feature Learning for Face Anti-spoofing
abstract
As face recognition is widely used in diverse security-critical applications, the study of face anti-spoofing (FAS) has attracted more and more attention. Several FAS methods have achieved promising performance if the attack types in the testing data are included in the training data, while the performance significantly degrades for unseen attack types. It is essential to learn more generalized and discriminative features to prevent overfitting to pre-defined spoof attack types. This paper proposes a novel dual-stage disentangled representation learning method that can efficiently untangle spoof-related features from irrelevant ones. Un-like previous FAS disentanglement works with one-stage architecture, we found that the dual-stage training design can improve the training stability and effectively encode the features to detect unseen attack types. Our experiments show that the proposed method provides superior accuracy than the state-of-the-art methods on several cross-type FAS benchmarks.
Yu-Chun Wang, Chien-Yi Wang, Shang-Hong Lai
WACV3
2021 Moving-Object-Aware Anomaly Detection in Surveillance Videos
abstract
Video anomaly detection plays a crucial role in automatically detecting abnormal actions or events from surveillance video, which can help to protect public safety. Deep learning techniques have been extensively employed and achieved excellent anomaly detection results recently. However, previous image-reconstruction-based models did not fully exploit foreground object regions for the video anomaly detection. Some recent works applied pre-trained object detectors to provide local context in the video surveillance scenario for anomaly detection. Nevertheless, these methods require prior knowledge of object types for the anomaly which is somewhat contradictory to the problem setting of unsupervised anomaly detection. In this paper, we propose a novel framework based on learning the moving-object feature prediction based on a convolutional autoencoder architecture. We train our anomaly detector to be aware of moving-object regions in a scene without using an object detector or requiring prior knowledge of specific object classes for the anomaly. The appearance and motion features in moving objects regions provide comprehensive information of moving foreground objects for unsupervised learning of video anomaly detector. Besides, the proposed latent representation learning scheme encourages the convolutional autoencoder model to learn a more convergent latent representation for normal training data, while anomalous data exhibits quite different representations. We also propose a novel anomaly scoring method based on the feature prediction errors of moving foreground object regions and the latent representation regularity. Our experimental results demonstrate that the proposed approach achieves competitive results compared with SOTA methods on three public datasets for video anomaly detection.
Chun-Lung Yang, Tsung-Hsuan Wu, Shang-Hong Lai
AVSS3
2021 High-Accuracy RGB-D Face Recognition via Segmentation-Aware Face Depth Estimation and Mask-Guided Attention Network
abstract
Deep learning approaches have achieved highly accurate face recognition by training the models with very large face image datasets. Unlike the availability of large 2D face image datasets, there is a lack of large 3D face datasets available to the public. Existing public 3D face datasets were usually collected with few subjects, leading to the over-fitting problem. This paper proposes two CNN models to improve the RGB-D face recognition task. The first is a segmentation-aware depth estimation network, called DepthNet, which estimates depth maps from RGB face images by including semantic segmentation information for more accurate face region localization. The other is a novel mask-guided RGB-D face recognition model that contains an RGB recognition branch, a depth map recognition branch, and an auxiliary segmentation mask branch with a spatial attention module. Our DepthN et is used to augment a large 2D face image dataset to a large RGB-D face dataset, which is used for training an accurate RGB-D face recognition model. Furthermore, the proposed mask-guided RGB-D face recognition model can fully exploit the depth map and segmentation mask information and is more robust against pose variation than previous methods. Our experimental results show that DepthNet can produce more reliable depth maps from face images with the segmentation mask. Our mask-guided face recognition model outperforms state-of-the-art methods on several public 3D face datasets.
Meng-Tzu Chiu, Hsun-Ying Cheng, Chien-Yi Wang, Shang-Hong Lai
FG4
2021 Y-Net: Learning Domain Robust Feature Representation for ground camera image and large-scale image-based point cloud registration
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Baiqi Lai, Xuelun Shen, Ming Cheng 0002, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001
Inf. Sci.8
2021 GAN-Based Day-to-Night Image Style Transfer for Nighttime Vehicle Detection
abstract
Data augmentation plays a crucial role in training a CNN-based detector. Most previous approaches were based on using a combination of general image-processing operations and could only produce limited plausible image variations. Recently, GAN (Generative Adversarial Network) based methods have shown compelling visual results. However, they are prone to fail at preserving image-objects and maintaining translation consistency when faced with large and complex domain shifts, such as day-to-night. In this paper, we propose AugGAN, a GAN-based data augmenter which could transform on-road driving images to a desired domain while image-objects would be well-preserved. The contribution of this work is three-fold: (1) we design a structure-aware unpaired image-to-image translation network which learns the latent data transformation across different domains while artifacts in the transformed images are greatly reduced; (2) we quantitatively prove that the domain adaptation capability of a vehicle detector is not limited by its training data; (3) our object-preserving network provides significant performance gain in the difficult day-to-night case in terms of vehicle detection. AugGAN could generate more visually plausible images compared to competing methods on different on-road image translation tasks across domains. In addition, we quantitatively evaluate different methods by training Faster R-CNN and YOLO with datasets generated from the transformed results and demonstrate significant improvement on the object detection accuracies by using the proposed AugGAN model.
Che-Tsung Lin, Sheng-Wei Huang, Yen-Yi Wu, Shang-Hong Lai
IEEE Trans. Intell. Transp. Syst.4
2020 Multimodal Structure-Consistent Image-to-Image Translation
abstract
Unpaired image-to-image translation is proven quite effective in boosting a CNN-based object detector for a different domain by means of data augmentation that can well preserve the image-objects in the translated images. Recently, multimodal GAN (Generative Adversarial Network) models have been proposed and were expected to further boost the detector accuracy by generating a diverse collection of images in the target domain, given only a single/labelled image in the source domain. However, images generated by multimodal GANs would achieve even worse detection accuracy than the ones by a unimodal GAN with better object preservation. In this work, we introduce cycle-structure consistency for generating diverse and structure-preserved translated images across complex domains, such as between day and night, for object detector training. Qualitative results show that our model, Multimodal AugGAN, can generate diverse and realistic images for the target domain. For quantitative comparisons, we evaluate other competing methods and ours by using the generated images to train YOLO, Faster R-CNN and FCN models and prove that our model achieves significant improvement and outperforms other methods on the detection accuracies and the FCN scores. Also, we demonstrate that our model could provide more diverse object appearances in the target domain through comparison on the perceptual distance metric.
Che-Tsung Lin, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong Lai
AAAI4
2020 3D Object Detection from Consecutive Monocular Images
Chia-Chun Cheng, Shang-Hong Lai
ACCV (1)2
2020 Unified Representation Learning for Cross Model Compatibility
Chien-Yi Wang, Ya-Liang Chang, Shang-Ta Yang, Shang-Hong Lai
BMVC5
2020 ByeGlassesGAN: Identity Preserving Eyeglasses Removal for Face Images
Yu-Hui Lee, Shang-Hong Lai
ECCV (29)2
2020 Stylized-Colorization for Line Arts
abstract
We address a novel problem of stylized-colorization which colorizes a given line art using a given coloring style in text. This problem can be stated as multi-domain image translation and is more challenging than the current colorization problem because it requires not only capturing the illustration distribution but also satisfying the required coloring styles specific to anime such as lightness, shading, or saturation. We propose a GAN-based end-to-end model for stylized-colorization where the model has one generator and two discriminators. Our generator is based on the U-Net architecture and receives a pair of a line art and a coloring style in text as its input to produce a stylized-colorization image of the line art. Two discriminators, on the other hand, share weights at early layers to judge the stylized-colorization image in two different aspects: one for color and one for style. One generator and two discriminators are jointly trained in an adversarial and end-to-end manner. Extensive experiments demonstrate the effectiveness of our proposed model.
Tzu-Ting Fang, Duc Minh Vo, Akihiro Sugimoto, Shang-Hong Lai
ICPR4
2020 Attention-Based Deep Metric Learning for Near-Duplicate Video Retrieval
abstract
Near-duplicate video retrieval (NDVR) is an important and challenging problem due to the increasing amount of videos uploaded to the Internet. In this paper, we propose an attention-based deep metric learning method for NDVR. Our method is based on well-established principles: We leverage two-stream networks to combine RGB and optical flow features, and incorporate an attention module to effectively deal with distractor frames commonly observed in near duplicate videos. We further aggregate the features corresponding to multiple video segments to enhance the discriminative power. The whole system is trained using a deep metric learning objective with a Siamese architecture. Our experiments show that the attention module helps eliminate redundant and noisy frames, while focusing on visually relevant frames for solving NVDR. We evaluate our approach on recent large-scale NDVR datasets, CC_WEB_VIDEO, VCDB, FIVR and SVD. To demonstrate the generalization ability of our approach, we report results in both within- and cross-dataset settings, and show that the proposed method significantly outperforms state-of-the-art approaches.
Kuan-Hsun Wang, Chia-Chun Cheng, Yi-Ling Chen 0004, Yale Song, Shang-Hong Lai
ICPR5
2020 Learning to Match Ground Camera Image and UAV 3-D Model-Rendered Image Based on Siamese Network With Attention Mechanism
abstract
Different domain image sensors or imaging mechanisms provide cross-domain images when sensing the same scene. There is a domain shift between cross-domain images so that the image gap between different domains is the major challenge for measuring the similarity of the feature descriptors extracted from different domain images. Specifically, matching ground camera images and unmanned aerial vehicle (UAV) 3-D model-rendered images, which are two kinds of extremely challenging cross-domain images, is a way to establish indirectly the spatial relationship between 2-D and 3-D spaces. This provides a solution for the virtual-real registration of augmented reality (AR) in outdoor environments. However, during matching, handcrafted descriptors and existing learning-based feature descriptors limit the rendered images. In this letter, first, to learn robust and invariant 128-D local feature descriptors for ground camera and rendered images, we present a novel network structure, SiamAM-Net, which embeds the autoencoders with an attention mechanism into the Siamese network. Then, to narrow the gap between the cross-domain images during the optimizing of SiamAM-Net, we design an adaptive margin for the loss function. Finally, we match the ground camera-rendered images by using the learned local feature descriptors and explore the outdoor AR virtual-real registration. Experiments show that the local feature descriptors, learned by SiamAM-Net, are robust and achieve state-of-the-art retrieval performance on the cross-domain image data set of ground camera and rendered images. In addition, several outdoor AR applications also demonstrate the usefulness of the proposed outdoor AR virtual-real registration.
Weiquan Liu, Cheng Wang 0003, Xuesheng Bian, Shangshu Yu, Xiuhong Lin, Shang-Hong Lai, Dongdong Weng, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.7
2020 Multi-task CNN for restoring corrupted fingerprint images1
Wei Jing Wong, Shang-Hong Lai
Pattern Recognit.2
2019 Audio Tempo Estimation Method Improved by Rhythm Pattern and Data Augmentation
abstract
Tempo is the intuitive attribute of audio music, since people could feel fast or slow expressively and detect salient pulses to form perceived tempo value naturally. Nonetheless, for some audio, the tempo value could be ambiguous due to complex metrical level, different composing habit and creating style. Even though most of audio have the predominant tempo with consensus between the listeners, the others could have two dominant tempi. The challenge and goal of tempo estimation is to discriminate the salient tempi, mostly one or two tempos, related to the metric level by analyzing the audio signal directly. In this study, we propose the rhythm patterns of long-term periodicity curve derived from tempogram to improve the saliency detection. Besides, the data augmentation method is also invented to conquer the deficiency and representative of the three training datasets. The performance is evaluated on three public datasets in which the accuracy of “GiantSteps” dataset even outperforms the state-of-the-art tempo estimator of convolutional neural network implementation.
Fu-Hai Frank Wu, Shang-Hong Lai
CoDIT2
2019 Object Detection in Curved Space for 360-Degree Camera
abstract
360° camera has recently become popular since it can capture the whole 360° scene. A large number of related applications have been springing up. In this paper, We propose a deep learning based object detector that can be applied directly on 360° images. The proposed detector is based on modifications of the faster RCNN model. Three modification schemes are proposed here, including (1) distortion data augmentation, (2) introducing muilti-kernel layers for improving accuracy for distorted object detection, and (3) adding position information into the model for learning spatial information. Additionally, we create two datasets, 360GoogleStreetView and 360Videos, and perform experiments on these two datasets to demonstrate that our object detector provides superior accuracy for object detection directly on 360° images.
Kuan-Hsun Wang, Shang-Hong Lai
ICASSP2
2019 Automatic Generation of Photorealistic Training Data for Detection of Industrial Components
abstract
With the prosperous development of deep learning, people pay more attention to the needs of different training data. In this paper, we propose a method to automatically generate realistic training data for industrial components detection. Our method can generate a large scale of various synthetic images associated with the corresponding precise instance segmentation masks through the concept of domain randomization and style transfer. Besides, we demonstrate that the proposed method is effective to generate images for training a wrench detector. Our method can enhance the performance of the wrench detection task obviously, and it can be easily extended to the detection of different kinds of industrial components. As far as we know, the proposed method is novel and effective for generating training data for training detectors for industrial components.
Yu-Hui Lee, Chu-Chun Chuang, Shang-Hong Lai, Zih-Jian Jhang
ICIP3
2019 Ground Camera Images and UAV 3D Model Registration for Outdoor Augmented Reality
abstract
This paper presents a novel virtual-real registration approach for augmented reality (AR) in large-scale outdoor environments. Essentially, it is a pose estimation for the mobile camera images (ground camera images) in 3D model recovered by Unmanned Aerial Vehicle (UAV) image sequence via Structure-From-Motion (SFM) technology. The approach considers to indirectly establish the spatial relationship between 2D and 3D space by inferring the transformation relationship between the ground camera images and the UAV 3D model rendered images. Specifically, the proposed approach can overcome the positioning errors, which are deterioration and drift in the GPS, and deviation of orientation. The experimental results demonstrate the possibility of the proposed virtual-real registration approach, and show that the approach is robust, efficient and intuitive for AR in large-scale outdoor environments.
Weiquan Liu, Cheng Wang 0003, Shang-Hong Lai, Dongdong Weng, Xuesheng Bian, Xiuhong Lin, Xuelun Shen, Jonathan Li 0001
VR4
2019 Image captioning by incorporating affective concepts learned from both visual and textual components
Jufeng Yang, Jie Liang 0007, Bo Ren 0003, Shang-Hong Lai
Neurocomputing5
2018 SegmentedFusion: 3D Human Body Reconstruction Using Stitched Bounding Boxes
abstract
This paper presents SegmentedFusion, a method possessing the capability of reconstructing non-rigid 3D models of a human body by using a single depth camera with skeleton information. Our method estimates a dense volumetric 6D motion field that warps the integrated model into the live frame by segmenting a human body into different parts and building a canonical space for each part. The key feature of this work is that a deformed and connected canonical volume for each part is created, and it is used to integrate data. The dense volumetric warp field of one volume is represented efficiently by blending a few rigid transformations. Overall, SegmentedFusion is able to scan a non-rigidly deformed human surface as well as to estimate the dense motion field by using a consumer-grade depth camera. The experimental results demonstrate that SegmentedFusion is robust against fast inter-frame motion and topological changes. Since our method does not require prior assumption, SegmentedFusion can be applied to a wide range of human motions.
Shih-Hsuan Yao, Diego Thomas, Akihiro Sugimoto, Shang-Hong Lai, Rin-Ichiro Taniguchi
3DV4
2018 AugGAN: Cross Domain Adaptation with GAN-Based Data Augmentation
Sheng-Wei Huang, Che-Tsung Lin, Shu-Ping Chen, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong Lai
ECCV (9)6
2018 Emotion-Preserving Representation Learning via Generative Adversarial Network for Multi-View Facial Expression Recognition
abstract
Face frontalization is one way to overcome the pose variation problem, which simplifies multi-view recognition into one canonical-view recognition. This paper presents a multi-task learning approach based on the generative adversarial network (GAN) that learns the emotion-preserving representations in the face frontalization framework. Taking advantage of adversarial relationship between the generator and the discriminator in GAN, the generator can frontalize input non-frontal face images into frontal face images while preserving the identity and expression characteristics; in the meantime, it can employ the learnt emotion-preserving representations to predict the expression class label from the input face. The proposed network is optimized by combining both synthesis and classification objective functions to make the learnt representations generative and discriminative simultaneously. Experimental results demonstrate that the proposed face frontalization system is very effective for expression recognition with large head pose variations.
Ying-Hsiu Lai, Shang-Hong Lai
FG2
2018 Indoor Scene Layout Estimation from a Single Image
abstract
With the popularity of the hand devices and intelligent agents, many aimed to explore machine's potential in interacting with reality. Scene understanding, among the many facets of reality interaction, has gained much attention for its relevance in applications such as augmented reality (AR). Scene understanding can be partitioned into several sub tasks (i.e., layout estimation, scene classification, saliency prediction, etc). In this paper, we propose a deep learning-based approach for estimating the layout of a given indoor image in real-time. Our method consists of a deep fully convolutional network, a novel layout-degeneration augmentation method, and a new training pipeline which integrate an adaptive edge penalty and smoothness terms into the training process. Unlike previous deep learning-based methods that depend on post-processing refinement (e.g., proposal ranking and optimization), our method motivates the generalization ability of the network and the smoothness of estimated layout edges without deploying postprocessing techniques. Moreover, the proposed approach is time-efficient since it only takes the model one forward pass to render accurate layouts. We evaluate our method on LSUN Room Layout and Hedau dataset and obtain estimation results comparable with the state-of-the-art methods.
Hung-Jin Lin, Sheng-Wei Huang, Shang-Hong Lai, Chen-Kuo Chiang
ICPR3
2018 Editorial for ACCV'16 award papers
Shang-Hong Lai, Vincent Lepetit, Ko Nishino, Yoichi Sato 0001
Comput. Vis. Image Underst.1
2017 General Deep Image Completion with Lightweight Conditional Generative Adversarial Networks
Ching Wei Tseng, Hung-Jin Lin, Shang-Hong Lai
BMVC3
2017 Hierarchical Structured Dictionary Learning for image categorization
abstract
A novel Hierarchical Structured Dictionary Learning (HSDL) algorithm is proposed in this paper. It aims to learn class-specific dictionaries for all classes simultaneously in a hierarchical structure. A discriminative term based on Fisher discrimination criterion is jointly considered for both the class-specific dictionaries in the lower level and the shared dictionaries in the upper level to enhance the discrimination of dictionaries. The experimental results evaluated on the ImageNet database have shown the superior performance of HSDL over the state-of-the-art dictionary learning methods.
Tzu-Chan Chuang, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP3
2017 A deep learning approach towards pore extraction for high-resolution fingerprint recognition
abstract
As high-resolution fingerprint images are becoming more common, the pores have been found to be one of the promising candidates in improving the performance of automated fingerprint identification systems (AFIS). This paper proposes a deep learning approach towards pore extraction. It exploits the feature learning and classification capability of convolutional neural networks (CNNs) to detect pores on fingerprints. Besides, this paper also presents a unique affine Fourier moment-matching (AFMM) method of matching and fusing the scores obtained for three different fingerprint features to deal with both local and global linear distortions. Combining the two aforementioned contributions, an EER of 3.66% can be observed from the experimental results.
Hong-Ren Su, Kuang-Yu Chen, Wei Jing Wong, Shang-Hong Lai
ICASSP4
2017 Real-time video stitching
abstract
This paper presents a real-time video stitching system which can stitch videos acquired from multiple moving cameras. Most conventional stitching methods were developed under the assumption of static scenes without considering moving objects. However, cameras could move freely for general video stitching and the homography matrices between different camera views could vary from frame to frame. The proposed algorithm estimates the refined homography in both spatial and temporal domains. To be more specific, we first estimate homography between images acquired by different cameras by using RANSAC in the spatial domain at some key frames separated by a fixed interval. Then, we obtain temporally smooth homography transformations by using temporally linear interpolation for the frames between the key frames. For the image stitching, we correct the exposure of the stitched image by linear blending in the overlapping region and generate a panoramic view with cylindrical warping. To achieve real-time video stitching, we speed up our system with CUDA parallel programming. The experimental results show that the stitched views by using the proposed algorithm compare favorably with other methods. In addition, our video stitching system can achieve real-time performance and it is much more efficient compared to the other methods included in our experimental comparison.
Shuo-Han Yeh, Shang-Hong Lai
ICIP2
2017 Edge-preserving disparity map estimation from stereo videos for bokeh synthesis
abstract
We present a new method of estimating disparity maps from stereo videos for bokeh effect synthesis. In this work, we develop an improved total variation regularization and the robust L1norm in the data fidelity term (TV-L1) [4] based method to estimate edge-preserving disparity map without stereo rectification. The proposed algorithm improves the TV-L1approach by incorporating structure edge detection, occlusion area detection, textureless region detection and applying the guided filter to alleviate the inconsistency problem between the disparity map and color image around object boundary. Furthermore, we propose a temporal filter to improve the temporal consistency of the disparity maps computed from the stereo videos. We use saliency map to focus the synthesis result on the objects which attract human attention most. Experimental comparisons on various real videos are shown to demonstrate that the proposed algorithm generates more visually pleasing bokeh video synthesis compared with those by using previous stereo matching methods.
Wei-Lun Lan, Shih-Hsuan Yao, Shang-Hong Lai
ICME3
2017 Video synthesis from stereo videos with iterative depth refinement
Chen-Hao Wei, Shang-Hong Lai, Chen-Kuo Chiang
J. Vis. Commun. Image Represent.2
2017 Multi-scale energy optimization for object proposal generation
Congchao Wang, Jufeng Yang, Kai Wang 0001, Shang-Hong Lai
Multim. Tools Appl.4
2016 Accurate and robust face recognition from RGB-D images with a deep learning approach
Yuan-Cheng Lee, Jiancong Chen, Ching Wei Tseng, Shang-Hong Lai
BMVC4
2016 Lighting estimation from a single image containing multiple planes
abstract
In this paper, we present a novel lighting estimation algorithm for the scene containing two or more planes. This paper focuses on near point light source estimation. We first detect planar markers to estimate the poses of the 3D planes in the scene. Then we estimate the shading image from the captured image. A near point light source lighting model is used to define an objective function for light source estimation in this paper. The output of the proposed method is the lighting parameters estimated from minimizing the objective function. In the experiments, we test the proposed algorithm on synthetic data and real dataset. Our experimental results show the proposed algorithm outperforms the state-of-the-art lighting estimation method. Moreover, we develop an augmented reality system that includes lighting estimation by using the proposed algorithm.
Ping-Cheng Kuo, Hsin-Yuan Huang, Shang-Hong Lai
MMSys3
2016 A Multiattribute Sparse Coding Approach for Action Recognition From a Single Unknown Viewpoint
abstract
We propose a novel approach for view-independent action recognition using multiattribute sparse representation enforced with group constraints. First, an oversegmentation-based background modeling and foreground detection approach is employed to extract silhouettes from action videos. Then multiple time intervals of motion history image are computed to capture motion and pose information in human activities. To obtain a more accurate and discriminative representation, we propose multiattribute sparse representation for multiview action video classification. Actions with multiple attributes can be represented by individual attribute matrices to describe group property for each action instance. These attribute matrices are incorporated into the formulation of l1-minimization. The sparsity property as well as the group constraints make the basis selection in sparse coding more efficient in terms of accuracy. Especially, our approach is able to operate under the condition of partially labeled attributes in the training data. Finally, we demonstrate the proposed algorithm through experiments on three multiview human action datasets to show the effectiveness and robustness of the proposed method.
Te-Feng Su, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.3
2015 Novel Facial Expression Recognition by Combining Action Unit Detection with Sparse Representation Classification
abstract
This paper presents a multi-attribute sparse coding approach for facial expression recognition by regarding Action-Units (AUs) as attributes. AUs describe the movements of individual facial muscles, which are detected from corresponding attribute masks in this work. They can not only be used to de scribe group property which enforces basis selection from groups with the same AUs as best as possible, but also penalize the selection of atoms with the AU distance far away from the target instance. The group constraint and the AU similarity constraint are incorporated into the formulation of l1-minimization to determine the optimal sparse representation for facial expression. Finally, we demonstrate the proposed algorithm through experiments on two facial expression datasets to show the effectiveness and robustness of the proposed method.
Te-Feng Su, Ching-Hua Weng, Shang-Hong Lai
COMPSAC3
2015 Non-rigid registration of images with geometric and photometric deformation by using local affine Fourier-moment matching
abstract
Registration between images taken with different cameras, from different viewpoints or under different lighting conditions is a challenging problem. It needs to solve not only the geometric registration problem but also the photometric matching problem. In this paper, we propose to estimate the integrated geometric and photometric transformations between two images based on a local affine Fourier-moment matching framework, which is developed to achieve deformable registration. We combine the local Fourier moment constraints with the smoothness constraints to determine the local affine transforms in a hierarchal block model. Our experimental results on registering some real images related by large color and geometric transformations show the proposed registration algorithm provides superior image registration results compared to the state-of-the-art image registration methods.
Hong-Ren Su, Shang-Hong Lai
CVPR2
2015 Using line consistency to estimate 3D indoor Manhattan scene layout from a single image
abstract
In this paper, an optimization approach is proposed to estimate the 3D indoor Manhattan scene layout from a single input image. The proposed system models the interior space as a three-dimensional box which includes ceiling, floor, and walls. The regions corresponding to different surfaces can be calculated by projecting the 3D box onto the two-dimensional image with suitable camera and box parameters. This paper also utilizes the consistency of coplanar lines and the boundary edges between different surfaces to design a cost function. The rotation, translation, and box parameters of the interior layout can be estimated with an energy minimization process. In the experimental results, we apply the proposed algorithm to a number of real images of interior scenes to demonstrate the effectiveness of the proposed system.
Hsing-Chun Chang, Szu-Hao Huang, Shang-Hong Lai
ICIP3
2014 Fast 3D Object Alignment from Depth Image with 3D Fourier Moment Matching on GPU
abstract
In this paper, we develop a fast and accurate 3D object alignment system which can be applied to detect objects and estimate their 3D pose from a depth image containing cluttered background. The proposed 3D alignment system consists of two main algorithms: the first is the 3D detection algorithm to detect the top-level object from a depth map of the cluttered 3D objects, and the second is the 3D Fourier based point-set alignment algorithm to estimate the 3D object pose from an input depth image. We also implement the proposed 3D alignment algorithm on a GPU computing platform to speed up the computation of the object detection and Fourier-based image alignment algorithms in order to align the 3D object in real time.
Hong-Ren Su, Hao-Yuan Kuo, Shang-Hong Lai, Chin-Chia Wu
3DV3
2014 Online facial expression recognition based on combining texture and geometric information
abstract
Automatic facial expression recognition is a challenging problem in human-computer interaction. In this paper, we develop an online facial expression recognition system based on utilizing texture and motion features extracted from a video. The combination of both types of features capture static and dynamic facial information which enhances the recognition accuracy. In addition, our method is capable of automatically detecting both the onset and the apex of the expression of a face from a video. Our proposed expression recognition system is fully automatic and achieves superior performance compared to other state-of-the-art algorithms.
Ching-Hua Weng, Shang-Hong Lai
ICIP2
2014 Exploring Depth Information for Object Segmentation and Detection
abstract
We propose a new framework for performing object segmentation and detection simultaneously. Our method leverages with an MRF graphical model that comprises two kinds of nodes and two types of labels for inference. Specifically, we decompose an image into super pixels and generate segment proposals from each super pixel. The super pixels are then duplicated to form the two types of nodes. For each segmentation node, the model is to predict the object class label, while it is to decide the label corresponding to the best segment proposal selection at each detection node. The former is clearly a segmentation problem and the latter a detection problem. We link the two tasks by establishing a unified energy function that has a joint energy term accounting for the compatibility of the segmentation and detection labelings. Marginalizing by fixing either type of variables, the energy function can be switched into the one specifically for detection or segmentation. This property enables an alternating procedure to conveniently obtain the optimal labelings. To better explain the geometry about the objects and the scene, we use the depth information so that 3-D distances between super pixels are available in computing each energy term. Experimental results on a dataset with depth information are provided to support the effectiveness of our method.
Tyng-Luh Liu, Kai-Yueh Chang, Shang-Hong Lai
ICPR3
2014 Guest editorial: Advances in 3D video processing
Shang-Hong Lai, Gene Cheung, Dinei A. F. Florêncio, Peter Eisert, Yo-Sung Ho
J. Vis. Commun. Image Represent.1
2014 A novel gradient attenuation Richardson-Lucy algorithm for image motion deblurring
Hao-Liang Yang, Po-Hao Huang, Shang-Hong Lai
Signal Process.3
2014 Spatio-Temporally Consistent View Synthesis From Video-Plus-Depth Data With Global Optimization
abstract
We propose a novel algorithm to generate a virtual-view video from a video-plus-depth sequence. The proposed method enforces the spatial and temporal consistency in the disocclusion regions by formulating the problem as an energy minimization problem in a Markov random field (MRF) framework. At the system level, we first recover the depth images and the motion vector maps after the image warping with the preprocessed depth map. Then, we formulate the energy function for the MRF with additional shift variables for each node. To reduce the high computational complexity of applying belief propagation (BP) to this problem, we present a multilevel BPs by using BP with smaller numbers of label candidates for each level. Finally, the Poisson image reconstruction is applied to improve the color consistency along the boundary of the disocclusion region in the synthesized image. Experimental results demonstrate the performance of the proposed method on several publicly available datasets.
Hsiao-An Hsu, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.3
2013 Multi-attributed Dictionary Learning for Sparse Coding
abstract
We present a multi-attributed dictionary learning algorithm for sparse coding. Considering training samples with multiple attributes, a new distance matrix is proposed by jointly incorporating data and attribute similarities. Then, an objective function is presented to learn category-dependent dictionaries that are compact (closeness of dictionary atoms based on data distance and attribute similarity), reconstructive (low reconstruction error with correct dictionary) and label-consistent (encouraging the labels of dictionary atoms to be similar). We have demonstrated our algorithm on action classification and face recognition tasks on several publicly available datasets. Experimental results with improved performance over previous dictionary learning methods are shown to validate the effectiveness of the proposed algorithm.
Chen-Kuo Chiang, Te-Feng Su, Chih Yen, Shang-Hong Lai
ICCV4
2013 Efficient vehicle detection with adaptive scan based on perspective geometry
abstract
Vehicle detection is an important research problem for Advanced Driver Assistance Systems to improve driving safety. Most existing methods are based on the sliding window search framework to locate vehicles in an image. However, such methods usually produce large numbers of false positives and are computationally intensive. In this paper, we propose an efficient vehicle detection algorithm that dramatically reduces the search space based on the perspective geometry of the road. In the training phase, we search a few images to locate all possible vehicle regions by using the standard HOG-based vehicle detector. Pairs of vehicle candidates that satisfy the projective geometry constraints are used to estimate the linear vehicle width model with respect to y coordinates in the image. Then an adaptive scan strategy based on the estimated vehicle width model is proposed to efficiently detect vehicles in an image. Experimental results show that the proposed algorithm provides improved performance in terms of both speed and accuracy compared to standard sliding-windows search strategy.
Yu-Chun Chen, Te-Feng Su, Shang-Hong Lai
ICIP3
2013 Rolling shutter correction for video with large depth of field
abstract
Rolling shutter correction has attracted considerable attention in recent years. Several algorithms have been proposed to correct the distortion. Previous methods on rolling shutter correction did not consider depth variations in the scene. In this work, we overcome the limitation of the previous works that the depth of field in the scene is small. We present a correction model for rectifying the rolling shutter video based on the depth maps estimated from the rolling shutter video. In addition, we propose a two-stage optimization algorithm to estimate the temporal camera motion and the associated depth maps. Experimental results show the improvement of the proposed rolling shutter correction algorithm that takes the depth information into account.
Yen-Hao Chiao, Tung-Ying Lee, Shang-Hong Lai
ICIP3
2013 Learning spatial weighting for facial expression analysis via constrained quadratic programming
Chia-Te Liao, Hui-Ju Chuang, Chih-Hsueh Duan, Shang-Hong Lai
Pattern Recognit.4
2013 Face Verification With Local Sparse Representation
abstract
In this letter, a local sparse representation is proposed for face components to describe the local structure and characteristics of the face image for face verification. We first learn a dictionary from collected local patches of face images. Then, a novel local descriptor is presented by using sparse coefficients obtained by the learned dictionary and local face patches from face components to represent the entire human face. We demonstrate the performance of the proposed local sparse representation method on several publicly available datasets. Extensive experiments on both CMU PIE dataset and the challenging LFW database have shown the effectiveness of the proposed method.
Chih-Hsueh Duan, Chen-Kuo Chiang, Shang-Hong Lai
IEEE Signal Process. Lett.3
2013 Learning Component-Level Sparse Representation for Image and Video Categorization
abstract
A novel component-level dictionary learning framework that exploits image/video group characteristics based on sparse representation is introduced in this paper. Unlike the previous methods that select the dictionaries to best reconstruct the data, we present an energy minimization formulation that jointly optimizes the learning of both sparse dictionary and component-level importance within one unified framework to provide a discriminative and sparse representation for image/video groups. The importance measures how well each feature component represents the group property with the dictionary. Then, the dictionary is updated iteratively to reduce the influence of unimportant components, thus refining the sparse representation for each group. In the end, by keeping the top K important components, a compact representation is obtained for the sparse coding dictionary. Experimental results on several public image and video data sets are shown to demonstrate the superior performance of the proposed algorithm compared with the-state-of-the-art methods.
Chen-Kuo Chiang, Chao-Hsien Liu, Chih-Hsueh Duan, Shang-Hong Lai
IEEE Trans. Image Process.4
2012 Parallelized Random Walk algorithm for background substitution on a multi-core embedded platform
abstract
Random Walk (RW) is a popular algorithm and can be applied to many applications in computer vision. In this paper, a fast algorithm is proposed to solve the large linear system in RW based on adapting the Gauss-Seidel method on a multi-core embedded system. Two tables, TYPE and INDEX, are introduced to fast locate the required data for the close-form solution. The computational overhead, along with the memory requirement, to solve the linear system can be reduced greatly, thus making the RW algorithm feasible to many applications on an embedded system. In addition, the proposed fast method is parallelized for a heterogeneous multi-core embedded platform to make the most use of the benefits of the system architecture. Experimental results show that the computational overhead can be significantly reduced by the proposed algorithm.
Yutzu Lee, Chen-Kuo Chiang, Yu-Wei Sun, Te-Feng Su, Shang-Hong Lai
ICASSP5
2012 Learning expression kernels for facial expression intensity estimation
abstract
Although many studies of facial expression analysis have been conducted, most previous works indeed focused on expression recognition. Different from previous works, this paper proposes a novel approach to learn the expression kernel for facial expression intensity estimation. The solution involves first aligning the optical flow to a neutral face to reduce inter-person variations in facial geometry, followed by solving an optimization problem with the ordinal ranking of expression intensities in temporal domain as constraints. Extensive experiments on the Cohn-Kanade database manifest that using the learned expression kernels leads to superior performance than the previous methods for facial expression intensity estimation.
Chia-Te Liao, Hui-Ju Chuang, Shang-Hong Lai
ICASSP3
2012 Parallelization of Belief Propagation on Cell Processors for Stereo Vision
abstract
Markov random field models provide a robust formulation for the stereo vision problem of inferring three-dimensional scene geometry from two images taken from different viewpoints. One of the most advanced algorithms for solving the associated energy minimization problem in the formulation is belief propagation (BP). Although BP provides very accurate results in solving stereo vision problems, the high computational cost of the algorithm hinders it from real-time applications. In recent years, multicore architectures have been widely adopted in various industrial application domains. The high computing power of multicore processors provides new opportunities to implement stereo vision algorithms. This article examines and extracts the parallelisms in the BP method for stereo vision on multicore processors. This article shows that parallelism of the algorithm can be efficiently utilized on multicore processors. The results show that parallelization on multicore processors provides a speedup for the BP algorithm of almost 15 times compared to the single-processor implementation on the PPE of the Cell BE. The experimental results also indicate that a frame rate of 6.5 frames/second is possible when implementing the parallelized BP algorithm on the multicore processor of Cell BE with one PPE and six SPEs.
Kun-Yuan Hsieh, Chi-Hua Lai, Shang-Hong Lai, Jenq Kuen Lee
ACM Trans. Embed. Comput. Syst.3
2011 From co-saliency to co-segmentation: An efficient and fully unsupervised energy minimization model
abstract
We address two key issues of co-segmentation over multiple images. The first is whether a pure unsupervised algorithm can satisfactorily solve this problem. Without the user's guidance, segmenting the foregrounds implied by the common object is quite a challenging task, especially when substantial variations in the object's appearance, shape, and scale are allowed. The second issue concerns the efficiency if the technique can lead to practical uses. With these in mind, we establish an MRF optimization model that has an energy function with nice properties and can be shown to effectively resolve the two difficulties. Specifically, instead of relying on the user inputs, our approach introduces a co-saliency prior as the hint about possible foreground locations, and uses it to construct the MRF data terms. To complete the optimization framework, we include a novel global term that is more appropriate to co-segmentation, and results in a submodular energy function. The proposed model can thus be optimally solved by graph cuts. We demonstrate these advantages by testing our method on several benchmark datasets.
Kai-Yueh Chang, Tyng-Luh Liu, Shang-Hong Lai
CVPR3
2011 Detecting moving objects from dynamic background with shadow removal
abstract
Background subtraction is commonly used to detect foreground objects in video surveillance. Traditional background subtraction methods are usually based on the assumption that the background is stationary. However, they are not applicable to dynamic background, whose background images change over time. In this paper, we propose an adaptive Local-Patch Gaussian Mixture Model (LPGMM) as the dynamic background model for detecting moving objects from video with dynamic background. Then, the SVM classification is employed to discriminate between foreground objects and shadow regions. Finally, we show some experimental results on several video sequences to demonstrate the effectiveness and robustness of the proposed method.
Shih-Chieh Wang, Te-Feng Su, Shang-Hong Lai
ICASSP3
2011 Fusing generic objectness and visual saliency for salient object detection
abstract
We present a novel computational model to explore the relatedness of objectness and saliency, each of which plays an important role in the study of visual attention. The proposed framework conceptually integrates these two concepts via constructing a graphical model to account for their relationships, and concurrently improves their estimation by iteratively optimizing a novel energy function realizing the model. Specifically, the energy function comprises the objectness, the saliency, and the interaction energy, respectively corresponding to explain their individual regularities and the mutual effects. Minimizing the energy by fixing one or the other would elegantly transform the model into solving the problem of objectness or saliency estimation, while the useful information from the other concept can be utilized through the interaction term. Experimental results on two benchmark datasets demonstrate that the proposed model can simultaneously yield a saliency map of better quality and a more meaningful objectness output for salient object detection.
Kai-Yueh Chang, Tyng-Luh Liu, Hwann-Tzong Chen, Shang-Hong Lai
ICCV4
2011 Learning component-level sparse representation using histogram information for image classification
abstract
A novel component-level dictionary learning framework which exploits image group characteristics within sparse coding is introduced in this work. Unlike previous methods, which select the dictionaries that best reconstruct the data, we present an energy minimization formulation that jointly optimizes the learning of both sparse dictionary and component level importance within one unified framework to give a discriminative representation for image groups. The importance measures how well each feature component represents the image group property with the dictionary by using histogram information. Then, dictionaries are updated iteratively to reduce the influence of unimportant components, thus refining the sparse representation for each image group. In the end, by keeping the top K important components, a compact representation is derived for the sparse coding dictionary. Experimental results on several public datasets are shown to demonstrate the superior performance of the proposed algorithm compared to the-state-of-the-art methods.
Chen-Kuo Chiang, Chih-Hsueh Duan, Shang-Hong Lai, Shih-Fu Chang
ICCV3
2011 People Localization in a Camera Network Combining Background Subtraction and Scene-Aware Human Detection
Tung-Ying Lee, Tsung-Yu Lin, Szu-Hao Huang, Shang-Hong Lai, Shang-Chih Hung
MMM (1)4
2011 Recovering Depth Map from Video with Moving Objects
Hsiao-Wei Chen, Shang-Hong Lai
PSIVT (2)2
2011 Pedestrian Image Segmentation via Shape-Prior Constrained Random Walks
Ke-Chun Li, Hong-Ren Su, Shang-Hong Lai
PSIVT (2)3
2011 CT-MR Image Registration in 3D K-Space Based on Fourier Moment Matching
Hong-Ren Su, Shang-Hong Lai
PSIVT (2)2
2011 Blind Image Deblurring with Modified Richardson-Lucy Deconvolution for Ringing Artifact Suppression
Hao-Liang Yang, Yen-Hao Chiao, Po-Hao Huang, Shang-Hong Lai
PSIVT (2)4
2011 Robust 3D object pose estimation from a single 2D image
abstract
In this paper, we propose a robust algorithm for 3D object pose estimation from a single 2D image. The proposed pose estimation algorithm is based on modifying the traditional image projection error function to a sum of squared image projection errors weighted by their associated distances. By using an Euler angle representation, we formulate the energy minimization for the pose estimation problem as searching a global minimum solution. Based on this framework, the proposed algorithm employs robust techniques to detect outliers in a coarse-to-fine fashion, thus providing very robust pose estimation. Our experiments show that the algorithm outperforms previous methods under noisy conditions.
Chia-Ming Cheng, Hsiao-Wei Chen, Tung-Ying Lee, Shang-Hong Lai, Ya-Hui Tsai
VCIP4
2011 Wide-angle distortion correction by Hough transform and gradient estimation
abstract
Wide-angle cameras have been widely used in surveillance and endoscopic imaging. An automatic distortion correction method is very useful for these applications. Traditional methods extract corners or curved straight lines for estimating distortion parameters. Hough transform is a powerful tool to assess straightness. However, previous methods usually require some human intervention or only focus on using a single curve. In this paper, we propose a new method based on Hough transform by considering all curves into the estimation of distortion parameters. By considering the relationship between distortion parameters and curves, our method is fully automatic and does not require manual selection of curves in an image. Experiments on synthetic and real datasets have been conducted. The results of our method are also compared with other Hough Transform based methods in quantitative measures. The experimental results show that the accuracy of the proposed automatic method is comparable to those of other manual line-based methods.
Tung-Ying Lee, Tzu-Shan Chang, Shang-Hong Lai, Kai-Che Liu, Hurng-Sheng Wu
VCIP3
2011 Accurate depth map estimation from video via MRF optimization
abstract
In this paper, we propose a novel system to estimate depth maps of outdoor scenes from a video sequence. According to the characteristics of a video, our approach considers more information in the temporal domain than the traditional depth reconstruction methods. We perform Structure From Motion (SfM) on consecutive image frames from a video from SIFT feature point correspondences, which provides some camera information, including 3D translation and rotation, for all the images. Then, we compute the constrained optical flow between selected scenes so that we can solve an over-constrained linear system to estimate the depth map for all pixels at each frame. In addition, mean shift image segmentation is incorporated to aggregate the depth estimation. Thus, this initial depth map is used as the data term of our pixel-based and region-based Markov Random Field (MRF) formulation for depth map estimation. The proposed MRF depth estimation not only imposes adaptive smoothness constraints but also includes sky detection in the final depth map estimation. By minimizing the associated MRF energy function for each frame, we obtain refined depth maps that achieve detail-preserving and temporally consistent depth estimation results.
Sheng-Po Tseng, Shang-Hong Lai
VCIP2
2011 Bipartite Polar Classification for Surface Reconstruction
abstract
Abstract In this paper, we propose bipartite polar classification to augment an input unorganized point set ℘ with two disjoint groups of points distributed around the ambient space of ℘ to assist the task of surface reconstruction. The goal of bipartite polar classification is to obtain a space partitioning of ℘ by assigning pairs of Voronoi poles into two mutually invisible sets lying in the opposite sides of ℘ through direct point set visibility examination. Based on the observation that a pair of Voronoi poles are mutually invisible, spatial classification is accomplished by carving away visible exterior poles with their counterparts simultaneously determined as interior ones. By examining the conflicts of mutual invisibility, holes or boundaries can also be effectively detected, resulting in a hole‐aware space carving technique. With the classified poles, the task of surface reconstruction can be facilitated by more robust surface normal estimation with global consistent orientation and off‐surface point specification for variational implicit surface reconstruction. We demonstrate the ability of the bipartite polar classification to achieve robust and efficient space carving on unorganized point clouds with holes and complex topology and show its application to surface reconstruction.
Yi-Ling Chen 0004, Tung-Ying Lee, Bing-Yu Chen 0004, Shang-Hong Lai
Comput. Graph. Forum4
2011 A novel robust kernel for visual learning problems
Chia-Te Liao, Shang-Hong Lai
Neurocomputing2
2011 Reconstructing 3D Face Model with Associated Expression Deformation from a Single Face Image via Constructing a Low-Dimensional Expression Deformation Manifold
abstract
Facial expression modeling is central to facial expression recognition and expression synthesis for facial animation. In this work, we propose a manifold-based 3D face reconstruction approach to estimating the 3D face model and the associated expression deformation from a single face image. With the proposed robust weighted feature map (RWF), we can obtain the dense correspondences between 3D face models and build a nonlinear 3D expression manifold from a large set of 3D facial expression models. Then a Gaussian mixture model in this manifold is learned to represent the distribution of expression deformation. By combining the merits of morphable neutral face model and the low-dimensional expression manifold, a novel algorithm is developed to reconstruct the 3D face geometry as well as the facial deformation from a single face image in an energy minimization framework. Experimental results on simulated and real images are shown to validate the effectiveness and accuracy of the proposed algorithm.
Shu-Fan Wang, Shang-Hong Lai
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Fast H.264 Encoding Based on Statistical Learning
abstract
H.264/AVC, the latest video coding standard of the Joint Video Team, greatly outperforms previous standards in terms of coding bitrate and video quality, because it adopts several new techniques. However, the computational complexity is also considerably increased due to these new components. In this paper, we propose fast algorithms based on statistical learning to reduce the computational cost involved in three main components in H.264 encoder, i.e., intermode decision, multi-reference motion estimation (ME), and intra-mode prediction. First, representative features are extracted to build the learning models. Then, an offline pre-classification approach is used to determine the best results from the extracted features, thus a significant amount of computation is reduced based on the classification strategy. The proposed statistical learning-based approach is applied to the aforementioned three main components in H.264 encoder to speed up the computation. Experimental results show that the ME time of the proposed system is significantly sped up with 12 times faster than the conventional fast ME algorithm of H.264, and the total encoding time of the proposed encoder is greatly reduced with about four times faster than the fast encoder EPZS in the H.264 reference code with negligible video quality degradation.
Chen-Kuo Chiang, Wei-Hau Pan, Chiuan Hwang, ShinShan Zhuang, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.5
2011 An Orientation Inference Framework for Surface Reconstruction From Unorganized Point Clouds
abstract
In this paper, we present an orientation inference framework for reconstructing implicit surfaces from unoriented point clouds. The proposed method starts from building a surface approximation hierarchy comprising of a set of unoriented local surfaces, which are represented as a weighted combination of radial basis functions. We formulate the determination of the globally consistent orientation as a graph optimization problem by treating the local implicit patches as nodes. An energy function is defined to penalize inconsistent orientation changes by checking the sign consistency between neighboring local surfaces. An optimal labeling of the graph nodes indicating the orientation of each local surface can, thus, be obtained by minimizing the total energy defined on the graph. The local inference results are propagated over the model in a front-propagation fashion to obtain the global solution. The reconstructed surfaces are consolidated by a simple and effective inspection procedure to locate the erroneously fitted local surfaces. A progressive reconstruction algorithm that iteratively includes more oriented points to improve the fitting accuracy and efficiently updates the RBF coefficients is proposed. We demonstrate the performance of the proposed method by showing the surface reconstruction results on some real-world 3-D data sets with comparison to those by using the previous methods.
Yi-Ling Chen 0004, Shang-Hong Lai
IEEE Trans. Image Process.2
2011 Compressibility-Aware Media Retargeting With Structure Preserving
abstract
A number of algorithms have been proposed for intelligent image/video retargeting with image content retained as much as possible. However, they usually suffer from some artifacts in the results, such as ridge or structure twist. In this paper, we present a structure-preserving media retargeting technique that preserves the content and image structure as best as possible. Different from the previous pixel or grid based methods, we estimate the image content saliency from the structure of the content. A block structure energy is introduced with a top-down strategy to constrain the image structure inside to deform uniformly in either x or y direction. However, the flexibilities for retargeting are quite different for different images. To cope with this problem, we propose a compressibility assessment scheme for media retargeting by combining the entropies of image gradient magnitude and orientation distributions. Thus, the resized media is produced to preserve the image content and structure as best as possible. Our experiments demonstrate that the proposed method provides resized images/videos with better preservation of content and structure than those by the previous methods.
Shu-Fan Wang, Shang-Hong Lai
IEEE Trans. Image Process.2
2011 A learning-based contrarian trading strategy via a dual-classifier model
abstract
Behavioral finance is a relatively new and developing research field which adopts cognitive psychology and emotional bias to explain the inefficient market phenomenon and some irrational trading decisions. Unlike the experts in this field who tried to reason the price anomaly and applied empirical evidence in many different financial markets, we employ the advanced binary classification algorithms, such as AdaBoost and support vector machines, to precisely model the overreaction and strengthen the portfolio compositions of the contrarian trading strategies. The novelty of this article is to discover the financial time-series patterns through a high-dimensional and nonlinear model which is constructed by integrated knowledge of finance and machine learning techniques. We propose a dual-classifier learning framework to select candidate stocks from the past results of original contrarian trading strategies based on the defined learning targets. Three different feature extraction methods, including wavelet transformation, historical return distribution, and various technical indicators, are employed to represent these learning samples in a 381-dimensional financial time-series feature space. Finally, we construct the classifier models with four different learning kernels and prove that the proposed methods could improve the returns dramatically, such as the 3-year return that improved from 26.79% to 53.75%. The experiments also demonstrate significantly higher portfolio selection accuracy, improved from 57.47% to 66.41%, than the original contrarian trading strategy. To sum up, all these experiments show that the proposed method could be extended to an effective trading system in the historical stock prices of the leading U.S. companies of S&P 100 index.
Szu-Hao Huang, Shang-Hong Lai, Shih-Hsien Tai
ACM Trans. Intell. Syst. Technol.2
2010 Over-Segmentation Based Background Modeling and Foreground Detection with Shadow Removal by Using Hierarchical MRFs
Te-Feng Su, Yi-Ling Chen 0004, Shang-Hong Lai
ACCV (3)3
2010 A learning-based system for generating exaggerative caricature from face images with expression
abstract
In this paper, we propose a learning-based system for generating exaggerative caricatures with expression. Most of the previous works can only deal with frontal face images with neutral expression without glasses or hats, and are unable to apply more than one drawing prototype which was learned from the caricatures drawn by one single cartoonist at a time. The proposed caricature generation system exaggerates face images with expressions and learns the drawing prototypes from training data as well. Experimental results show that our system can capture the features selected by the artist and exaggerate them in similar ways.
Ting-Ting Yang, Shang-Hong Lai
ICASSP2
2010 MRI motion artifact correction based on spectral extrapolation with generalized series
abstract
Motion artifact is a serious problem to MRI diagnosis and analysis due to periodic motion or patient motion during the image acquisition. A metric-based method, EXTRACT, used the information by using extrapolation of the k-space data and a metric based on correlation between the extrapolated k-space data and the motion corrupted k-space data. In this paper, we propose to use the finite support technique in conjunction with the generalized series in the spectral extrapolation to determine the motion that causes phase shift in the corrupted k-space data, thus improving the performance of MRI motion correction. Some experimental results are given to demonstrate the performance of the proposed algorithm.
Hong-Ren Su, Tung-Ying Lee, Shang-Hong Lai, Ti-Chiun Chang
ICIP3
2010 Fingerprint compression: An adaptive and fast DCT-based approach
abstract
Fingerprints have been used to accurately identify people since the nineteenth century. However, as more and more persons included into the repository, the size of database grows explosively. Storing the high-quality fingerprint images in low bit-rate thus becomes necessary. In this paper, a novel DCT-based coder is developed for fingerprint compression by using the specific energy distributions of fingerprint patterns. An adaptive scheme, which utilizes spatial-oriented tree (SOT) construction on the transformed image blocks, is proposed to improve the compression process. Consequently, the proposed coding scheme yields the enhancement of 0.5 to 2.5 dB over WSQ, SPIHT and JPEG2000 for fingerprint patterns at the same compression ratio. The computation complexity of this method is O(n).
Yu-Lin Wang, Chia-Te Liao, Alvin Wen-Yu Su, Shang-Hong Lai
ICIP4
2010 A novel structure-from-motion strategy for refining depth map estimation and multi-view synthesis in 3DTV
abstract
The video-plus-depth format has been widely used for representing the 3D scene due to its main advantage of compatibility to image format. In practice, the depth inconsistency may lead to unsatisfactory view synthesis results. In this paper, we propose a new structure-from-motion (SfM) technique, called locally temporal bundle adjustment (LTBA), to handle the dynamic scenes as well as the static camera motion, which violates the conventional structure from motion assumption. By integrating the camera information, depth map, and video temporally, we develop a geometric quadrilateral filter to reduce noise in the depth map and enhance the spatio-temporal consistency to improve the quality of depth maps. We show the improved quality of dynamic depth maps by using the proposed algorithm through experiments on real video-plus-depth sequences.
Chia-Ming Cheng, Xiao-An Hsu, Shang-Hong Lai
ICME3
2010 Image compressibility assessment and the application of structure-preserving image retargeting
abstract
A number of algorithms have been proposed for intelligent image/video retargeting with important content retained as much as possible. In some cases, we can notice that they suffer from artifacts in the resized results, such as ridge or structure twist. In this paper, we suggest that the compressibility of an image should be estimated properly first by analyzing the image structure to determine the optimal scaling factors for the resizing algorithm. To cope with this problem, we propose a compressibility assessment scheme by combining the entropies of image gradient magnitude and orientation distributions. In order to further improve the result, we also present a structure-preserving media retargeting technique that preserves the content and image structure as best as possible. Since we focus on protecting the content structure, a block structure energy is introduced with a top-down strategy to constrain the image structure inside to scale uniformly in either x or y direction. Our experiments demonstrate that the proposed compressibility assessment scheme provides better preservation of content and structure in the resized images/videos than those by the previous methods.
Shu-Fan Wang, Shang-Hong Lai
ICME2
2010 Robust Fourier-Based Image Alignment with Gradient Complex Image
abstract
The paper proposes a robust image alignment framework based on Fourier transform of a gradient complex image. The proposed Fourier-based algorithm can handle translation, rotation, and scaling, and it is robust against noise and non-uniform illumination. The proposed alignment algorithm is further extended to work under occlusion by partitioning the template and performing the Fourier-based alignment for all partitioned sub-templates in a voting framework. Our experiments show superior alignment results by using the proposed robust Fourier-based alignment over the previous related methods.
Hong-Ren Su, Shang-Hong Lai, Ya-Hui Tsai
ICPR2
2010 Binary Orientation Trees for Volume and Surface Reconstruction from Unoriented Point Clouds
abstract
Abstract Given a complete unoriented point set, we propose a binary orientation tree (BOT) for volume and surface representation, which roughly splits the space into the interior and exterior regions with respect to the input point set. The BOTs are constructed by performing a traditional octree subdivision technique while the corners of each cell are associated with a tag indicating thein/outrelationship with respect to the input point set. Starting from the root cell, a growing stage is performed to efficiently assign tags to the connected empty sub‐cells. The unresolved tags of the remaining cell corners are determined by examining their visibility via the hidden point removal operator. We show that the outliers accompanying the input point set can be effectively detected during the construction of the BOTs. After removing the outliers and resolving thein/outtags, the BOTs are ready to support any volume or surface representation techniques. To represent the surfaces, we also present a modified MPU implicits algorithm enabled to reconstruct surfaces from the input unoriented point clouds by taking advantage of the BOTs.
Yi-Ling Chen 0004, Bing-Yu Chen 0004, Shang-Hong Lai, Tomoyuki Nishita
Comput. Graph. Forum3
2010 Manifold-Based 3D Face Caricature Generation with Individualized Facial Feature Extraction
abstract
Abstract Caricature is an interesting art to express exaggerated views of different persons and things through drawing. The face caricature is popular and widely used for different applications. To do this, we have to properly extract unique/specialized features of a person's face. A person's facial feature not only depends on his/her natural appearance, but also the associated expression style. Therefore, we would like to extract the neutural facial features and personal expression style for different applicaions. In this paper, we represent the 3D neutral face models in BU–3DFE database by sparse signal decomposition in the training phase. With this decomposition, the sparse training data can be used for robust linear subspace modeling of public faces. For an input 3D face model, we fit the model and decompose the 3D model geometry into a neutral face and the expression deformation separately. The neutral geomertry can be further decomposed into public face and individualized facial feature. We exaggerate the facial features and the expressions by estimating the probability on the corresponding manifold. The public face, the exaggerated facial features and the exaggerated expression are combined to synthesize a 3D caricature for a 3D face model. The proposed algorithm is automatic and can effectively extract the individualized facial features from an input 3D face model to create 3D face caricature.
S. F. Wang, Shang-Hong Lai
Comput. Graph. Forum2
2010 An Optical Flow-Based Approach to Robust Face Recognition Under Expression Variations
abstract
Face recognition is one of the most intensively studied topics in computer vision and pattern recognition, but few are focused on how to robustly recognize faces with expressions under the restriction of one single training sample per class. A constrained optical flow algorithm, which combines the advantages of the unambiguous correspondence of feature point labeling and the flexible representation of optical flow computation, has been developed for face recognition from expressional face images. In this paper, we propose an integrated face recognition system that is robust against facial expressions by combining information from the computed intraperson optical flow and the synthesized face image in a probabilistic framework. Our experimental results show that the proposed system improves the accuracy of face recognition from expressional face images.
Chao-Kuei Hsieh, Shang-Hong Lai, Yung-Chang Chen
IEEE Trans. Image Process.2
2009 Learning partially-observed hidden conditional random fields for facial expression recognition
abstract
This paper describes a novel graphical model approach to seamlessly coupling and simultaneously analyzing facial emotions and the action units. Our method is based on the hidden conditional random fields (HCRFs) where we link the output class label to the underlying emotion of a facial expression sequence, and connect the hidden variables to the image frame-wise action units. As HCRFs are formulated with only the clique constraints, their labeling for hidden variables often lacks a coherent and meaningful configuration. We resolve this matter by introducing a partially-observed HCRF model, and establish an efficient scheme via Bethe energy approximation to overcome the resulting difficulties in training. For real-time applications, we also propose an online implementation to perform incremental inference with satisfactory accuracy.
Kai-Yueh Chang, Tyng-Luh Liu, Shang-Hong Lai
CVPR3
2009 Compensation of motion artifacts in MRI via graph-based optimization
abstract
In two-dimensional Fourier transform magnetic resonance imaging (2DFT-MRI), patient/object motion during the image acquisition results in ghosting and blurring. These motion artifacts are commonly considered as a major limitation in the MRI community. To correct these artifacts without resorting to additional navigator echoes, most existing methods perform image quality measure to estimate motion; but they may easily fail when the motion is large. Viewed as a blind image restoration problem where the motion point spread function (PSF) is unknown, state-of-the-art restoration algorithms can not be easily applied because they cannot handle a complex PSF kernel that has the same size as the image. To overcome these challenges, we propose a novel approach that exploits the image structure to segment the kernel into several fragments. Based on this kernel representation, determining a kernel fragment can be formulated as a binary optimization problem, where each binary variable represents whether a segment in MR signals is corrupted by a certain motion or not. We establish a graphical model for these variables and estimate the kernel by minimizing an energy functional associated with the model. Experimental results show that the proposed method can provide satisfactory compensation of motion artifacts even when large motions are involved in the MR images.
Tung-Ying Lee, Hong-Ren Su, Shang-Hong Lai, Ti-Chiun Chang
CVPR3
2009 A novel robust kernel for applications to images
abstract
Robustness is an essential issue to computer vision and pattern recognition in developing multimedia applications. In this work, we present a robust kernel approach that is highly robust against random noises and intra-class deformations. By incorporating the robust error function used in robust statistics together with a deformation-invariant distance measure, the derived robust kernel is shown to be insensitive to the influence of outliers and robust to intra-class deformations. In the experiments, we justify our robust kernel with different kernel machines with applications to handwritten digit recognition and data visualization on the USPS database.
Chia-Te Liao, Shang-Hong Lai
ICASSP2
2009 Fast structure-preserving image retargeting
abstract
Several different methods have been proposed for image/video retargeting while retaining the content. However, they sometimes produce some artifacts, such as ridge or structure twist. In this paper, we present a structure-preserving image resizing technique for the image retargeting applications. Based on the warping-based retargeting technique proposed by Wolf et al.[13], we propose an efficient and adaptive image resizing algorithm that preserves the content and image structure as best as possible. We first downsample the size of the original image by using bilinear interpolation. In order to preserve the content, we introduce the structure constraints derived from the line detection into the large linear system. Then, the mapping matrices are enlarged to the original size by joint-bilateral upsampling and the resized image can be produced to preserve the content and structure as best as possible. Most of the computation is on the low-resolution layer and therefore it can be very efficient. From our experiments, the proposed method can provide resized images with higher image quality and faster speed than that in [13].
Shu-Fan Wang, Shang-Hong Lai
ICASSP2
2009 Image deblurring by exploiting inherent bi-level regions
abstract
In this paper, we propose an image restoration framework for restoring an image degraded by unknown motion blur. Our approach takes advantage of inherent bi-level regions of an image to estimate a blur kernel. The framework contains three parts: bi-level region searching, initial blur kernel estimation and iterative maximum a posteriori (MAP) image restoration. Firstly, candidate bi-level regions are located around the detected corners. We use four image features to score each region and choose the best N regions for estimating an initial blur kernel. Finally, an alternating minimization algorithm is developed to iteratively refine both the blur kernel and the restored image. Experimental results of synthetic and real blurred images are shown to demonstrate the performance of the proposed algorithm.
Po-Hao Huang, Yu-Mo Lin, Hao-Liang Yang, Shang-Hong Lai
ICIP4
2009 A hierarchical image kernel with application to pedestrian identification for video surveillance
abstract
Video surveillance usually requires multiple cameras to monitor objects of interest, such as people. However, different appearances acquired from different cameras of the same people often make the construction of a robust individualized appearance model very challenging. In this paper, we present a kernel-based method that maps the bag-of-feature based image features to a hierarchical representation. The image comparison is performed through summing the weighted similarities of nodes in the hierarchical structure. The kernel is also proven to be positive-definite, making it valid for use in other kernel-based learning algorithms. In the experiments we show the classifier embedded with our kernel function is robust against view-point and scaling variations, and it is more accurate compared to other related approaches.
Chia-Te Liao, Shang-Hong Lai
ICIP2
2009 Face detection directly from h.264 compressed video with convolutional neural network
abstract
Human faces provide a useful cue in indexing video content. In this paper, we propose a novel face detection algorithm based on a convolutional neural network architecture that can rapidly detect human face regions in video sequences encoded by H.264/AVC. By detecting faces directly in the compressed domain, we use the discrete cosine transform (DCT) coefficients in H.264 intra coding as the feature vector for face detection, thus it is not necessary to carry out additional DCT transform during the encoding or decoding process. With the face detector inside the video encoding process, we can adjust the coding parameters adaptively and allocate more resources to the macroblocks corresponding to the face regions. Some experimental results of applying the face detector on the H.264 intra coded images are given to demonstrate the performance of the proposed algorithm.
ShinShan Zhuang, Shang-Hong Lai
ICIP2
2009 Efficient multiple virtual view generation based on reduced depth stereo image for advanced autostereoscopic displays
abstract
Recent development of autostereoscopic displays demands the synthesis of more virtual views in wider baseline, an inevitable trend for future 3DTV systems. However, current standard file formats, e.g. multi-view coding (MVC) and image plus depth, encounter challenging problems to meet the request. Therefore, we propose a new file format, called reduced depth stereo image (RDSI), which saves the color and depth images of the left view and the disoccluded regions in the right view. Based on RDSI, rendering virtual images from parallel viewpoints along the baseline can be simplified as view interpolation that is not only very efficient but also consistent in disocclusion regions. Through experiments on both simulated and real data, we demonstrate the superior performance with several quantitative assessments for the adopted file format and the rendering algorithm. The results suggest RDSI as a better choice to meet the demands for the online synthesis of many virtual views in wide baseline for advanced autostereoscopic displays.
Chia-Ming Cheng, Shu-Jyuan Lin, Shang-Hong Lai, Jenq Kuen Lee
ICME3
2009 Fast multi-reference motion estimation via statistical learning for H.264/AVC
abstract
In the H.264/AVC coding standard, motion estimation (ME) is allowed to use multiple reference frames to make full use of reducing temporal redundancy in a video sequence. Although it can further reduce the motion compensation errors, it introduces tremendous computational complexity as well. In this paper, we propose a statistical learning approach to reduce the computation involved in the multireference motion estimation. Some representative features are extracted in advance to build a learning model. Then, an off-line pre-classification approach is used to determine the best reference frame number according to the run-time features. It turns out that motion estimation will be performed only on the necessary reference frames based on the learning model. Experimental results show that the computation complexity is about three times faster than the conventional fast ME algorithm while the video quality degradation is negligible.
Chen-Kuo Chiang, Shang-Hong Lai
ICME2
2009 2D expression-invariant face recognition with constrained optical flow
abstract
Face recognition is one of the most intensively studied topics in computer vision and pattern recognition. A constrained optical flow algorithm, which combines the advantages of the unambiguous correspondence of feature point labeling and the flexible representation of optical flow computation, has been developed for face recognition from expressional face images. In this paper, we propose an integrated face recognition system that is robust against facial expressions by combining information from the computed intra-person optical flow and the synthesized face image in a probabilistic framework. Our experimental results show that the proposed system improves the accuracy of face recognition from expressional face images.
Chao-Kuei Hsieh, Shang-Hong Lai, Yung-Chang Chen
ICME2
2009 A novel color-context descriptor and its applications
abstract
This paper presents a new descriptor for object categorization and pedestrian identification applications. One of the main drawbacks of shape-context descriptor is its vulnerability and distinctness to color images. We propose a spherical descriptor that simultaneously adopts the spatial and color information as a discriminative representation. Based on the descriptor, this paper also contributes a bag-of-features framework to pedestrian identification for video surveillance. In contrast to the previous works, the proposed scheme does not require background subtraction stage. Thus the potential problems, such as the susceptibility to shadows and highlights from the background subtraction procedure, are avoided. Experiments validate the discriminant power of the proposed descriptor in object categorization on COIL- 100 database and pedestrian identification in surveillance videos.
Chia-Te Liao, Yu-Lin Wang, Shang-Hong Lai, Chiou-Ting Hsu
ICME3
2009 Integrated Expression-Invariant Face Recognition with Constrained Optical Flow
Chao-Kuei Hsieh, Shang-Hong Lai, Yung-Chang Chen
PSIVT2
2009 Robust surface reconstruction from defective point clouds by using orientation inference and volumetric regularization
abstract
Surface reconstruction is a critical stage in the 3D data acquisition and model creation system. Most existing reconstruction algorithms are designed for oriented data, i.e. point sets with surface normals. However, in some applications, explicit orientation information may not be available, e.g. Shape from Contour (SfC). Besides, the point sets recovered from images and camera calibration are typically noisy and contains defects, e.g. holes or non-uniform sampling. We present a robust method that achieves smooth surface approximation from unoriented and defective point sets by orientation inference and volumetric regularization.
Yi-Ling Chen 0004, Shang-Hong Lai, Tomoyuki Nishita
SIGGRAPH ASIA Sketches2
2009 Surface simplification by image retargeting
abstract
Surface simplification aims to reduce the complexity of a 3D model while maintaining a good approximation to the original model. In this work, we propose a novel combination of geometry images and content-aware image resizing to achieve efficient surface simplification. There are two main advantages to simplify surface based on geometry images. First, it is relatively simple to simplify the surface in the parameterized 2D space because the features of a 3D surface can be easily represented by the gradient energy. Second, the regularity and features on 3D surface can also be preserved without additional effort. The proposed retargeting algorithm performs well both on real images and 3D surface simplification.
Shu-Fan Wang, Yi-Ling Chen 0004, Chen-Kuo Chiang, Shang-Hong Lai
SIGGRAPH ASIA Sketches4
2009 A consensus sampling technique for fast and robust model fitting
Chia-Ming Cheng, Shang-Hong Lai
Pattern Recognit.2
2009 Fast JND-Based Video Carving With GPU Acceleration for Real-Time Video Retargeting
abstract
A recently developed image resizing technique, seam carving, has been proved to be a useful tool for content-adaptive spatially nonuniform image resizing with the purpose of optimal display on a screen of reduced resolution or different aspect ratio. In this paper, we present a fast algorithm for real-time content-aware video retargeting based on the improved seam carving method proposed in this paper. The proposed algorithm is designed to be highly parallelizable and suitable for running on a multicore architecture. First, two novel operators, i.e., seam update and seam split, are introduced to analyze an image for detecting the local and global seams with minimum costs very efficiently. With these operators, parallel processing can be achieved to determine multiple seams simultaneously. In addition, the saliency measure is extended with a just-noticeable-distortion model which makes the resized video more consistent with human perception. We demonstrate the efficiency of the above new components with a graphics processing unit (GPU) implementation. In addition, the proposed fast seam carving algorithm is extended for video retargeting. To the best of our knowledge, this is the first paper based on the seam carving method to achieve real-time video retargeting on a GPU. Experimental results on video sequences of various characteristics are demonstrated to show the superior performance of the proposed algorithm in comparison with the existing content-adaptive image/video resizing methods.
Chen-Kuo Chiang, Shu-Fan Wang, Yi-Ling Chen 0004, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.4
2009 Learning-Based Vertebra Detection and Iterative Normalized-Cut Segmentation for Spinal MRI
abstract
Automatic extraction of vertebra regions from a spinal magnetic resonance (MR) image is normally required as the first step to an intelligent spinal MR image diagnosis system. In this work, we develop a fully automatic vertebra detection and segmentation system, which consists of three stages; namely, AdaBoost-based vertebra detection, detection refinement via robust curve fitting, and vertebra segmentation by an iterative normalized cut algorithm. In order to produce an efficient and effective vertebra detector, a statistical learning approach based on an improved AdaBoost algorithm is proposed. A robust estimation procedure is applied on the detected vertebra locations to fit a spine curve, thus refining the above vertebra detection results. This refinement process involves removing the false detections and recovering the miss-detected vertebrae. Finally, an iterative normalized-cut segmentation algorithm is proposed to segment the precise vertebra regions from the detected vertebra locations. In our implementation, the proposed AdaBoost-based detector is trained from 22 spinal MR volume images. The experimental results show that the proposed vertebra detection and segmentation system can achieve nearly 98% vertebra detection rate and 96% segmentation accuracy on a variety of testing spinal MR images. Our experiments also show the vertebra detection and segmentation accuracies by using the proposed algorithm are superior to those of the previous representative methods. The proposed vertebra detection and segmentation system is proved to be robust and accurate so that it can be used for advanced research and application on spinal MR images.
Szu-Hao Huang, Yi-Hong Chu, Shang-Hong Lai, Carol L. Novak
IEEE Trans. Medical Imaging3
2009 Expression-Invariant Face Recognition With Constrained Optical Flow Warping
abstract
Face recognition is one of the most intensively studied topics in computer vision and pattern recognition, but few are focused on how to robustly recognize expressional faces with one single training sample per class. In this paper, we modify the regularization-based optical flow algorithm by imposing constraints on some given point correspondences to compute precise pixel displacements and intensity variations. By using the optical flow computed for the input expression variant face with respect to a reference neutral face image, we remove the expression from the face image by elastic image warping to recognize the subject with facial expression. Experimental validation is given to show that the proposed expression normalization algorithm significantly improves the accuracy of face recognition on expression variant faces.
Chao-Kuei Hsieh, Shang-Hong Lai, Yung-Chang Chen
IEEE Trans. Multim.2
2009 Creating MPU implicit surfaces from unoriented point sets with orientation inference
Yi-Ling Chen 0004, Shang-Hong Lai
Vis. Comput.2
2008 Silhouette-based camera calibration from sparse views under circular motion
abstract
In this paper, we propose a new approach to camera calibration from silhouettes under circular motion with minimal data. We exploit the mirror symmetry property and derive a common homography that relates silhouettes with epipoles under circular motion. With the epipoles determined, the homography can be computed from the frontier points induced by epipolar tangencies. On the other hand, given the homography, the epipoles can be located directly from the bi-tangent lines of silhouettes. With the homography recovered, the image invariants under circular motion and camera parameters can be determined. If the epipoles are not available, camera parameters can be determined by a low-dimensional search of the optimal homography in a bounded region. In the degenerate case, when the camera optical axes intersect at one point, we derive a closed-form solution for the focal length to solve the problem. By using the proposed algorithm, we can achieve camera calibration simply from silhouettes of three images captured under circular motion. Experimental results on synthetic and real images are presented to show its performance.
Po-Hao Huang, Shang-Hong Lai
CVPR2
2008 Efficient NCC-Based Image Matching in Walsh-Hadamard Domain
Wei-Hau Pan, Shou-Der Wei, Shang-Hong Lai
ECCV (3)3
2008 Estimating 3D Face Model and Facial Deformation from a Single Image Based on Expression Manifold Optimization
Shu-Fan Wang, Shang-Hong Lai
ECCV (1)2
2008 A hybrid motion estimation approach based on normalized cross correlation for video compression
abstract
In this paper we propose a new hybrid approach for block based motion estimation (ME) by adaptively using the normalized cross correlation (NCC) and sum of absolute differences (SAD) measures. We use the SAD value and gradient sum as the criterion to determine which similarity measure to be used for motion estimation. In general, using the NCC as the similarity measure in the motion estimation leads to more uniform residuals than those of using the SAD. This leads to larger DC terms and smaller AC terms, which yields less information loss after DCT quantization. However, NCC is not suitable for homogeneous regions since the best match may have a high NCC value but with large average gray level difference. Thus, we propose to alternatively use the SAD and NCC as the ME criterion for homogeneous and inhomogeneous blocks. Experimental results show the proposed hybrid motion estimation algorithms can provide superior PSNR and SSIM values than the traditional SAD-based ME method.
Wei-Hau Pan, Shou-Der Wei, Shang-Hong Lai
ICASSP3
2008 Expressional face image analysis with constrained optical flow
abstract
Face recognition is one of the most intensively studied topics in computer vision and pattern recognition. A constrained optical flow algorithm, which combines the advantages of the unambiguous correspondence of feature point labeling and the flexible representation of optical flow computation has been proposed in our pervious work. Facial expression normalization, from expressive to neutral facial images, based on optical flow analysis is discussed in this paper. In addition, we propose a new algorithm for nonnegative coefficient projection algorithm for projecting optical flow onto a facial expression subspace. Experimental validation is given to show that the proposed systems improve the accuracy of face and expression recognition on expressional face images.
Chao-Kuei Hsieh, Shang-Hong Lai, Yung-Chang Chen
ICME2
2008 Video-based face recognition based on view synthesis from 3D face model reconstructed from a single image
abstract
Most of the face recognition algorithms were proposed based on training numerous still examples of face images to accommodate different face variations, such as pose and illumination variations. However, it is not practical to collect lots of face images under different variations for each subject in a real authentication system. In this paper, we propose a novel face recognition system with only one single image for each individual in the training dataset. The proposed face recognition system applies the 3D face model reconstructed from the single face to synthesize different views for effectively training, thus leading to robustness against poses variations. The proposed system integrates the temporal face recognition results from the video in a probabilistic framework to make reliable decision when enough evidence is accumulated. In addition, it rejects imposters with the notion of locally linear embedding. The experiment results on FG-Net video database are shown to validate the effectiveness and reliability of the proposed algorithm.
Chia-Te Liao, Shufang Wang, Yun-Jen Lu, Shang-Hong Lai
ICME4
2008 Improved novel view synthesis from depth image with large baseline
abstract
In this paper, a new algorithm is developed for recovering the large disocclusion regions in depth image based rendering (DIBR) systems on 3DTV. For the DIBR systems, undesirable artifacts occur in the disocclusion regions by using the conventional view synthesis techniques especially with large baseline. Three techniques are proposed to improve the view synthesis results. The first is the preprocessing of the depth image by using the bilateral filter, which helps to sharpen the discontinuous depth changes as well as to smooth the neighboring depth of similar color, thus restraining noises from appearing on the warped images. Secondly, on the warped image of a new viewpoint, we fill the disocclusion regions on the depth image with the background depth levels to preserve the depth structure. For the color image, we propose the depth-guided exemplar-based image inpainting that combines the structural strengths of the color gradient to preserve the image structure in the restored regions. Finally, a trilateral filter, which simultaneous combines the spatial location, the color intensity, and the depth information to determine the weighting, is applied to enhance the image synthesis results. Experimental results are shown to demonstrate the superior performance of the proposed novel view synthesis algorithm compared to the traditional methods.
Chia-Ming Cheng, Shu-Jyuan Lin, Shang-Hong Lai, Jinn-Cherng Yang
ICPR3
2008 Image deblurring with blur kernel estimation from a reference image patch
abstract
In this paper, we propose a new approach for image deblurring from two images, non-blurred and blurred, in different poses by exploiting the co-existing planar object in both views. We focus on the problem of aligning the corresponding image patches, which are the co-existing planar object, in both images and propose an iterative two-stage algorithm for patch alignment and kernel estimation. In the first stage, we extend the intensity-based alignment method to find the geometric transformation between patches, and then the aligned image patches are used for blur kernel estimation in the second stage. These two stages are repeated until convergence. Furthermore, the proposed algorithm can also be used when the geometric relationship between the two images is a homography or an approximate homography, such as images from image mosaic. Experimental results on real images are given to demonstrate its performance.
Po-Hao Huang, Yu-Mo Lin, Shang-Hong Lai
ICPR3
2008 A novel robust kernel for appearance-based learning
abstract
Robustness is one of the most critical issues in the appearance-based learning strategies. In this work, we propose a novel kernel that is robust against data corruption for various visual learning problems. By incorporating a robust rho-function to relieve the influence of outliers, the proposed kernel is shown to be robust against various types of outliers. By incorporating the proposed kernel into different kernel-based approaches, we verify the robustness of the proposed kernel on various applications, including face recognition and data visualization. Our experiments on these visual learning problems demonstrate the superior performance of the proposed kernel compared to the conventional kernels.
Chia-Te Liao, Shang-Hong Lai
ICPR2
2008 Fast Intermode Decision Via Statistical Learning for H.264 Video Coding
Wei-Hau Pan, Chen-Kuo Chiang, Shang-Hong Lai
MMM3
2008 A Novel Motion Estimation Method Based on Normalized Cross Correlation for Video Compression
Shou-Der Wei, Wei-Hau Pan, Shang-Hong Lai
MMM3
2008 Real-Time Video Surveillance Based on Combining Foreground Extraction and Human Detection
Hui-Chi Zeng, Szu-Hao Huang, Shang-Hong Lai
MMM3
2008 Image-based three-dimensional model reconstruction for Chinese treasure - Jadeite Cabbage with Insects
Chia-Ming Cheng, Shu-Fan Wang, Chin-Hung Teng, Shang-Hong Lai
Comput. Graph.4
2008 Reconstructing 3D Shape, Albedo and Illumination from a Single Face Image
abstract
Abstract The morphable model has been employed to efficiently describe 3D face shape and the associated albedo with a reduced set of basis vectors. The spherical harmonics (SH) model provides a compact basis to well approximate the image appearance of a Lambertian object under different illumination conditions. Recently, the SH and morphable models have been integrated for 3D face shape reconstruction. However, the reconstructed 3D shape is either inconsistent with the SH bases or obtained just from landmarks only. In this work, we propose a geometrically consistent algorithm to reconstruct the 3D face shape and the associated albedo from a single face image iteratively by combining the morphable model and the SH model. The reconstructed 3D face geometry can uniquely determine the SH bases, therefore the optimal 3D face model can be obtained by minimizing the error between the input face image and a linear combination of the associated SH bases. In this way, we are able to preserve the consistency between the 3D geometry and the SH model, thus refining the 3D shape reconstruction recursively. Furthermore, we present a novel approach to recover the illumination condition from the estimated weighting vector for the SH bases in a constrained optimization formulation independent of the 3D geometry. Experimental results show the effectiveness and accuracy of the proposed face reconstruction and illumination estimation algorithm under different face poses and multiple‐light‐source illumination conditions.
Shu-Fan Wang, Shang-Hong Lai
Comput. Graph. Forum2
2008 Fast Optimal Motion Estimation Based on Gradient-Based Adaptive Multilevel Successive Elimination
abstract
In this paper, we propose a fast and optimal solution for block motion estimation based on an adaptive multilevel successive elimination algorithm. This algorithm is accomplished by applying a modified multilevel successive elimination algorithm (SEA) with the elimination order determined by the sum of the gradient magnitudes of each subblock and the elimination process terminated by comparing the above sum with a threshold. In addition a fast approximate motion estimation method and the accumulated distortion scheme are employed to make the proposed algorithm even more efficiently. Experimental results show that the proposed adaptive multilevel successive elimination strategy (AdaMSEA) algorithm significantly outperforms other previous optimal motion estimation algorithms, including SEA, MSEA, and FGSE on a wide variety of video sequences. Finally, we modify the proposed AdaMSEA to an approximate motion estimation algorithm to achieve very fast computational speed, and the experimental results show superior performance of this approximate algorithm over some fast motion estimation algorithms.
Shao-Wei Liu, Shou-Der Wei, Shang-Hong Lai
IEEE Trans. Circuits Syst. Video Technol.3
2008 Fast Template Matching Based on Normalized Cross Correlation With Adaptive Multilevel Winner Update
abstract
In this paper, we propose a fast pattern matching algorithm based on the normalized cross correlation (NCC) criterion by combining adaptive multilevel partition with the winner update scheme to achieve very efficient search. This winner update scheme is applied in conjunction with an upper bound for the cross correlation derived from Cauchy-Schwarz inequality. To apply the winner update scheme in an efficient way, we partition the summation of cross correlation into different levels with the partition order determined by the gradient energies of the partitioned regions in the template. Thus, this winner update scheme in conjunction with the upper bound for NCC can be employed to skip unnecessary calculation. Experimental results show the proposed algorithm is very efficient for image matching under different lighting conditions.
Shou-Der Wei, Shang-Hong Lai
IEEE Trans. Image Process.2
2007 Camera Calibration from Silhouettes Under Incomplete Circular Motion with a Constant Interval Angle
Po-Hao Huang, Shang-Hong Lai
ACCV (1)2
2007 Efficient Normalized Cross Correlation Based on Adaptive Multilevel Successive Elimination
Shou-Der Wei, Shang-Hong Lai
ACCV (1)2
2007 A Robust Kernel Based on Robust ρ-Function
abstract
Noise-resistance capability is a very important issue to signal processing systems as well as machine learning applications. In this work, we present a new kernel that is highly robust against outliers and random noises. By incorporating a robust ρ-function into the distance metric, the derived robust kernel was shown to be very insensitive to the influence of outlier elements. In the experiments, we show that the proposed kernel brought significant improvement to the support vector machines (SVM) classifier in face recognition accuracy and outperformed several traditional kernels for corrupted data. We also applied our kernel to the kernel principal component analysis (PCA) and evaluate the efficiency in recovering contaminated face images. Experiments show our robust kernel also brings benefits in noise-reduction applications.
Chia-Te Liao, Shang-Hong Lai
ICASSP (2)2
2007 Adaptive Multi-Reference Downhill Simplex Search Based on Spatial-Temporal Motion Smoothness Criterion
abstract
Multi-reference frame motion estimation improves the accuracy of motion compensation in video coding. However, it also increases computational complexity dramatically. In this paper, we propose a different approach for multi-reference motion estimation via downhill simplex search. Additionally, an adaptive reference frame selection algorithm is developed based on spatial and temporal smoothness of motion vectors. We first apply single-reference downhill simplex search to the previous frame. Then, temporal smoothness of motion vectors in collocated blocks is calculated to decide the number of reference frames to be included for motion estimation. Spatial smoothness of motion vectors in the neighboring blocks is used as a criterion for termination. Experimental results show that the proposed algorithm provides better PSNR than that of original multi-reference downhill simplex search in all testing sequences with similar computational speed. In addition, it outperforms several representative single-reference frame block matching methods in terms of estimation speed and coding quality.
Wei-Hau Pan, Chen-Kuo Chiang, Shang-Hong Lai
ICASSP (1)3
2007 Fast Template Matching by Applying Winner-Update on Walsh-Hadamard Domain
abstract
Fast template matching is strongly demanded for many practical applications related to computer vision and image processing. In this paper, we propose a fast template matching method by applying the winner-update strategy on the Walsh-Hadamard domain. By taking advantage of the nice energy packing property of the Walsh-Hadamard transformation, we can just apply the winner-update process with a small number of Walsh-Hadamard coefficients to reduce the computational burden for template matching in an image. Experimental results demonstrate the efficiency and robustness of the proposed template matching algorithm under different noise levels.
Shou-Der Wei, Shao-Wei Liu, Shang-Hong Lai
ICASSP (1)3
2007 Adaptive Foreground Object Extraction for Real-Time Video Surveillance with Lighting Variations
abstract
In this paper we present an adaptive foreground object extraction algorithm for real-time video surveillance. The proposed algorithm improves the previous Gaussian mixture background models (GMMs) by applying a two-stage foreground/background classification procedure to remove the undesirable subtraction results due to shadow, automatic white balance, and sudden illumination change. The traditional background subtraction technique usually cannot work well for situations with lighting variations in the scene. In the proposed two-stage classification, an adaptive classifier is applied to the foreground pixels in a pixel-wise manner based on the normalized color and brightness gain information. Secondly, the remaining foreground candidate pixels are grouped into regions and the corresponding background regions are compared to check if they are foreground regions. Experimental results on some real surveillance video are shown to demonstrate the robustness of the proposed adaptive foreground extraction algorithm under a variety of different environments with lighting variations.
Hui-Chi Zeng, Shang-Hong Lai
ICASSP (1)2
2007 Efficient Intra Mode Selection using Image Structure Tensor for H.264/AVC
abstract
Intra mode decision and motion estimation for spatial and temporal prediction play important roles for achieving high video compression ratio in the latest video coding standard H.264/AVC. However, both components take most of the computational cost in the video encoding process. In this paper, we propose an efficient intra mode prediction algorithm based on image structure tensor analysis. The image structure tensor can provide the local image structure information for making the intra-mode decision. We show significant reduction of the computation time with negligible video quality degradation for H.264 video encoding by implementation into JM reference program .
Chiuan Hwang, ShinShan Zhuang, Shang-Hong Lai
ICIP (5)3
2007 Iterative Blind Image Motion Deblurring via Learning a No-Reference Image Quality Measure
abstract
In this paper, we propose a learning-based image restoration algorithm for restoring images degraded by uniform motion blurs. The motion blur parameters are first approximately estimated from the robust global motion estimation result. Then, we present a novel framework to refine the image restoration iteratively based on recursively adjusting the motion blur parameters for image restoration to achieve the best image quality measure. Note that a no-reference image quality assessment model is learned by training a RBF neural network from a collection of representative training images simulated with different motion blurs. Experimental results blurred on real videos are given to demonstrate the performance of the proposed blind motion deblurring algorithm.
Wen-Hao Lee, Shang-Hong Lai, Chia-Lun Chen
ICIP (4)2
2007 Reproducibility Analysis of Event-Related fMRI Experiments Using Laguerre Polynomials
Hong-Ren Su, Michelle Liou, Philip E. Cheng, John A. D. Aston, Shang-Hong Lai
ICONIP (1)5
2007 Temporally Integrated Pedestrian Detection from Non-stationary Video
Chi-Jiunn Wu, Shang-Hong Lai
MMM (1)2
2007 A Partition-of-Unity Based Algorithm for Implicit Surface Reconstruction Using Belief Propagation
abstract
In this paper, we propose a new algorithm for the fundamental problem of reconstructing surfaces from a large set of unorganized 3D data points. The local shapes of the surface are recovered by variational implicit surface represented as a weighted combination of radial basis functions. The variational implicit patches are then combined together to form the overall surface via a set of blending functions, which is also referred to as the partition-of-unity method. The reconstruction algorithm first partitions the input point set by octree subdivision and surface normal estimation is performed so as to orientate the local variational implicit patches. A new graph optimization scheme based on the belief propagation framework is proposed to determine the global consistent orientation for the entire set of data points. To achieve multi-scale reconstruction, we propose a novel progressive reconstruction algorithm which utilizes the Schur complement formula to reduce the computational cost of iteratively updating the radial basis function coefficients. Finally, we demonstrate the performance of the proposed algorithm by showing experimental results on some real-world 3D data sets.
Yi-Ling Chen 0004, Shang-Hong Lai
Shape Modeling International2
2007 Hybrid image matching combining Hausdorff distance with normalized gradient matching
Chyuan-Huei Thomas Yang, Shang-Hong Lai, Long-Wen Chang
Pattern Recognit.2
2006 Adaptive Object Tracking with Online Statistical Model Update
KaiYeuh Chang, Shang-Hong Lai
ACCV (2)2
2006 Efficient 3D Face Reconstruction from a Single 2D Image by Combining Statistical and Geometrical Information
Shu-Fan Wang, Shang-Hong Lai
ACCV (2)2
2006 Contour-Based Structure from Reflection
abstract
In this paper, we propose a novel contour-based algorithm for 3D object reconstruction from a single uncalibrated image acquired under the setting of two plane mirrors. With the epipolar geometry recovered from the image and the properties of mirror reflection, metric reconstruction of an arbitrary rigid object is accomplished without knowing the camera parameters and the mirror poses. For this mirror setup, the epipoles can be estimated from the correspondences between the object and its reflection, which can be established automatically from the tangent lines of their contours. By using the property of mirror reflection as well as the relationship between the mirror plane normal with the epipole and camera intrinsic, we can estimate the camera intrinsic, plane normals and the orientation of virtual cameras. The positions of the virtual cameras are determined by minimizing the distance between the object contours and the projected visual cone for a reference view. After the camera parameters are determined, the 3D object model is constructed via the image-based visual hulls (IBVH) technique. The 3D model can be refined by integrating the multiple models reconstructed from different views. The main advantage of the proposed contour-based Structure from Reflection (SfR) algorithm is that it can achieve metric reconstruction from an uncalibrated image without feature point correspondences. Experimental results on synthetic and real images are presented to show its performance.
Po-Hao Huang, Shang-Hong Lai
CVPR (1)2
2006 Automatic Multi-Layer Red-Eye Detection
abstract
Red-eye is a frequently encountered problem caused by flash reflection bouncing back into the camera from a person's retina. In this paper, we propose a two-stage automatic red-eye detection method. At the first stage, a series of heuristic filters are used to rapidly remove impossible regions according to constraints on color, smoothness, and size. A multi-layer process is applied at this stage to prevent the influence of surrounding redness of the eye. Moreover, at each layer, the approximate size of pupil can be decided automatically without multi-scaling. At the second stage, an SVM classifier that was trained for eye detection is applied to confirm the remaining candidate regions. Some experimental results on real images show the performance of the propose algorithm.
Po-Hao Huang, Yu-Chieh Chien, Shang-Hong Lai
ICIP3
2006 Fast and Optimal Block Motion Estimation via Adaptive Successive Elimination
abstract
In this paper we propose a fast and optimal solution for block motion estimation based on an adaptive successive elimination algorithm (SEA). We first apply an fast approximate method likes adaptive rood pattern search (ARPS) method to obtain a good initial motion vector as well as a tight initial bound of distortion measure to be used in SEA. Then, we apply the multi-level SEA with the elimination order determined by the sum of the gradient magnitudes of each sub-block. Experimental results of applying different motion estimation methods to video compression show that the proposed adaptive MSEA method can achieve the same PSNR with full search with the speed faster than the diamond search (DS).
Shou-Der Wei, Shao-Wei Liu, Shang-Hong Lai
ICIP3
2006 Fast Multi-Reference Frame Motion Estimation via Downhill Simplex Search
abstract
Multi-reference frame motion estimation improves the accuracy of motion compensation in video compression, but it also dramatically increases computational complexity. Based on tracing motion vector trajectories, fast approximated motion estimation results can be obtained for multi-reference frames. In this paper, we extend the downhill simplex search to multiple reference frames and propose several enhanced schemes to improve its efficiency and accuracy. Experimental results show that the proposed algorithm outperforms several representative single-reference frame block matching methods
Chen-Kuo Chiang, Shang-Hong Lai
ICME2
2006 Modified Winner Update with Adaptive Block Partition for Fast Motion Estimation
abstract
Motion estimation (ME) plays an important role in video compression. Block-based ME has been adopted in most video compression standards due to its efficiency. In this paper, we propose a novel and fast block-based ME algorithm based on applying the modified winner-update scheme in conjunction with the adaptive partition order of macroblock. The partition order is determined from the block gradient distribution. Experimental results show the proposed algorithm achieves the optimal motion estimation very efficiently
Shou-Der Wei, Shao-Wei Liu, Shang-Hong Lai
ICME3
2006 Learning-Based Interactive Video Retrieval System
abstract
This paper presents an interactive video event retrieval system based on improved adaboost learning. This system consists of three main steps. Firstly, a long video sequence is partitioned into several video clips by using a distribution-based approach instead of detecting shot transition boundaries. Secondly, audiovisual features (i.e., color, motion and audio features) are extracted from video sequences for video clip representation. Finally, the modified AdaBoost learning algorithm is employed for interactive video retrieval with relevance feedback. This AdaBoost learning algorithm differs from conventional AdaBoost learning methods mainly in the selection of paired video features for the weak classifiers. Experimental results show improved performance of video retrieval by using the proposed system
Chi-Jiunn Wu, Hui-Chi Zeng, Szu-Hao Huang, Shang-Hong Lai
ICME4
2006 A robust real-time video stabilization algorithm
Hung-Chang Chang, Shang-Hong Lai, Kuang-Rong Lu
J. Vis. Commun. Image Represent.2
2006 Improved AdaBoost-based image retrieval with relevance feedback via paired feature learning
Szu-Hao Huang, Qi-Jiunn Wu, Shang-Hong Lai
Multim. Syst.3
2006 Robust and Efficient Image Alignment Based on Relative Gradient Matching
abstract
In this paper, we present a robust image alignment algorithm based on matching of relative gradient maps. This algorithm consists of two stages; namely, a learning-based approximate pattern search and an iterative energy-minimization procedure for matching relative image gradient. The first stage finds some candidate poses of the pattern from the image through a fast nearest-neighbor search of the best match of the relative gradient features computed from training database of feature vectors, which are obtained from the synthesis of the geometrically transformed template image with the transformation parameters uniformly sampled from a given transformation parameter space. Subsequently, the candidate poses are further verified and refined by matching the relative gradient images through an iterative energy- minimization procedure. This approach based on the matching of relative gradients is robust against nonuniform illumination variations. Experimental results on both simulated and real images are shown to demonstrate superior efficiency and robustness of the proposed algorithm over the conventional normalized correlation method.
S.-D. Wei, Shang-Hong Lai
IEEE Trans. Image Process.2
2005 An adaptive window width/center adjustment system with online training capabilities for MR images
Shang-Hong Lai
Artif. Intell. Medicine1
2005 Accurate optical flow computation under non-uniform brightness variations
Chin-Hung Teng, Shang-Hong Lai, Yung-Sheng Chen, Wen-Hsing Hsu
Comput. Vis. Image Underst.2
2004 An accurate and adaptive optical flow estimation algorithm
Chin-Hung Teng, Shang-Hong Lai, Yung-Sheng Chen, Wen-Hsing Hsu
ICIP2
2004 A robust and efficient video stabilization algorithm
abstract
The acquisition of digital video usually suffers from undesirable camera jitter due to unstable random camera motion, which is produced by a hand-held camera or a camera in a vehicle moving on a non-smooth road or terrain. We propose a real-time robust video stabilization algorithm to remove undesirable motion jitter and produce a stabilized video. We first compute the optical flow between successive frames, followed by estimating the camera motion by fitting the computed optical flow field to a simplified affine motion model with a trimmed least squares method. Then, the computed camera motions are smoothed temporally to reduce the motion vibrations by using a regularization method. Finally, we transform all frames of the video based on the original and smoothed motions to obtain a stabilized video. Experimental results are given to demonstrate the stabilization performance and the efficiency of the proposed algorithm.
Hung-Chang Chang, Shang-Hong Lai, Kuang-Rong Lu
ICME2
2004 Real-Time Face Detection in Color Video
abstract
In this paper, we propose a novel and fast face detection algorithm for detecting face in color video sequences. This algorithm can be integrated into a real-time surveillance or a video retrieval in the design of the algorithm. A set of multiresolution Haar wavelet coefficients pairs is selected by the proposed learning algorithm to determine if a particular region is a face. We apply an ID3-like balanced decision tree for the wavelet coefficients quantization, to reduce the quantization error. For each pair of quantized features, we estimate the associated conditional joint probability density function from a large set of face and non-face training data. Then, we compute the Kullback Leibler (KL) distance to measure the discrimination between the face and non-face conditional density functions for each feature pair. The feature pairs with larger KL-distance are selected as the feature candidates. It is an effective feature dimension reduction method and helps to speedup the Adaboost training algorithm when considering the spatial relationship between all coefficient pairs. Aided by an automatic skin color judgment method and a Gaussian face location model both in temporal and spatial domain, the experiments show that the proposed algorithm runs faster than 4 times the video rate with good detection accuracy.
Szu-Hao Huang, Shang-Hong Lai
MMM2
2004 Robust camera motion estimation and classification for video analysis
abstract
Camera motion estimation is very important for indexing and retrieving video information. In this paper, we propose a robust camera motion estimation and classification algorithm. Our camera motion estimation algorithm consists of optical flow estimation, iterative RANSAC (RANdom SAmple Consensus) multiple motion estimation, and long-term camera motion estimation through a shortest-path search. In this approach, we first estimate multiple global affine motions from the computed optical flow field for every frame in the video sequence. Then, the long-term camera motion is determined from searching a shortest path in a graph of cascaded nodes of global motions. After the camera motion is determined for the whole video, we apply an artificial neural network to classify the camera motion type. This neural network is trained from a large set of different types of camera motion data. We show accurate camera motion classification results through experiments on real videos.
Hung-Chang Chang, Shang-Hong Lai
VCIP2
2004 Computation of optical flow under non-uniform brightness variations
Shang-Hong Lai
Pattern Recognit. Lett.1
2002 Robust face matching under different lighting conditions
abstract
We focus on the problem of comparing face images under different lighting conditions. A new and robust face image matching algorithm is developed based on an accumulated consistency measure of corresponding normalized gradients at face contour locations between two face images. The proposed new image matching approach is motivated by the characteristics of high image gradient along the face contour. To amend the matching problem due to lighting changes between face images, we define a new consistency measure to be the inner product between two normalized gradient vectors at the corresponding locations in two images. The normalized gradient is obtained by dividing the computed gradient vector by a maximal gradient magnitude in a local neighborhood centered at the pixel of computation. Then we compute the summation of the individual consistency measure from normalized gradients at all the contour pixels to be the robust matching measure between two face images. To alleviate the problem due to shadow and intensity saturation, we introduce an intensity weighting function for each individual consistency measure to form a weighted consistency measure. We test the proposed image matching algorithm on the Yale Face Database, which contains 15 persons captured under three very different lighting conditions. The proposed method can achieve satisfactory recognition results in the experiments.
Chyuan-Huei Thomas Yang, Shang-Hong Lai, Long-Wen Chang
ICME (2)2
2000 A hierarchical neural network algorithm for robust and automatic windowing of MR images
Shang-Hong Lai
Artif. Intell. Medicine1
2000 Robust Image Matching under Partial Occlusion and Spatially Varying Illumination Change
Shang-Hong Lai
Comput. Vis. Image Underst.1
1999 Robust and Efficient Image Alignment with Spatially Varying Illumination Models
abstract
Image alignment is one of the most important task in computer vision. In this paper, we explicitly model spatial illumination variations by low-order polynomial functions in an energy minimization framework. Data constraints for the alignment and illumination parameters are derived from the first-order Taylor approximation of a generalized brightness assumption. We formulate the parameter estimation problem in a weighted least-square framework by using the influence function from robust estimation to derive an iterative re-weighted least-square algorithm. A dynamic weighting scheme, which combines the factors from influence function, consistency of image gradients and nonlinear image intensity sensing, is used to improve the robustness of the image matching. In addition, a constraint sampling scheme and an estimation-warping alternative strategy are used in the proposed algorithm to improve its efficiency and accuracy. Experimental results are shown to demonstrate the robustness, efficiency and accuracy of the algorithm.
Shang-Hong Lai
CVPR1
1999 Efficient hybrid search for visual reconstruction problems
Shang-Hong Lai, Baba C. Vemuri
Image Vis. Comput.1
1999 A new variational shape-from-orientation approach to correcting intensity inhomogeneities in magnetic resonance images
Shang-Hong Lai
Medical Image Anal.1
1998 Reliable and Efficient Computation of Optical Flow
Shang-Hong Lai, Baba C. Vemuri
Int. J. Comput. Vis.1
1998 Fast numerical algorithms for fitting multiresolution hybrid shape models to brain MRI
Baba C. Vemuri, Yanlin Guo, Shang-Hong Lai, Christiana Morison Leonard
Medical Image Anal.3
1997 Fast numerical algorithms for fitting multiresolution hybrid shape models to brain MRI
Baba C. Vemuri, Yanlin Guo, Christiana Morison Leonard, Shang-Hong Lai
Medical Image Anal.4
1997 Physically Based Adaptive Preconditioning for Early Vision
abstract
Several problems in early vision have been formulated in the past in a regularization framework. These problems, when discretized, lead to large sparse linear systems. In this paper, we present a novel physically based adaptive preconditioning technique which can be used in conjunction with a conjugate gradient algorithm to dramatically improve the speed of convergence for solving the aforementioned linear systems. A preconditioner, based on the membrane spline, or the thin plate spline, or a convex combination of the two, is termed a physically based preconditioner for obvious reasons. The adaptation of the preconditioner to an early vision problem is achieved via the explicit use of the spectral characteristics of the regularization filter in conjunction with the data. This spectral function is used to modulate the frequency characteristics of a chosen wavelet basis, and these modulated values are then used in the construction of our preconditioner. We present the preconditioner construction for three different early vision problems namely, the surface reconstruction, the shape from shading, and the optical flow computation problems. Performance of the preconditioning scheme is demonstrated via experiments on synthetic and real data sets.
Shang-Hong Lai, Baba C. Vemuri
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 A Fast Gibbs Sampler for Synthesizing Constrained Fractals
abstract
It is well known that the spatial frequency spectrum of membrane and thin plate splines exhibit self-affine characteristics and, hence, behave as fractals. This behavior was exploited in generating the constrained fractal surfaces, which were generated by using a Gibbs sampler algorithm in the work of Szeliski and Terzopoulos (1989). The algorithm involves locally perturbing a constrained spline surface with white noise until the spline surface reaches an equilibrium state. We introduce a fast generalized Gibbs sampler that combines two novel techniques, namely, a preconditioning technique in a wavelet basis for constraining the splines and a perturbation scheme in which, unlike the traditional Gibbs sampler, all sites (surface nodes) that do not share a common neighbor are updated simultaneously. In addition, we demonstrate the capability to generate arbitrary order fractal surfaces without resorting to blending techniques. Using this fast Gibbs sampler algorithm, we demonstrate the synthesis of realistic terrain models from sparse elevation data.
Baba C. Vemuri, Chhandomay Mandal, Shang-Hong Lai
IEEE Trans. Vis. Comput. Graph.3
1994 An O(N) iterative solution to the Poisson equation in low-level vision problems
abstract
In this paper, we present a novel iterative numerical solution to the Poisson equation whose solution is needed in a variety of low-level vision problems. Our algorithm is an O(N) (N being the number of discretization points) iterative technique and does not make any assumptions on the shape of the input domain unlike the polyhedral domain assumption in the proof of convergence of multigrid techniques. We present two major results namely, a generalized version of the capacitance matrix theorem and a theorem on O(N) convergence of the alternating direction implicit method (ADI) used in our algorithm. Using this generalized theorem, we express the linear system corresponding to the discretized Poisson equation as a Lyapunov and a capacitance matrix equation. The former is solved using the ADI method while the solution to the later is obtained using a modified bi-conjugate gradient algorithm. We demonstrate the algorithm performance on synthesized data for the surface reconstruction and the SFS problems.>
Shang-Hong Lai, Baba C. Vemuri
CVPR1
1992 A Generalized Depth Estimation Algorithm with a Single Image
abstract
A depth estimation algorithm proposed by A.P. Pentland (1987) is generalized. In the proposed algorithm, the raw image data in the vicinity of the edge is used to estimate the depth from defocus. Since no differentiation operation on the image data is required before the optimization process, the method is less sensitive to the noise disturbance of measurements. Furthermore, the edge orientation that was critical in Pentland's approach will not be required in the case. This algorithm is then applied to synthetic images containing various amounts of noise to test its performance. Experimental results indicate that the depth estimation errors are kept within 5% of true values on the average when it is applied to real images.>
Shang-Hong Lai, Chang-Wu Fu, Shyang Chang
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Estimation of 3-D translational motion parameters via hadamard transform
Shang-Hong Lai, Shyang Chang
Pattern Recognit. Lett.1