EDBT 2026 Demo / reviewers in the wild / expert
Chiou-Shann Fuh
dblp:95/2319
· DBLP profile ↗
52ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0002-6174-2556ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment ModelabstractThis study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multitask learning. The 3QUEST metrics, namely Speech-MOS (S-MOS), Noise-MOS (N-MOS), and General-MOS (G-MOS), are the assessment targets. The pretrained MOSA-Net model is utilized to estimate three pseudo labels: perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), and speech distortion index (SDI). Multi-task learning is then employed to train MTQ-Net by combining a supervised loss (derived from the difference between the estimated score and the ground-truth label) and a semi-supervised loss (derived from the difference between the estimated score and the pseudo label), where the Huber loss is employed as the loss function. Experimental results first demonstrate the advantages of MPL compared to training a model from scratch and using a direct knowledge transfer mechanism. Second, the benefit of the Huber loss for improving the predictive ability of MTQ-Net is verified. Finally, the MTQ-Net with the MPL approach exhibits higher overall predictive power compared to other SSL-based speech assessment models. Ryandhimas E. Zezario, Bo-Ren Bai, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 3 |
| 2024 | A Study On Incorporating Whisper For Robust Speech AssessmentabstractThis research introduces an enhanced version of the multi-objective speech assessment model–MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper’s embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems. Ryandhimas E. Zezario, Yuwen Chen 0006, Szu-Wei Fu, Yu Tsao 0001, Hsin-Min Wang, Chiou-Shann Fuh |
ICME | 6 |
| 2024 | Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 3 |
| 2023 | Contrastive Feature Decoupling for Weakly-Supervised Disease Detection
Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh |
MICCAI (5) | 3 |
| 2023 | Deep Learning-Based Non-Intrusive Multi-Objective Speech Assessment Model With Cross-Domain FeaturesabstractThis study proposes a cross-domain multi-objective speech assessment model, called MOSA-Net, which can simultaneously estimate the speech quality, intelligibility, and distortion assessment scores of an input speech signal. MOSA-Net comprises a convolutional neural network and bidirectional long short-term memory architecture for representation extraction, and a multiplicative attention layer and a fully connected layer for each assessment metric prediction. Additionally, cross-domain features (spectral and time-domain features) and latent representations from self-supervised learned (SSL) models are used as inputs to combine rich acoustic information to obtain more accurate assessments. Experimental results show that in both seen and unseen noise environments, MOSA-Net can improve the linear correlation coefficient (LCC) scores in perceptual evaluation of speech quality (PESQ) prediction, compared to Quality-Net, an existing single-task model for PESQ prediction, and improve LCC scores in short-time objective intelligibility (STOI) prediction, compared to STOI-Net, an existing single-task model for STOI prediction. Moreover, MOSA-Net can be used as a pre-trained model to be effectively adapted to an assessment model for predicting subjective quality and intelligibility scores with a limited amount of training data. Experimental results show that MOSA-Net can improve LCC scores in mean opinion score (MOS) predictions, compared to MOS-SSL, a strong single-task model for MOS prediction. We further adopt the latent representations of MOSA-Net to guide the speech enhancement (SE) process and derive a quality-intelligibility (QI)-aware SE (QIA-SE) approach. Experimental results show that QIA-SE outperforms the baseline SE system with improved PESQ scores in both seen and unseen noise environments over a baseline SE model. Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Self-supervised Sparse Representation for Video Anomaly Detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ECCV (13) | 4 |
| 2022 | MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibility Prediction Model for Hearing Aids
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 3 |
| 2022 | MTI-Net: A Multi-Target Speech Intelligibility Prediction ModelabstractRecently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention.Many studies report that these DL-based models yield satisfactory assessment performance and good flexibility, but their performance in unseen environments remains a challenge.Furthermore, compared to quality scores, fewer studies elaborate deep learning models to estimate intelligibility scores.This study proposes a multi-task speech intelligibility prediction model, called MTI-Net, for simultaneously predicting human and machine intelligibility measures.Specifically, given a speech utterance, MTI-Net is designed to predict human subjective listening test results and word error rate (WER) scores.We also investigate several methods that can improve the prediction performance of MTI-Net.First, we compare different features (including low-level features and embeddings from self-supervised learning (SSL) models) and prediction targets of MTI-Net.Second, we explore the effect of transfer learning and multi-tasking learning on training MTI-Net.Finally, we examine the potential advantages of fine-tuning SSL embeddings.Experimental results demonstrate the effectiveness of using cross-domain features, multi-task learning, and fine-tuning SSL embeddings.Furthermore, it is confirmed that the intelligibility and WER scores predicted by MTI-Net are highly correlated with the ground-truth scores. Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2021 | Learning Unsupervised Metaformer for Anomaly DetectionabstractAnomaly detection (AD) aims to address the task of classification or localization of image anomalies. This paper addresses two pivotal issues of reconstruction-based approaches to AD in images, namely, model adaptation and reconstruction gap. The former generalizes an AD model to tackling a broad range of object categories, while the latter provides useful clues for localizing abnormal regions. At the core of our method is an unsupervised universal model, termed as Metaformer, which leverages both meta-learned model parameters to achieve high model adaptation capability and instance-aware attention to emphasize the focal regions for localizing abnormal regions, i.e., to explore the reconstruction gap at those regions of interest. We justify the effectiveness of our method with SOTA results on the MVTec AD dataset of industrial images and highlight the adaptation flexibility of the universal Metaformer with multi-class and few-shot scenarios. Jhih-Ciang Wu, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ICCV | 3 |
| 2021 | One-class anomaly detection via novelty normalization
Jhih-Ciang Wu, Sherman Lu, Chiou-Shann Fuh, Tyng-Luh Liu |
Comput. Vis. Image Underst. | 3 |
| 2021 | TsanKit: artificial intelligence for solder ball head-in-pillow defect inspection
Ting-Chen Tsan, Teng-Fu Shih, Chiou-Shann Fuh |
Mach. Vis. Appl. | 3 |
| 2018 | Use 3D Convolutional Neural Network to Inspect Solder Ball Defects
Bing-Jhang Lin, Ting-Chen Tsan, Tzu-Chia Tung, You-Hsien Lee, Chiou-Shann Fuh |
ICONIP (1) | 5 |
| 2018 | Solder Ball 3D Reconstruction with X-Ray Images Using Filtered Back ProjectionabstractWe aim to reconstruct the 3D model of solder balls on the PCB (Printed Circuit Board) according to many projections of X-ray images of those solder balls. We mainly use a tomographic reconstruction algorithm FBP (Filtered Back Projection) to obtain the cross-section of different levels without breaking the PCB and inspect their inner structure. We also optimize our execution times and compare our results with Volume Graphics. Ting-Chen Tsan, Bing-Jhang Lin, You-Hsien Lee, Tzu-Chia Tung, Chiou-Shann Fuh |
TENCON | 5 |
| 2018 | Bone-conducted speech enhancement using deep denoising autoencoder
Hung-Ping Liu, Yu Tsao 0001, Chiou-Shann Fuh |
Speech Commun. | 3 |
| 2017 | Pin Defect Inspection with X-ray Images
Hsien-Pei Kao, Tzu-Chia Tung, Hong-Yi Chen, Cheng-Shih Wong, Chiou-Shann Fuh |
ISNN (2) | 5 |
| 2014 | Attributed hypergraph matching on a Riemannian manifold
Jung Ming Wang, Sei-Wang Chen, Chiou-Shann Fuh |
Mach. Vis. Appl. | 3 |
| 2012 | Generation of Environmental Representation of a Large Indoor Parking Lot
Jung Ming Wang, Chih-Fan Hsu, Sei-Wang Chen, Chiou-Shann Fuh |
ICONIP (2) | 4 |
| 2011 | Coregulation of transcription factors and microRNAs in human transcriptional regulatory networkabstractBACKGROUND: MicroRNAs (miRNAs) are small RNA molecules that regulate gene expression at the post-transcriptional level. Recent studies have suggested that miRNAs and transcription factors are primary metazoan gene regulators; however, the crosstalk between them still remains unclear. METHODS: We proposed a novel model utilizing functional annotation information to identify significant coregulation between transcriptional and post-transcriptional layers. Based on this model, function-enriched coregulation relationships were discovered and combined into different kinds of functional coregulation networks. RESULTS: We found that miRNAs may engage in a wider diversity of biological processes by coordinating with transcription factors, and this kind of cross-layer coregulation may have higher specificity than intra-layer coregulation. In addition, the coregulation networks reveal several types of network motifs, including feed-forward loops and massive upstream crosstalk. Finally, the expression patterns of these coregulation pairs in normal and tumour tissues were analyzed. Different coregulation types show unique expression correlation trends. More importantly, the disruption of coregulation may be associated with cancers. CONCLUSION: Our findings elucidate the combinatorial and cooperative properties of transcription factors and miRNAs regulation, and we proposes that the coordinated regulation may play an important role in many biological processes. Cho-Yi Chen, Shui-Tein Chen, Chiou-Shann Fuh, Hsueh-Fen Juan, Hsuan-Cheng Huang |
BMC Bioinform. | 3 |
| 2011 | Multiple Kernel Learning for Dimensionality ReductionabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. The resulting data representations are typically high-dimensional and assume diverse forms. Hence, finding a way of transforming them into a unified space of lower dimension generally facilitates the underlying tasks such as object recognition or clustering. To this end, the proposed approach (termed MKL-DR) generalizes the framework of multiple kernel learning for dimensionality reduction, and distinguishes itself with the following three main contributions: first, our method provides the convenience of using diverse image descriptors to describe useful characteristics of various aspects about the underlying data. Second, it extends a broad set of existing dimensionality reduction techniques to consider multiple kernel learning, and consequently improves their effectiveness. Third, by focusing on the techniques pertaining to dimensionality reduction, the formulation introduces a new class of applications with the multiple kernel learning framework to address not only the supervised learning problems but also the unsupervised and semi-supervised ones. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Clustering Complex Data with Group-Dependent Feature Selection
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (6) | 3 |
| 2009 | Automatic Skin Color Beautification
Chih-Wei Chen, Da-Yuan Huang, Chiou-Shann Fuh |
ArtsIT | 3 |
| 2009 | Video stabilization for a hand-held camera based on 3D motion modelabstractIn this paper, a video stabilization technique is presented. There are four steps in the proposed approach. We begin with extracting feature points from the input image using the Lowe SIFT (scale invariant feature transform) point detection technique. This set of feature points is then matched against the set of feature points detected in the previous image using the Wyk et al. RKHS (reproducing kernel Hilbert space) graph matching technique. We can calculate the camera motion between the two images with the aid of a 3D motion model. Expected and unexpected components are separated using a motion taxonomy method. Finally, a full-frame technique to fill up blank image areas is applied to the transformed image. Jung Ming Wang, Han-Ping Chou, Sei-Wang Chen, Chiou-Shann Fuh |
ICIP | 4 |
| 2008 | Foreground Object Detection Using Two Successive ImagesabstractDetecting foreground object often need to face the problems of illumination change and image noise. In this paper, we propose an object detection method using two successive image frames. Illumination change would be very small in such short time, and then we can handle the first problem more easily. Image noise will confuse the detection of an object boundary. To handle this problem, we apply level set method to enclose the foreground object regions. The experiments show that our method can be applied to extract foreground objects in various environments and different cameras. Jung Ming Wang, Shen Cherng, Chiou-Shann Fuh, Sei-Wang Chen |
AVSS | 3 |
| 2008 | Dimensionality Reduction for Data in Multiple Feature RepresentationsabstractIn solving complex visual learning tasks, adopting multiple descriptors to more precisely characterize the data has been a feasible way for improving performance. These representations are typically high dimensional and assume diverse forms. Thus finding a way to transform them into a unified space of lower dimension generally facilitates the underlying tasks, such as object recognition or clustering. We describe an approach that incorporates multiple kernel learning with dimensionality reduction (MKL-DR). While the proposed framework is flexible in simultaneously tackling data in various feature representations, the formulation itself is general in that it is established upon graph embedding. It follows that any dimensionality reduction techniques explainable by graph embedding can be generalized by our method to consider data in multiple feature representations. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
NIPS | 3 |
| 2007 | Local Ensemble Kernel Learning for Object Category RecognitionabstractThis paper describes a local ensemble kernel learning technique to recognize/classify objects from a large number of diverse categories. Due to the possibly large intraclass feature variations, using only a single unified kernel-based classifier may not satisfactorily solve the problem. Our approach is to carry out the recognition task with adaptive ensemble kernel machines, each of which is derived from proper localization and regularization. Specifically, for each training sample, we learn a distinct ensemble kernel constructed in a way to give good classification performance for data falling within the corresponding neighborhood. We achieve this effect by aligning each ensemble kernel with a locally adapted target kernel, followed by smoothing out the discrepancies among kernels of nearby data. Our experimental results on various image databases manifest that the technique to optimize local ensemble kernels is effective and consistent for object recognition. Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
CVPR | 3 |
| 2007 | Fast and versatile algorithm for nearest neighbor search based on a lower bound tree
Yong-Sheng Chen, Yi-Ping Hung, Ting-Fang Yen, Chiou-Shann Fuh |
Pattern Recognit. | 4 |
| 2006 | Segmenting Highly Articulated Video Objects with Weak-Prior Random Forests
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (4) | 3 |
| 2006 | A New Sampling Method of Auto Focus for Voice Coil Motor in Camera Modules
Wei Hsu, Chiou-Shann Fuh |
PSIVT | 2 |
| 2005 | Learning Effective Image Metrics from Few Pairwise ExamplesabstractWe present a new approach to learning image metrics. The main advantage of our method lies in a formulation that requires only a few pairwise examples. Apparently, based on the little amount of side-information, it would take a very effective learning scheme to yield a useful image metric. Our algorithm achieves this goal by addressing two key issues. First, we establish a global-local (glocal) image representation that induces two structure-meaningful vector spaces to respectively describe the global and the local image properties. Second, we develop a metric optimization framework that finds an optimal bilinear transform to best explain the given side-information. We emphasize it is the glocal image representation that makes the use of bilinear transform more powerful. Experimental results on classifications of face images and visual tracking are included to demonstrate the contributions of the proposed method. Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
ICCV | 3 |
| 2005 | Tone Reproduction: A Perspective from Luminance-Driven Perceptual Grouping
Hwann-Tzong Chen, Tyng-Luh Liu, Chiou-Shann Fuh |
Int. J. Comput. Vis. | 3 |
| 2004 | Fast Object Detection with Occlusions
Yen-Yu Lin, Tyng-Luh Liu, Chiou-Shann Fuh |
ECCV (1) | 3 |
| 2004 | An automatic road sign recognition system based on a computational model of human recognition processing
Chiung-Yao Fang, Chiou-Shann Fuh, P. S. Yen, Shen Cherng, Sei-Wang Chen |
Comput. Vis. Image Underst. | 2 |
| 2004 | Image Reconstruction With Improved Super-Resolution AlgorithmabstractIn this paper we propose a technique that reconstructs high-resolution images with improved super-resolution algorithms, based on Irani and Peleg iterative method, and employs our suggested initial interpolation, robust image registration, automatic image selection and image enhancement post-processing. When the target of reconstruction is a moving object with respect to a stationary camera, high-resolution images can still be reconstructed, whereas previous systems only work well when we move the camera and the displacement of the whole scene is the same. Chien-Yu Chen 0001, Yu-Chuan Kuo, Chiou-Shann Fuh |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2003 | A Road Sign Recognition System Based on Dynamic Visual ModelabstractWe propose a computational model motivated by human cognitive processes for detecting changes of driving environments. The model, called dynamic visual model, consists of three major components: sensory, perceptual, and conceptual components. The proposed model is used as the underlying framework in which a system for detecting and recognizing road signs is developed. Chiung-Yao Fang, Chiou-Shann Fuh, Sei-Wang Chen, P. S. Yen |
CVPR (1) | 2 |
| 2003 | Automatic change detection of driving environments in a vision-based driver assistance systemabstractDetecting critical changes of environments while driving is an important task in driver assistance systems. In this paper, a computational model motivated by human cognitive processing and selective attention is proposed for this purpose. The computational model consists of three major components, referred to as the sensory, perceptual, and conceptual analyzers. The sensory analyzer extracts temporal and spatial information from video sequences. The extracted information serves as the input stimuli to a spatiotemporal attention (STA) neural network embedded in the perceptual analyzer. If consistent stimuli repeatedly innervate the neural network, a focus of attention will be established in the network. The attention pattern associated with the focus, together with the location and direction of motion of the pattern, form what we call a categorical feature. Based on this feature, the class of the attention pattern and, in turn, the change in driving environment corresponding to the class are determined using a configurable adaptive resonance theory (CART) neural network, which is placed in the conceptual analyzer. Various changes in driving environment, both in daytime and at night, have been tested. The experimental results demonstrated the feasibilities of both the proposed computational model and the change detection system. Chiung-Yao Fang, Sei-Wang Chen, Chiou-Shann Fuh |
IEEE Trans. Neural Networks | 3 |
| 2001 | Fast Algorithm for Nearest Neighbor Search Based on a Lower Bound Tree
Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICCV | 3 |
| 2001 | Fast Search Algorithms for Industrial InspectionabstractThis paper presents an efficient general purpose search algorithm for alignment and an applied procedure for IC print mark quality inspection. The search algorithm is based on normalized cross-correlation and enhances it with a hierarchical resolution pyramid, dynamic programming, and pixel over-sampling to achieve subpixel accuracy on one or more targets. The general purpose search procedure is robust with respect to linear change of image intensity and thus can be applied to general industrial visual inspection. Accuracy, speed, reliability, and repeatability are all critical for the industrial use. After proper optimization, the proposed procedure was tested on the IC inspection platforms in the Mechanical Industry Research Laboratories (MIRL), Industrial Technology Research Institute (ITRI), Taiwan. The proposed method meets all these criteria and has worked well in field tests on various IC products. Ming-Ching Chang, Chiou-Shann Fuh, Hsien-Yei Chen |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2001 | Simple and efficient method of calibrating a motorized zoom lens
Yong-Sheng Chen, Sheng-Wen Shih, Yi-Ping Hung, Chiou-Shann Fuh |
Image Vis. Comput. | 4 |
| 2001 | Three-dimensional ego-motion estimation from motion fields observed with multiple cameras
Yong-Sheng Chen, Lin-Gwo Liou, Yi-Ping Hung, Chiou-Shann Fuh |
Pattern Recognit. | 4 |
| 2001 | Fast block matching algorithm based on the winner-update strategyabstractBlock matching is a widely used method for stereo vision, visual tracking, and video compression. Many fast algorithms for block matching have been proposed in the past, but most of them do not guarantee that the match found is the globally optimal match in a search range. This paper presents a new fast algorithm based on the winner-update strategy which utilizes an ascending lower bound list of the matching error to determine the temporary winner. Two lower bound lists derived by using partial distance and by using Minkowski's inequality are described. The basic idea of the winner-update strategy is to avoid, at each search position, the costly computation of the matching error when there exists a lower bound larger than the global minimum matching error. The proposed algorithm can significantly speed up the computation of the block matching because: 1) computational cost of the lower bound we use is less than that of the matching error itself; 2) an element in the ascending lower bound list will be calculated only when its preceding element has already been smaller than the minimum matching error computed so far; 3) for many search positions, only the first several lower bounds in the list need to be calculated. Our experiments have shown that, when applying to motion vector estimation for several widely-used test videos, 92% to 98% of operations can be saved while still guaranteeing the global optimality. Moreover, the proposed algorithm can be easily modified either to meet the limited time requirement or to provide an ordered list of best candidate matches. Our source codes of the proposed algorithm are available at http://smart.iis.sinica.edu.tw/html/winup.html. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
IEEE Trans. Image Process. | 3 |
| 2000 | Winner-Update Algorithm for Nearest Neighbor SearchabstractThis paper presents an algorithm, called the winner-update algorithm, for accelerating the nearest neighbor search. By constructing a hierarchical structure for each feature point in the l/sub p/ metric space, this algorithm can save a large amount of computation at the expense of moderate preprocessing and twice the memory storage. Given a query point, the cost for computing the distances from this point to all the sample points can be reduced by using a lower bound list of the distance established from Minkowski's inequality. Our experiments have shown that the proposed algorithm can save a large amount of computation, especially when the distance between the query point and its nearest neighbor is relatively small. With slight modification, the winner-update algorithm can also speed up the search for k nearest neighbors, neighbors within a specified distance threshold, and neighbors close to the nearest neighbor. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICPR | 3 |
| 2000 | Camera Calibration with a Motorized Zoom LensabstractThis paper presents a simple and efficient method of calibrating the intrinsic camera parameters for all the lens settings of a motorized zoom lens. We fix the aperture setting and perform the camera calibration, adaptively, over the ranges of the room and focus settings. Bilinear interpolation is used to provide the values of the intrinsic camera parameters for those lens settings where no observations are taken. Our experiments show that the proposed method can provide accurate intrinsic camera parameters for all the lens settings, even though camera calibration is performed only for a small number of sampled lens settings. A calibration object suitable for zoom lens calibration is also presented. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh, Sheng-Wen Shih |
ICPR | 3 |
| 2000 | Hierarchical color image region segmentation for content-based image retrieval systemabstractIn this work, we propose a model of a content-based image retrieval system by using the new idea of combining a color segmentation with relationship trees and a corresponding tree-matching method. We retain the hierarchical relationship of the regions in an image during segmentation. Using the information of the relationships and features of the regions, we can represent the desired objects in images more accurately. In retrieval, we compare not only region features but also region relationships. Chiou-Shann Fuh, Shun-Wen Cho, Kai Essig |
IEEE Trans. Image Process. | 1 |
| 1998 | Free-hand pointer by use of an active stereo vision systemabstractWe developed a system of free-hand pointer by tracking the finger of the speaker with an active stereo vision system and then computing the projection direction in either of the following two modes: the finger-orientation mode and the eye-to-fingertip mode. We prefer the latter for its robustness. In order to allow the speaker to move around in a wider 3D space without reducing the pointing resolution, we utilize a well-calibrated active stereo vision system which has a relatively small field of view but can control its stereo cameras to fixate at the moving finger. Our experiments have successfully demonstrated the feasibility of developing a free-handpointer using an active stereo vision system. Yi-Ping Hung, Yao-Strong Yang, Yong-Sheng Chen, Ing-Bor Hsieh, Chiou-Shann Fuh |
ICPR | 5 |
| 1998 | Near Point Light Sources for Shape from ShadingabstractIn this paper, a linear algorithm3,4 is proposed to recover shape information from multiple images, each of them is taken under the environment that all of the object surfaces are illuminated by a known near point light source. In this method, an approximate range of the distance for the objects to the viewer (e.g. camera) is previously defined. Using this predefined value, the absolute depth map of the objects can be found out. Sheng-Liang Kao, Chiou-Shann Fuh |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1998 | Projection for pattern recognition
Chiou-Shann Fuh, Horng-Bin Liu |
Image Vis. Comput. | 1 |
| 1998 | The Fourier slice theorem for range data reconstruction
Chiou-Shann Fuh, Shih-Schon Lin |
Image Vis. Comput. | 1 |
| 1998 | Multipass hierarchical stereo matching for generation of digital terrain models from aerial images
Yi-Ping Hung, Chu-Song Chen, Kuan-Chung Hung, Yong-Sheng Chen, Chiou-Shann Fuh |
Mach. Vis. Appl. | 5 |
| 1997 | Ego-Motion Estimation Using Optical Flow Fields Observed from Multiple CamerasabstractIn this paper, we consider a multi-camera vision system mounted on a moving object in a static three-dimensional environment. By using the motion flow fields seen by all of the cameras, an algorithm which does not need to solve the point-correspondence problem among the cameras is proposed to estimate the 3D ego-motion parameters of the moving object. Our experiments have shown that using multiple optical flow fields obtained from different cameras can be very helpful for ego-motion estimation. An-Ting Tsao, Chiou-Shann Fuh, Yi-Ping Hung, Yong-Sheng Chen |
CVPR | 2 |
| 1996 | Range data reconstruction using Fourier slice theoremabstractThis paper proposes a new approach to resolve the ambiguity problem in multistriping laser triangulation systems. Our approach is based on the Fourier slice theorem which is briefly described in this paper. This theorem also forms the basis of X-ray CT (computed tomography) reconstruction. Shih-Schon Lin, Chiou-Shann Fuh |
ICPR | 2 |
| 1991 | Affine models for image matching and motion detectionabstractA model is developed for detecting the displacement field in spatiotemporal image sequences. It allows for affine shape deformations of corresponding spatial regions and for affine transformations of the image intensity range. The model includes the block matching method as a special case. A least-squares algorithm is used to find the model parameters. It is experimentally demonstrated that the affine matching model performs better than other standard approaches. The resulting 2-D motion estimates are then used by a 3-D affine model and a least-squares algorithm that recover 3-D rigid body motion and depth from two perspective views.> Chiou-Shann Fuh, Petros Maragos |
ICASSP | 1 |
| 1989 | Region-based optical flow estimationabstractA correspondence method is developed for determining optical flow where the primitive motion tokens to be matched between consecutive time frames are regions. The computation of optical flow consists of three stages: region extraction, region matching, and optical flow smoothing. The computation is completed by smoothing the initial optical flow, where the sparse velocity data are either smoothed with a vector median filter or interpolated to obtain dense velocity estimates by using a motion-coherence regularization. The proposed region-based method for optical flow is simple, computationally efficient, and more robust than iterative gradient methods, especially for medium-range motion.> Chiou-Shann Fuh, Petros Maragos |
CVPR | 1 |