VLDB 2026 Research / reviewers in the wild / expert
Yi-Ping Hung
dblp:21/1331
· DBLP profile ↗
156ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 108 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 52 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 31 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7Systems, architecture and hardware · 6 · 1 first-authorComputer networks · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real-Time Facial Animation of Gaussian Head Avatars via Mocap-to-Parametric Expression MappingabstractWe present a real-time system for animating photorealistic 3D Gaussian head avatars driven by motion data from consumer facial mocap systems. At the core of our framework is an efficient regression model that maps mocapderived blendshape parameters to the expression space of parametric face models, enabling direct control over Gaussian avatars without iterative optimization. Our design incorporates a lightweight expression regularization mechanism that improves stability and expressiveness by encouraging semantically disentangled, identity-specific deformations. Through extensive evaluations-implemented using ARKit as the mocap source and FLAME as the target parametric model-we show that our method outperforms both fitting-based and regression-based baselines in animation quality and latency. The system supports both live streaming and offline reenactment, enabling efficient, real-time avatar control for virtual meetings and social telepresence. I-Hsin Chen, Sheng-Yen Huang, Yi-Ping Hung |
3DV | 3 |
| 2025 | MindFlow: Breathing-Integrated Progressive Muscle Relaxation with a Full-Body Self-Avatar in Virtual Reality
Hangcheng Yang, Yuan-An Chan, Bin Yu 0004, Yi-Ping Hung, Panos Markopoulos 0001, Rong-Hao Liang |
Conference on Designing Interactive Systems | 4 |
| 2023 | Leap to the Eye: Implicit Gaze-based Interaction to Reveal Invisible Objects for Virtual Environment ExplorationabstractCinematic virtual reality (CVR) brings viewers a novel and immersive movie-watching experience. However, they may miss story events and scene transitions that the director has designed as key points hidden in the VR scene. In this paper, we introduce implicit gaze-based interaction for enhancing the exploration experience in CVR. In contrast to most research on gaze-based selection of objects or explicit guidance of attention using visual cues, we focus on implicit interaction that utilizes the user’s natural gaze and attention to explore the scene. We design and implement different gaze trigger methods for implicit interaction, making the interaction more intuitive and natural when users reveal the invisible objects. We implemented the adaptive collider technique, offering users a better sense of exploration than raycasting and spotlight techniques. We have also conducted user studies to compare animation sequences for visual feedback, with each animation sequence offering different storytelling techniques. One of the sequences is better suited for describing spaces in the virtual world, while the other sequence offers users the feeling of constructing a world through their gaze. Yang-Sheng Chen, Chiao-En Hsieh, Miguel Ying Jie Then, Ping-Hsuan Han, Yi-Ping Hung |
ISMAR | 5 |
| 2023 | Domain-Adaptive Mean Teacher for Category-Level Object Pose EstimationabstractCategory-level object pose estimation aims at predicting 6-DoF object poses for previously unseen objects. Current methods mostly rely on ground-truth labels such as object poses and CAD models. However, annotating these labels manually is time-consuming and error-prone in the real-world scenario. Hence, we propose a novel method to solve unsupervised domain adaptation (UDA) for category-level object pose estimation. We adopt a teacher-student framework to utilize both labeled synthetic data and unlabeled real-world data. The student and the teacher are trained to make consistent predictions under different perturbations. Furthermore, we introduce domain adversarial training to bridge the domain gap between synthetic and real-world data. To prevent false feature alignment between domains, we adopt multiple discriminators instead of a single one and perform category-aware alignments. Extensive experiments show that our method achieves state-of-the-art performance on the REAL275 dataset. Through ablation studies, we also demonstrate that our method is not restricted to certain network architecture and can serve as a general UDA method for category-level object pose estimation. I-Ju Hsieh, Yo-Chung Lau, Peng-Yuan Kao, Shih-Ping Hung, Yi-Ping Hung |
MMAsia | 5 |
| 2021 | 3D Video Stabilization With Depth Estimation by CNN-Based OptimizationabstractVideo stabilization is an essential component of visual quality enhancement. Early methods rely on feature tracking to recover either 2D or 3D frame motion, which suffer from the robustness of local feature extraction and tracking in shaky videos. Recently, learning-based methods seek to find frame transformations with high-level information via deep neural networks to overcome the robustness issue of feature tracking. Nevertheless, to our best knowledge, no learning-based methods leverage 3D cues for the transformation inference yet; hence they would lead to artifacts on complex scene-depth scenarios. In this paper, we propose Deep3D Stabilizer, a novel 3D depth-based learning method for video stabilization. We take advantage of the recent self-supervised framework on jointly learning depth and camera ego-motion estimation on raw videos. Our approach requires no data for pre-training but stabilizes the input video via 3D reconstruction directly. The rectification stage incorporates the 3D scene depth and camera motion to smooth the camera trajectory and synthesize the stabilized video. Unlike most one-size-fits-all learning-based methods, our smoothing algorithm allows users to manipulate the stability of a video efficiently. Experimental results on challenging benchmarks show that the proposed solution consistently outperforms the state-of-the-art methods on almost all motion categories. Yao-Chih Lee, Kuan-Wei Tseng, Yu-Ta Chen, Chien-Cheng Chen, Chu-Song Chen, Yi-Ping Hung |
CVPR | 6 |
| 2021 | PixStabNet: Fast Multi-Scale Deep Online Video Stabilization with Pixel-Based WarpingabstractOnline video stabilizaton is increasingly needed for real-time applications such as live streaming, drone remote control, and video communication. We propose a multi-scale convolutional neural network (PixStabNet) which stabilizes video in real time without using future frames. Instead of calculating a global homography or multiple homographies, we estimate a pixel-based warping map to make the transformation of each pixel to achieve more precise modelling. In addition, we propose well-designed loss functions along with a two-stage training scheme to enhance network robustness. The quantitative result shows that our method outperforms other learning-based online methods in terms of stability with excellent geometric and temporal consistency. Moreover, to the best of our knowledge, the proposed algorithm is the most efficient approach for video stabilization. The models and results are available at: https://yu-ta-chen.github.io/PixStabNet. Yu-Ta Chen, Kuan-Wei Tseng, Yao-Chih Lee, Yi-Ping Hung |
ICIP | 5 |
| 2021 | aBio: Active Bi-Olfactory Display Using Subwoofers for Virtual RealityabstractIncluding olfactory cues in virtual reality (VR) would enhance user immersion in the virtual environment, and precise control of smell would facilitate a more realistic experience for users. In this paper, we present aBio, an active bi-olfactory display system that delivers scents precisely to specific locations rather than diffusing scented air into the atmosphere. aBio provides users with a natural olfactory experience in free air by colliding two vortex rings launched from dual speaker-based vortex generators, which also has the effect of cushioning the force of air impact. According to the various requests of different applications, the collision point of the vortex rings can be positioned anywhere in front of the user's nose. To verify the effectiveness of our device and understand user sensations when using different parameters in our system, we conduct a series of experiments and user studies. The results show that the proposed system is effective in the sense that users perceive smell without sensible haptic disturbance while the system consumes only a very small amount of fragrant essential oil. We believe that aBio has great potential for increasing the level of presence in VR by delivering smells with high efficiency. Youyang Hu, Yao Fu Jan, Kuan-Wei Tseng, You-Shin Tsai, Hung-Ming Sung, Jin-Yao Lin, Yi-Ping Hung |
ACM Multimedia | 7 |
| 2020 | Activity Recognition Using First-Person-View Cameras Based on Sparse Optical FlowsabstractFirst-person-view (FPV) cameras are finding wide use in daily life to record activities and sports. In this paper, we propose a succinct and robust 3D convolutional neural network (CNN) architecture accompanied with an ensemble-learning network for activity recognition with FPV videos. The proposed 3D CNN is trained on low-resolution (32 × 32) sparse optical flows using FPV video datasets consisting of daily activities. According to the experimental results, our network achieves an average accuracy of 90%. Peng Yua Kao, Yan-Jing Lei, Chu-Song Chen, Ming-Sui Lee, Yi-Ping Hung |
ICPR | 6 |
| 2019 | A Kinect-Based Augmented Reality Game for Lower Limb ExerciseabstractAugmented reality (AR) is where 3D virtual objects are integrated into a 3D real environment in real time. The augmented reality applications such as medical visualization, maintenance and repair, robot path planning, entertainment, military aircraft navigation, and targeting applications have been proposed. This paper introduces the development of an augmented reality game which allows the user to carry out lower limb exercise using a natural user interface based on Microsoft Kinect. The system has been designed as an augmented game where users can see themselves in a world augmented with virtual objects generated by computer graphics. The player sitting in a chair just has to step on a mole that appears and disappears by moving upward and downward randomly. It encourages the activities of a large number of lower limb muscles which will help prevent falls. It is also suitable for rehabilitation. Yoshimasa Tokuyama, R. P. C. Janaka Rajapakse, Sachiyo Yamabe, Kouichi Konno, Yi-Ping Hung |
CW | 5 |
| 2019 | Feature Fusion of Face and Body for Engagement Intensity DetectionabstractOnline learning has grown rapidly in recent years. Automatically detecting student engagement plays a vital role in gauging the learning progress of each student.In this paper we propose a novel approach to detect student engagement. We fuse facial and body features into a single long short-term memory (LSTM) model to detect the temporal dynamics of student engagement. In contrast to other CNN models that use only facial or body features, we enhance detection accuracy with a compact feature set by merging facial and body features. Our single model generates state-of-the-art results on an engagement database from the EmotiW 2018 Challenge, where it achieves a 0.0439 mean squared error on the validation set, competitive with ensemble methods. Yan-Ying Li, Yi-Ping Hung |
ICIP | 2 |
| 2019 | Augmented Chair: Exploring the Sittable Chair in Immersive Virtual Reality for Seamless InteractionabstractVirtual reality (VR) has been a promising technique to provide an immersive experience. Comparing to traditional multimedia, when users want to take a rest or change the position such as sitting down, the chair might not be sittable because of the inconsistency between the physical and the virtual chair. In this work, we utilized a tracker attached to a physical chair and conducted a user study to explore the sitting behavior when they interact with different forms of the virtual chair. Results indicate that visualizing each part of the chair could provide different information and affect trust and preference. Ping-Hsuan Han, Ling Tsai, Jia-Wei Lin, Yuan-An Chan, Jhih-Hong Hsu, Wan-Ting Huang, Chiao-En Hsieh, Yi-Ping Hung |
VR | 8 |
| 2019 | On Learning Weight Distribution of Tai Chi Chuan Using Pressure Sensing Insoles and MR-HMDabstractTai Chi Chuan (TCC) is a famous Chinese martial art. In addition to reading instruction books and watching demonstration videos, the traditional way of learning TCC for most people is to observe and mimic the movements of the coach. However, it is not easy for a novice to determine the weight distribution of the coach, and it is also difficult for a novice to strike a pose with the correct weight distribution. To help people to learn correct weight distribution when practicing TCC, we proposed a TCC augmented reality (AR)/mixed reality (MR) learning system that consists of a mixed reality head-mounted display (MR-HMD) and a pair of pressure sensing insoles. We have designed three kinds of visual hints to help people strike poses with correct weight distributions. Two user studies have shown that our TCC learning system is helpful for people to learn the weight distribution correctly and helpful for people to improve their proprioceptive sensitivity on weight distribution. Peng-Yuan Kao, Ping-Hsuan Han, Yao Fu Jan, Chun-Hsien Li, Yi-Ping Hung |
VR | 6 |
| 2019 | Archaeological Excavation Simulation for Interaction in Virtual RealityabstractWe propose a real-time excavation simulation system for interactive gameplay in Virtual Reality. In order to increase the player's immersion, our simulation system will produce actual potholes and clods according to the depth and angle of the players excavation. We divide the process into three phases: ground deformation, clod generation and clod fragmentation. In ground deformation, we describe how to simulate the topographic changes before and after excavation. In clod generation, we describe how to generate the clod which mesh matches the depth and angle of the players excavation action. In clod fragmentation, the clods are broken and fall as the shovel lifts. This simulation system can create excavation effects on different geologies by changing the material of the ground and clods. Da-Chung Yi, Yang-Sheng Chen, Ping-Hsuan Han, Hao-Cheng Wang, Yi-Ping Hung |
VR | 5 |
| 2018 | Haptic around: multiple tactile sensations for immersive environment and interaction in virtual realityabstractIn this paper, we present Haptic Around, a hybrid-haptic feedback system, which utilizes fan, hot air blower, mist creator and heat light to recreate multiple tactile sensations in virtual reality for enhancing the immersive environment and interaction. This system consists of a steerable haptic device rigged on the top of the user head and a handheld device also with haptics feedbacks to simultaneously provide tactile sensations to the users in a 2m x 2m space. The steerable haptic device can enhance the immersive environment for providing full body experience, such as heat in the desert or cold in the snow mountain. Additionally, the handheld device can enhance the immersive interaction for providing partial body experience, such as heating the iron or quenching the hot iron. With our system, the users can perceive visual, auditory and haptic when they are moving around in virtual space and interacting with virtual object. In our study, the result has shown the potential of the hybrid-haptic feedback system, which the participants rated the enjoyment, realism, quality, immersion higher than the other. Ping-Hsuan Han, Yang-Sheng Chen, Kong-Chang Lee, Hao-Cheng Wang, Chiao-En Hsieh, Jui-Chun Hsiao, Chien-Hsing Chou, Yi-Ping Hung |
VRST | 8 |
| 2018 | One-Handed Input Through Rotational Motion for SmartwatchesabstractOne-handed input for smartwatches is crucial when users’ hands are occupied. A rotational motion is leveraged, which can be detected by built-in motion sensors in most smartwatches, as input to provide item selection. Users control a cursor by rotating the smartwatch into different orientations using the proposed item selection gestures. To understand human wrist dexterity and the ability to perform rotational motion input along each degree of freedom on smartwatches, a human-factor study was performed. Except dexterity, users’ consensus and agreement of user-defined rotational motion input gestures for selection and commitment on smartwatches in an exploratory user study were also observed. Based on both the studies, gestures roll and tilt are proposed as selection gestures, and shake and nod are used as commitment gestures. The item selection performance using the proposed gestures in 2D layouts in a user study was further evaluated. The study results showed that tilt-nod and tilt-shake are suitable for item numbers (4 × 3) and (5 × 5), respectively. Furthermore, due to high error rate for tilt-shake in (2 × 2) caused by false positive in shake detection, roll-shake with better social acceptance can be used as an alternative for smartwatches without a built-in front camera in (2 × 2). Hsin-Ruey Tsai, Po-Chang Chen, Li-Wei Chan 0001, Yi-Ping Hung |
Int. J. Hum. Comput. Interact. | 4 |
| 2017 | Learning and inferring human actions with temporal pyramid features based on conditional random fieldsabstractFinding an effective way to represent human actions is yet an open problem because it usually requires taking evidences extracted from various temporal resolutions into account. A conventional way of representing an action employs temporally ordered fine-grained movements, e.g., key poses or subtle motions. Many existing approaches model actions by directly learning the transitional relationships between those fine-grained features. Yet, an action data may have many similar observations with occasional and irregular changes, which make commonly used fine-grained features less reliable. This paper presents a set of temporal pyramid features that enriches action representation with various levels of semantic granularities. For learning and inferring the proposed pyramid features, we adopt a discriminative model with latent variables to capture the hidden dynamics in each layer of the pyramid. Our method is evaluated on a Tai-Chi Chun dataset and a daily activities dataset. Both of them are collected by us. Experimental results demonstrate that our approach achieves more favorable performance than existing methods. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 4 |
| 2017 | SaFeplay: a lightweight portable sensing system to estimate knee adduction momentabstractTraditional lower limb biomechanics detection system are not only expansive but difficult to monitor user in the whole day. It is not friendly to the large scale of patients with knee joint disease such as knee osteoarthritis. Therefore, we design a lightweight portable lower limbs motion monitoring, biomechanics measurement, and analysis system, Smart Footwear Platform (SaFePlay). We apply SaFePlay to estimate the index of knee joint loading, knee adduction moment (KAM), by the lever-arm approach. We get information about user's foot pressure and knee bending angle, then we can derive KAM with body parameters. We conduct an experiment to verify the accuracy of the proposed method by comparing with motion capture and 3DoF force plates systems as ground truth. The result shows that the curve of KAM from SaFePlay is similar to the ground truth. The proposed method not only makes KAM evaluation portable but also requires only lightweight devices. Shih-Yao Wei, Yi-Ping Lo, Chia-Yi Lin, Tse-Yu Lin, Yin-Yu Chou, Jung-Tang Huang, Rong-Sen Yang, Yi-Ping Hung |
MUM | 8 |
| 2017 | An Ensemble of Invariant Features for Person ReidentificationabstractThis paper proposes an ensemble of invariant features (EIFs), which can properly handle the variations of color difference and human poses/viewpoints for matching pedestrian images observed in different cameras with nonoverlapping field of views. Our proposed method is a direct reidentification (re-id) method, which requires no prior domain learning based on prelabeled corresponding training data. The novel features consist of the holistic and region-based features. The holistic features are extracted by using a publicly available pretrained deep convolutional neural network used in generic object classification. In contrast, the region-based features are extracted based on our proposed two-way Gaussian mixture model fitting, which overcomes the self-occlusion and pose variations. To make a better generalization during recognizing identities without additional learning, the ensemble scheme aggregates all the feature distances using the similarity normalization. The proposed framework achieves robustness against partial occlusion, pose, and viewpoint changes. Moreover, the evaluation results show that our method outperforms the state-of-the-art direct re-id methods on the challenging benchmark viewpoint invariant pedestrian recognition and 3D people surveillance data sets. Young-Gun Lee, Shen-Chi Chen, Jenq-Neng Hwang, Yi-Ping Hung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Vision-Based Positioning for Internet-of-VehiclesabstractThis paper presents an algorithm for ego-positioning by using a low-cost monocular camera for systems based on the Internet-of-Vehicles. To reduce the computational and memory requirements, as well as the communication load, we tackle the model compression task as a weighted k-cover problem for better preserving the critical structures. For real-world vision-based positioning applications, we consider the issue of large scene changes and introduce a model update algorithm to address this problem. A large positioning data set containing data collected for more than a month, 106 sessions, and 14275 images is constructed. Extensive experimental results show that submeter accuracy can be achieved by the proposed ego-positioning algorithm, which outperforms existing vision-based approaches. Kuan-Wen Chen, Chun-Hsin Wang, Qiao Liang 0002, Chu-Song Chen, Ming-Hsuan Yang 0001, Yi-Ping Hung |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2017 | Recognizing Human Actions with Outlier Frames by Observation Filtering and CompletionabstractThis article addresses the problem of recognizing partially observed human actions. Videos of actions acquired in the real world often contain corrupt frames caused by various factors. These frames may appear irregularly, and make the actions only partially observed. They change the appearance of actions and degrade the performance of pretrained recognition systems. In this article, we propose an approach to address the corrupt-frame problem without knowing their locations and durations in advance. The proposed approach includes two key components: outlier filtering and observation completion . The former identifies and filters out unobserved frames, and the latter fills up the filtered parts by retrieving coherent alternatives from training data. Hidden Conditional Random Fields (HCRFs) are then used to recognize the filtered and completed actions. Our approach has been evaluated on three datasets, which contain both fully observed actions and partially observed actions with either real or synthetic corrupt frames. The experimental results show that our approach performs favorably against the other state-of-the-art methods, especially when corrupt frames are present. Shih-Yao Lin 0001, Yen-Yu Lin, Chu-Song Chen, Yi-Ping Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2016 | DigitSpace: Designing Thumb-to-Fingers Touch Interfaces for One-Handed and Eyes-Free InteractionsabstractThumb-to-fingers interfaces augment touch widgets on fingers, which are manipulated by the thumb. Such interfaces are ideal for one-handed eyes-free input since touch widgets on the fingers enable easy access by the stylus thumb. This study presents DigitSpace, a thumb-to-fingers interface that addresses two ergonomic factors: hand anatomy and touch precision. Hand anatomy restricts possible movements of a thumb, which further influences the physical comfort during the interactions. Touch precision is a human factor that determines how precisely users can manipulate touch widgets set on fingers, which determines effective layouts of the widgets. Buttons and touchpads were considered in our studies to enable discrete and continuous input in an eyes-free manner. The first study explores the regions of fingers where the interactions can be comfortably performed. According to the comfort regions, the second and third studies explore effective layouts for button and touchpad widgets. The experimental results indicate that participants could discriminate at least 16 buttons on their fingers. For touchpad, participants were asked to perform unistrokes. Our results revealed that since individual participant performed a coherent writing behavior, personalized $1 recognizers could offer 92% accuracy on a cross-finger touchpad. A series of design guidelines are proposed for designers, and a DigitSpace prototype that uses magnetic-tracking methods is demonstrated. Da-Yuan Huang, Li-Wei Chan 0001, Rong-Hao Liang, De-Nian Yang, Yi-Ping Hung, Bing-Yu Chen 0004 |
CHI | 7 |
| 2016 | iKneeBraces: knee adduction moment evaluation measured by motion sensors in gait detectionabstractWe propose light-weight wearable devices, iKneeBraces, to prevent knee osteoarthritis (OA) using knee adduction moment (KAM) evaluation. iKneeBrace consists of two inertial measurement units (IMUs) to measure shin and thigh angles. KAM is estimated by ground force reaction (GRF), knee position and center of pressure position. Instead of heavy and bulky 3DoF force plates conventionally used, we propose to build a 2D input regression model using shin and thigh angles from iKneeBrace as input to infer GRF direction and further estimate KAM. We perform an experiment to evaluate the method. The results show that iKneeBrace can infer KAM similar to the ground truth in the first peak, the most important part to prevent knee OA. Furthermore, the proposed method can infer KAM in all parts if better IMUs used in iKneeBrace in the future. The proposed method not only makes KAM evaluation portable but also requires only light-weight devices. Hsin-Ruey Tsai, Shih-Yao Wei, Jui-Chun Hsiao, Ting-Wei Chiu, Yi-Ping Lo, Chi-Feng Keng, Yi-Ping Hung, Jin-Jong Chen |
UbiComp | 7 |
| 2016 | Visual enhancement using sparsity-based image decomposition for low backlight displaysabstractWe propose a power-constrained image enhancement system to maintain human visual perception when the LCD or LED display is under low backlight. Adopting the low backlight mode can save the electricity and lengthen the battery using time. First, we deduce the relationship between the image and the backlight for maintaining the same visual perceptual quality. Then, we propose a sparsity-based image decomposition to separate the intensity image into base layer and detail layer. Afterwards, we refer to the image-backlight relationship to compensate the base layer, while we also adopt texture-aw are boosting to enhance the detail layer. Experimental simulated results show that our system outperforms than the compared systems. Chih-Tsung Shen, Zongqing Lu 0001, Yi-Ping Hung, Soo-Chang Pei |
ISCAS | 3 |
| 2016 | Nail+: sensing fingernail deformation to detect finger force touch interactions on rigid surfacesabstractForce sensing has been widely used for bringing the touch from binary to multiple states, creating new abilities on surface interactions. However, prior proposed force sensing techniques mainly focus on enabling force-applied gestures on certain devices. This paper presents Nail+, a technique using fingernail deformation to enable force touch sensing interactions on everyday rigid surfaces. Our prototype, 3x3 0.2mm strain sensor array mounted on a fingernail, was implemented and conducted with a 12-participant study for evaluating the feasibility of this sensing approach. Result showed that the accuracy for sensing normal and force-applied tapping and swiping can achieve 84.67% on average. We finally proposed two example applications using Nail+ prototype for controlling the interfaces of head-mounted display (HMD) devices and remote screens. Min-Chieh Hsiu, Chiuan Wang, Da-Yuan Huang, Jhe-Wei Lin, Yu-Chih Lin, De-Nian Yang, Yi-Ping Hung, Mike Y. Chen |
MobileHCI | 7 |
| 2016 | CircuitStack: Supporting Rapid Prototyping and Evolution of Electronic CircuitsabstractFor makers and developers, circuit prototyping is an integral part of building electronic projects. Currently, it is common to build circuits based on breadboard schematics that are available on various maker and DIY websites. Some breadboard schematics are used as is without modification, and some are modified and extended to fit specific needs. In such cases, diagrams and schematics merely serve as blueprints and visual instructions, but users still must physically wire the breadboard connections, which can be time-consuming and error-prone. We present CircuitStack, a system that combines the flexibility of breadboarding with the correctness of printed circuits, for enabling rapid and extensible circuit construction. This hybrid system enables circuit reconfigurability, component reusability, and high efficiency at the early stage of prototyping development. Chiuan Wang, Hsuan-Ming Yeh, Bryan Wang, Te-Yen Wu, Hsin-Ruey Tsai, Rong-Hao Liang, Yi-Ping Hung, Mike Y. Chen |
UIST | 7 |
| 2015 | Location-aware object detection via coherent region groupingabstractWe present a scene adaptation algorithm for object detection. Our method discovers scene-dependent features discriminative to classifying foreground objects into different categories. Unlike previous works suffering from insufficient training data collected online, our approach incorporated with a similarity grouping procedure can automatically gather more consistent training examples from a neighbour area. Experimental results show that the proposed method outperforms several related works with higher detection accuracies. Shen-Chi Chen, Chu-Song Chen, Yi-Ping Hung |
ICASSP | 4 |
| 2015 | 3D Printing and Camera Mapping: Dialectic of Virtual and RealityabstractProjection Mapping, the superimposing of virtual images upon actual objects, is already extensively used in performance arts. Applications of it are already quite mature, therefore, here we wish to achieve the opposite, or specifically speaking, the superimposing of actual objects into virtual images. This method of reverse superimposition is called "camera mapping." Through cameras, camera mapping captures actual objects, and introduces them into a virtual world. Then using superimposition, this allows for actual objects to be rendered as virtual objects. However, the actual objects here must have refined shapes so that they may be superimposed back into the camera. Through the proliferation of 3D printing, virtual 3D models in computers can be created in reality, thereby providing a framework for the limits and demands of "camera mapping." The new media artwork Digital Buddha combines 3D Printing and camera mapping. This work was created by 3-D deformable modeling through a computer, then transforming the model into a sculpture using 3D printing, and then remapping the materially produced sculpture back into the camera. Finally, it uses the already known algorithm to convert the model back into that of the original non-deformed sculpture. From this creation project, in the real world, audiences will see a deformed, abstract sculpture; and in the virtual world, through camera mapping, they will see a concrete sculpture (Buddha). In its representation, this piece of work pays homage to the work TV Buddha produced by video art master Nam June Paik. Using the influence television possesses over people, this work extends into the most important concepts of the digital era, "coding" and "decoding," simultaneously addressing the shock and insecurity people in the digital era feel toward images. He-Lin Luo, I-Chun Chen, Yi-Ping Hung |
ACM Multimedia | 3 |
| 2015 | An ensemble of invariant features for person re-identificationabstractWe propose an ensemble of invariant features for person re-identification. The proposed method requires no domain learning and can effectively overcome the issues created by the variations of human poses and viewpoint between a pair of different cameras. Our ensemble model utilizes both holistic and region-based features. To avoid the misalignment problem, the test human object sample is used to generate multiple virtual samples, by applying slight geometric distortion. The holistic features are extracted from a publically available pre-trained deep convolutional neural network. On the other hand, the region-based features are based on our proposed Two-Way Gaussian Mixture Model Fitting and the Completed Local Binary Pattern texture representations. To make better generalization during the matching without additional learning processes for the feature aggregation, the ensemble scheme combines all three feature distances using distances normalization. The proposed framework achieves robustness against partial occlusion, pose and viewpoint changes. In addition, the experimental results show that our method exceeds the state of the art person re-identification performance based on the challenging benchmark 3DPeS. Shen-Chi Chen, Young-Gun Lee, Jenq-Neng Hwang, Yi-Ping Hung, Jang-Hee Yoo |
MMSP | 4 |
| 2015 | Hybrid Method for 3-D Gaze Tracking Using Glint and Contour FeaturesabstractGlint features have important roles in gaze-tracking systems. However, when the operation range of a gaze-tracking system is enlarged, the performance of glint-feature-based (GFB) approaches will be degraded mainly due to the curvature variation problem at around the edge of the cornea. Although the pupil contour feature may provide complementary information to help estimating the eye gaze, existing methods do not properly handle the cornea refraction problem, leading to inaccurate results. This paper describes a contour-feature-based (CFB) 3-D gaze-tracking method that is compatible to cornea refraction. We also show that both the GFB and CFB approaches can be formulated in a unified framework and, thus, they can be easily integrated. Furthermore, it is shown that the proposed CFB method and the GFB method should be integrated because the two methods provide complementary information that helps to leverage the strength of both features, providing robustness and flexibility to the system. Computer simulations and real experiments show the effectiveness of the proposed approach for gaze tracking. Chih-Chuan Lai, Sheng-Wen Shih, Yi-Ping Hung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Abandoned Object Detection via Temporal Consistency Modeling and Back-Tracing Verification for Visual SurveillanceabstractThis paper presents an effective approach for detecting abandoned luggage in surveillance videos. We combine short- and long-term background models to extract foreground objects, where each pixel in an input image is classified as a 2-bit code. Subsequently, we introduce a framework to identify static foreground regions based on the temporal transition of code patterns, and to determine whether the candidate regions contain abandoned objects by analyzing the back-traced trajectories of luggage owners. The experimental results obtained based on video images from 2006 Performance Evaluation of Tracking and Surveillance and 2007 Advanced Video and Signal-based Surveillance databases show that the proposed approach is effective for detecting abandoned luggage, and that it outperforms previous methods. Shen-Chi Chen, Chu-Song Chen, Daw-Tung Lin, Yi-Ping Hung |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2015 | Viewing-Distance Aware Super-Resolution for High-Definition DisplayabstractIn this paper, we propose a novel algorithm for high-definition displays to enlarge low-resolution images while maintaining perceptual constancy (i.e., the same field-of-view, perceptual blur radius, and the retinal image size in viewer's eyes). We model the relationship between a viewer and a display by considering two main aspects of visual perception, i.e., scaling factor and perceptual blur radius. As long as we enlarge an image while adjust its image blur levels on the display, we can maintain viewer's perceptual constancy. We show that the scaling factor should be set in proportion to the viewing distance and the blur levels on the display should be adjusted according to the focal length of a viewer. Toward this, we first refer to edge directions to interpolate a low-resolution image with the increasing of viewing distance and the scaling factor. After images are interpolated, we utilize a local contrast to estimate the spatially varying image blur levels of the interpolated image. We then further adjust the image blur levels using a parametric deblurring method, which combines L1 as well as L2 reconstruction errors, and Tikhonov with total variation regularization terms. By taking these factors into account, high-resolution images adaptive to viewing distance on a display can be generated. Experimental results on both natural image metric and user subjective studies across image scales demonstrate that the proposed super-resolution algorithm for high-definition displays performs favorably against the state-of-the-art methods. Chih-Tsung Shen, Hung-Hsun Liu, Ming-Hsuan Yang 0001, Yi-Ping Hung, Soo-Chang Pei |
IEEE Trans. Image Process. | 4 |
| 2014 | A spatiotemporal background extractor using a single-layer codebook modelabstractBackground subtraction is a crucial component in visual surveillance, which has been studied over years. However, an efficient algorithm that can tolerate the environment changes such as dynamic backgrounds and sudden changes of illumination is still demanding. In this paper, we design an innovative framework called the spatiotemporal background extractor (SBE) from a single-layer codebook model. Two main extractors, the background extractor (BE) and the background gradient extractor (BGE), are constructed to extract the foreground objects. The background extractor is built for each single frame with spatial information propagated from the neighbor locations, which is useful for handling dynamic background and sudden lighting changes. The background gradient extractor is also constructed and updated, and we design a propagation forbidden policy for background updating, so as to keep the completeness of foreground shape via the background gradient information. The proposed method can efficiently capture the foreground and eliminates the noise of background. The performance of the proposed method is compared with MoG [3], Codebook [4] and ViBe [8] on the Wallflower [1] and Perception [2] datasets. Chih-Wei Lin 0004, Wei-Jie Liao, Chu-Song Chen, Yi-Ping Hung |
AVSS | 4 |
| 2014 | TouchSense: expanding touchscreen input vocabulary using different areas of users' finger padsabstractWe present TouchSense, which provides additional touchscreen input vocabulary by distinguishing the areas of users' finger pads contacting the touchscreen. It requires minimal touch input area and minimal movement, making it especially ideal for wearable devices such as smart watches and smart glasses. For example, users of a calculator application on a smart watch could tap normally to enter numbers, and tap with the right side of their fingers to enter the operators (e.g. , -, =). Results from two human-factor studies showed that users could tap a touchscreen with five or more distinct areas of their finger pads. Also, they were able to tap with more distinct areas closer to their fingertips. We developed a TouchSense smart watch prototype using inertial measurement sensors, and developed two example applications: a calculator and a text editor. We also collected user feedback via an explorative study. Da-Yuan Huang, Ming-Chang Tsai, Ying-Chao Tung, Min-Lun Tsai, Yen-Ting Yeh 0001, Li-Wei Chan 0001, Yi-Ping Hung, Mike Y. Chen |
CHI | 7 |
| 2014 | A sleep monitoring system based on audio, video and depth information for detecting sleep eventsabstractThe purpose of this study is to develop a non-invasive sleep monitoring system to distinguish sleep disturbances based on multiple sensors. Unlike clinical sleep monitoring which records biological information such as EEG, EOG, and EMG, in this study, we aim to identify occurrences of events from a sleep environment. A device with an infrared depth sensor, a RGB camera, and a four-microphone array is used to detect three types of events: motion events, lighting events, and sound events. Given streams of depth signals and color images, we build two background models to detect movements and lighting effects, and audio signals are scored simultaneously. Moreover, we classify events by an epoch approach algorithm and provide a graphical sleep diagram for browsing corresponding video clips. Experimental results in sleep condition show the efficiency and reliability of our system, and it is convenient and cost-effective to be used in home context. Lyn Chao-ling Chen, Kuan-Wen Chen, Yi-Ping Hung |
ICME | 3 |
| 2014 | I-m-Cave: An interactive tabletop system for virtually touring Mogao CavesabstractThis paper presents i-m-Cave, a multitouch tabletop system for virtually touring the Mogao Caves. To recreate the experience of such a tour, a field study was conducted, and identified two key design considerations - exploration and restoration. For exploration, an innovative tangible figurine that can perform human-like neck extension/flexion is developed, and users can control the figurine to freely visit every corner of the virtual caves. Furthermore, users can view authorized animations to better understand the story behind important artifacts. With respect to restoration, users can restore digital artifacts and observe the rejuvenation and aging of artifacts with midair hand gestures and mobile devices, as if the users were shuttling back and forth between the present and the past. The i-m-Cave system presents the relationships between presently damaged artifacts and their virtually restored versions, which cannot be experienced in a real tour. Finally, the system was evaluated by interviewing experts. User feedback was positive, indicating the value of the system and its design. Da-Yuan Huang, Shen-Chi Chen, Li-Erh Chang, Po-Shiun Chen, Yen-Ting Yeh 0001, Yi-Ping Hung |
ICME | 6 |
| 2014 | Appearance-Based Gaze Tracking with Free Head MovementabstractIn this work, we develop an appearance-based gaze tracking system allowing user to move their head freely. The main difficulty of the appearance-based gaze tracking method is that the eye appearance is sensitive to head orientation. To overcome the difficulty, we propose a 3-D gaze tracking method combining head pose tracking and appearance-based gaze estimation. We use a random forest approach to model the neighbor structure of the joint head pose and eye appearance space, and efficiently select neighbors from the collected high dimensional data set. L1-optimization is then used to seek for the best solution for regression from the selected neighboring samples. Experiment results shows that it can provide robust binocular gaze tracking results with less constraints but still provides moderate estimation accuracy of gaze estimation. Chih-Chuan Lai, Kuan-Wen Chen, Shen-Chi Chen, Sheng-Wen Shih, Yi-Ping Hung |
ICPR | 6 |
| 2014 | 3-D Gaze Tracking Using Pupil Contour FeaturesabstractGlint features have important roles in gaze tracking systems. But when the operation range of a gaze tracking system is enlarged, the performance of glint-feature-based (GFB) approaches will be degraded mainly due to the curvature variation problem at around the edge of the cornea. Although the pupil contour feature may provide complementary information to help estimating the eye gaze, existing methods do not properly handle the cornea refraction problem, leading to inaccurate results. This paper describes a contour-feature-based (CFB) 3-D gaze tracking method that is compatible to cornea refraction. Experiments show the effectiveness of the proposed approach for gaze tracking. Chih-Chuan Lai, Sheng-Wen Shih, Hsin-Ruey Tsai, Yi-Ping Hung |
ICPR | 4 |
| 2014 | Left-Luggage Detection from Finite-State-Machine Analysis in Static-Camera VideosabstractWe present an abandoned object detection system in this paper. A finite-state-machine model is introduced to extract stationary foregrounds in a scene for visual surveillance, where the state value of each pixel is inferred via the cooperation of short-term and long-term background models constructed in the proposed approach. To identify the left-luggage event, we then verify whether the static foregrounds are abandoned objects through the analysis of owner's moving trajectory back-tracked to the static foreground locations. Experimental results reveal that the proposed approach tackles the problem well on publicly available datasets. Shen-Chi Chen, Chu-Song Chen, Daw-Tung Lin, Yi-Ping Hung |
ICPR | 5 |
| 2014 | Opportunities for Persuasive Technology to Motivate Heavy Computer Users for Stretching Exercise
Yong-Xiang Chen, Siek-Siang Chiang, Shu-Yun Chih, Wen-Ching Liao, Shih-Yao Lin 0001, Shang-Hua Yang, Shun-Wen Cheng, Shih-Sung Lin, Yu-Shan Lin, Ming-Sui Lee, Jau-Yih Tsauo, Cheng-Min Jen, Chia-Shiang Shih, King-Jen Chang, Yi-Ping Hung |
PERSUASIVE | 15 |
| 2014 | Large-Area, Multilayered, and High-Resolution Visual Monitoring Using a Dual-Camera SystemabstractLarge-area, high-resolution visual monitoring systems are indispensable in surveillance applications. To construct such systems, high-quality image capture and display devices are required. Whereas high-quality displays have rapidly developed, as exemplified by the announcement of the 85-inch 4K ultrahigh-definition TV by Samsung at the 2013 Consumer Electronics Show (CES), high-resolution surveillance cameras have progressed slowly and remain not widely used compared with displays. In this study, we designed an innovative framework, using a dual-camera system comprising a wide-angle fixed camera and a high-resolution pan-tilt-zoom (PTZ) camera to construct a large-area, multilayered, and high-resolution visual monitoring system that features multiresolution monitoring of moving objects. First, we developed a novel calibration approach to estimate the relationship between the two cameras and calibrate the PTZ camera. The PTZ camera was calibrated based on the consistent property of distinct pan-tilt angle at various zooming factors, accelerating the calibration process without affecting accuracy; this calibration process has not been reported previously. After calibrating the dual-camera system, we used the PTZ camera and synthesized a large-area and high-resolution background image. When foreground targets were detected in the images captured by the wide-angle camera, the PTZ camera was controlled to continuously track the user-selected target. Last, we integrated preconstructed high-resolution background and low-resolution foreground images captured using the wide-angle camera and the high-resolution foreground image captured using the PTZ camera to generate a large-area, multilayered, and high-resolution view of the scene. Chih-Wei Lin 0004, Kuan-Wen Chen, Shen-Chi Chen, Cheng-Wu Chen, Yi-Ping Hung |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2013 | Target-driven video summarization in a camera networkabstractNowadays, ever expanding camera network makes it difficult to find the suspect from lengthy video records. This paper proposes a target-driven video summarization framework which provides two-step Filtered Summarized Video (FSV) for tracing suspects. Before the target is identified, users can find the target efficiently using the firststep FSV of any arbitrary camera. The first-step FSV filters all the attributes of the target including the time information and the target's categories. After identifying the target, the second-step FSV with additional spatio-temporal & appearance cues are triggered in the neighbor cameras. To enhance the accuracy of the object classification for FSV, we propose a Perspective Dependent Model (PDM) which consists of many grid-based models. Finally, the experimental results show that grid-based model is more robust than general detectors and the user study demonstrates better performance for target finding and tracking in camera network for surveillance. Shen-Chi Chen, Shih-Yao Lin 0001, Kuan-Wen Chen, Chih-Wei Lin 0004, Chu-Song Chen, Yi-Ping Hung |
ICIP | 7 |
| 2013 | Real-time camera tampering detection using two-stage scene matchingabstractWe propose a tampering detection method using two-stage scene matching for real application with high efficiency and low false alarm rate. In the first stage, we use the intensity of edges as the main cue to detect the camera tampering events. Instead of using the entire edge points of the images, we sample the most significant edge points to represent the scene. Analyzing the edge variation with only the sample points, we discover that the events of camera tampering can be detected with low computation cost. Whenever the first stage detects the tampering event, the second stage is triggered to reduce false alarms. In the second stage, we propose an illumination change detector which can check the consistency of the scene structure using cell-based matching method. The experimental results demonstrate that our system can detect the camera tampering precisely and minimize false alarm even when the illumination changes dramatically or large crowds passing through the scene. Chao-Ching Shih, Shen-Chi Chen, Cheng-Feng Hung, Kuan-Wen Chen, Shih-Yao Lin 0001, Chih-Wei Lin 0004, Yi-Ping Hung |
ICME | 7 |
| 2013 | Spatially-varying super-resolution for HDTVabstractWe propose a system to up-sample visual signals for high definition televisions (HDTVs). Although the original visual signals are degraded and limited, we still try to solve these problems by using super-resolution. First, we interpolate the visual signals with a edge taper to the desired size. Then, we combine L1 data term, L2 data term, Tikhonov-like regularizer and total variation regularizer with a saliency weighting for our L12TTV deblurring method. Adopting our L12TTV deblurring, we can remove the spatially-varying blurry conditions of the interpolated signals. Experimental results show that our system outperforms than comparisons in both image and video cases. Chih-Tsung Shen, Hung-Hsun Liu, Ming-Sui Lee, Yi-Ping Hung, Soo-Chang Pei |
ISCAS | 4 |
| 2013 | Rapid selection of hard-to-access targets by thumb on mobile touch-screensabstractCurrent touch-based UIs commonly employ regions near the corners and/or edges of the display to accommodate essential functions. As the screen size of mobile phones is ever increasing, such regions become relatively distant from the thumb and hard to reach for single-handed use. In this paper, we present two techniques: CornerSpace and BezelSpace, designed to accommodate quick access to screen targets outside the thumb's normal interactive range. Our techniques automatically determine the thumb's physical comfort zone and only require minimal thumb movement to reach distant targets on the edge of the screen. A controlled experiment shows that BezelSpace is significantly faster and more accurate. Moreover, both techniques are application-independent, and instantly accommodate either hand, left or right. Neng-Hao Yu, Da-Yuan Huang, Jia-Jyun Hsu, Yi-Ping Hung |
Mobile HCI | 4 |
| 2013 | AirTouch panel: a re-anchorable virtual touch panelabstractTo achieve maximum mobility, device-less approaches for home appliance remote control have received increasing attention in recent years. In this paper, we propose a screen-less virtual touch panel, called AirTouch Panel, which can be positioned at any place with various orientations around users. The proposed virtual touch panel provides a potential ability to remotely control the home appliances, such as television, air conditioner, and so on. The proposed system allows users to anchor the panel at the place with comfortable poses. If the users want to change panel's position or orientation, they only need to re-anchor it, and then the panel will be reset. In this paper, our main contribution is to design a re-anchorable virtual panel for digital home remote control. Most importantly, we explore the design of such imaginary interface through two user studies. In our user studies, we analyze task completion time, satisfaction rate, and the number of miss-clicks. We are interested in the feasibility issues, for example, proper click gesture, panel size and button size, etc. Moreover, based on the AirTouch Panel, we also developed an intelligent TV to demonstrate the usability for controlling home appliance. Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Yi-Ping Hung |
ACM Multimedia | 4 |
| 2013 | Intensity Rank Estimation of Facial Expressions Based on a Single ImageabstractIn this paper, we propose a framework that estimates the discrete intensity rank of a facial expression based on a single image. For most people, judging whether an expression is more intense than others is easier than determining its real-valued intensity degree, and hence the relative order of two expressions is more distinguishable than the exact difference between them. We utilize the relative order to construct an image-based ranking approach for inferring the discrete ranks. The challenge of image-based approaches is to conduct a representation for subtle expression changes. We employ an efficient descriptor, scattering transform, which is translation invariant and can linearize deformations. This scattering representation recovers the lost high frequencies and retains discrimination under invariant property. Our experimental results demonstrate that the proposed framework with scattering transform outperforms other compared feature descriptors and algorithms. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
SMC | 3 |
| 2013 | Facial Trait CodeabstractWe propose a facial trait code (FTC) to encode human facial images, and apply it to face recognition. Extracted from an exhaustive set of local patches cropped from a large stack of faces, the facial traits and the associated trait patterns can accurately capture the appearance of a given face. The extraction has two phases. The first phase is composed of clustering and boosting upon a training set of faces with neutral expression, even illumination, and frontal pose. The second phase focuses on the extraction of the facial trait patterns from the set of faces with variations in expression, illumination, and poses. To apply the FTC to face recognition, two types of codewords, hard and probabilistic, with different metrics for characterizing the facial trait patterns are proposed. The hard codeword offers a concise representation of a face, while the probabilistic codeword enables matching with better accuracy. Our experiments compare the proposed FTC to other algorithms on several public datasets, all showing promising results. Ping-Han Lee, Gee-Sern Hsu, Tsuhan Chen, Yi-Ping Hung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | 2D Face Alignment and Pose Estimation Based on 3D Facial ModelsabstractFace alignment and head pose estimation has become a thriving research field with various applications for the past decade. Several approaches process on 2D texture image but most of them perform decently only with small pose variation. Recently, many approaches apply depth information to align objects. However, applications are restricted because depth cameras are more expensive than common cameras, and many original image resources contain no depth information. Therefore, we propose a 3D face alignment algorithm in 2D image based on Active Shape Model, and use Speeded-Up Robust Features (SURF) descriptors as local texture model. We train a 3D shape model with different view-based local texture models from a 3D database, and then fit a face in a 2D image by these models. We also improve the performance by two-stage search strategy. Furthermore, the head pose can be estimated by the alignment result of the proposed 3D model. Finally, we demonstrate some applications applied by our method. Shen-Chi Chen, Chia-Hsiang Wu, Shih-Yao Lin 0001, Yi-Ping Hung |
ICME | 4 |
| 2012 | Applying scattering operators for face recognition: A comparative study
Kuang-Yu Chang, Cheng-Fu Lin, Chu-Song Chen, Yi-Ping Hung |
ICPR | 4 |
| 2012 | Human action recognition using Action Trait Code
Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Ming-Sui Lee, Yi-Ping Hung |
ICPR | 5 |
| 2012 | Action recognition for human-marionette interactionabstractIn this paper, we propose a human-marionette interaction system based on a human action recognition approach for applications to interactive artistic puppetry and a mimicking-marionette game. We developed an intelligent marionette called "i-marionette" that is controlled by a sophisticated control device to achieve various human actions. Moreover, we utilized an action recognition approach to enable the i-marionette to learn and recognize complex dance movements. The idea of artistic puppetry is to present a conflict scenario between two different cultural worlds: the performer is active and represents the culture of modern technology based in the real world. In contrast, the i-marionette represents traditional culture and is passive and based in a virtual world. The active performer guides the passive i-marionette to form a space-time connection between the real world and the virtual world. The i-marionette mimics the performer's action, while the performer also mimics the i-marionette's action. The performance represents an artistic conception in which humans invent technology and the i-marionette is manipulated by human control. However, in this interactive circle, the human is implicitly affected by the i-marionette. In our mimicking-marionette game, a player mimics the i-marionette's action. Subsequently, our human action recognition system measures the action similarity between the player and the i-marionette, and our system provides a similarity score. Shih-Yao Lin 0001, Chuen-Kai Shie, Shen-Chi Chen, Yi-Ping Hung |
ACM Multimedia | 4 |
| 2012 | Interactive art "maelstrom&vortex": the body's speed - a race between digital and analog speedsabstractWith an explosion of accessible information, the digital era has been integrated into our daily lives. At any time, people can operate handheld computers (cell phones) to connect to the online world. The thrill of speed is slowly replacing the analog speed of the human body. Even as the human body attempts to catch up with the speed of a computer, it tries to flee from it at the same time. This type of hope returns us to an original and natural sense of ambivalence, so that, in post-digital times, people will attempt to find the divine light of the analog era. The work, "Maelstrom&Vortex", utilizes interactive art to try to explain analog and digital speeds. Through the interaction, people can experience differences in speed between physical sensations and computer data as it re-combines them with a sense of space. In the work, the analog movement of mechanical motors, the digital capture of cameras, and digital computation of computers are combined, and then intervened with the analog-like human body. This allows participants to discover the relationship amongst analog, digital, and the body. He-Lin Luo, Yi-Ping Hung |
ACM Multimedia | 2 |
| 2012 | MagMobile: enhancing social interactions with rapid view-stitching games of mobile devicesabstractMost mobile games are designed for users to only focus on their own screens thus lack of face-to-face interaction even users are sitting together. Prior work shows that the shared information space created by multiple mobile devices can encourage users to communicate to each other naturally. The aim of this work is to provide a fluent view-stitching technique for mobile phone users to establish their information-shared view. We present MagMobile: a new spatial interaction technique that allows users to stitch views by simply putting multiple mobile devices close to each other. We describe the design of spatial-aware sensor module which is low cost and easy to be obtained into phones. We also propose two collaborative games to engage social interactions in the co-located place. Da-Yuan Huang, Chien-Pang Lin, Yi-Ping Hung, Tzu-Wen Chang, Neng-Hao Yu, Min-Lun Tsai, Mike Y. Chen |
MUM | 3 |
| 2012 | Illumination Compensation Using Oriented Local Histogram Equalization and its Application to Face RecognitionabstractIllumination compensation and normalization play a crucial role in face recognition. The existing algorithms either compensated low-frequency illumination, or captured high-frequency edges. However, the orientations of edges were not well exploited. In this paper, we propose the orientated local histogram equalization (OLHE) in brief, which compensates illumination while encoding rich information on the edge orientations. We claim that edge orientation is useful for face recognition. Three OLHE feature combination schemes were proposed for face recognition: 1) encoded most edge orientations; 2) more compact with good edge-preserving capability; and 3) performed exceptionally well when extreme lighting conditions occurred. The proposed algorithm yielded state-of-the-art performance on AR, CMU PIE, and extended Yale B using standard protocols. We further evaluated the average performance of the proposed algorithm when the images lighted differently were observed, and the proposed algorithm yielded the promising results. Ping-Han Lee, Szu-Wei Wu, Yi-Ping Hung |
IEEE Trans. Image Process. | 3 |
| 2012 | Subject-Specific and Pose-Oriented Facial Features for Face Recognition Across PosesabstractMost face recognition scenarios assume that frontal faces or mug shots are available for enrollment to the database, faces of other poses are collected in the probe set. Given a face from the probe set, one needs to determine whether a match in the database exists. This is under the assumption that in forensic applications, most suspects have their mug shots available in the database, and face recognition aims at recognizing the suspects when their faces of various poses are captured by a surveillance camera. This paper considers a different scenario: given a face with multiple poses available, which may or may not include a mug shot, develop a method to recognize the face with poses different from those captured. That is, given two disjoint sets of poses of a face, one for enrollment and the other for recognition, this paper reports a method best for handling such cases. The proposed method includes feature extraction and classification. For feature extraction, we first cluster the poses of each subject's face in the enrollment set into a few pose classes and then decompose the appearance of the face in each pose class using Embedded Hidden Markov Model, which allows us to define a set of subject-specific and pose-priented (SSPO) facial components for each subject. For classification, an Adaboost weighting scheme is used to fuse the component classifiers with SSPO component features. The proposed method is proven to outperform other approaches, including a component-based classifier with local facial features cropped manually, in an extensive performance evaluation study. Ping-Han Lee, Gee-Sern Hsu, Yun-Wen Wang, Yi-Ping Hung |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | TUIC: enabling tangible interaction on capacitive multi-touch displaysabstractWe present TUIC, a technology that enables tangible interaction on capacitive multi-touch devices, such as iPad, iPhone, and 3M's multi-touch displays, without requiring any hardware modifications. TUIC simulates finger touches on capacitive displays using passive materials and active modulation circuits embedded inside tangible objects, and can be used with multi-touch gestures simultaneously. TUIC consists of three approaches to sense and track objects: spatial, frequency, and hybrid (spatial plus frequency). The spatial approach, also known as 2D markers, uses geometric, multi-point touch patterns to encode object IDs. Spatial tags are straightforward to construct and are easily tracked when moved, but require sufficient spacing between the multiple touch points. The frequency approach uses modulation circuits to generate high-frequency touches to encode object IDs in the time domain. It requires fewer touch points and allows smaller tags to be built. The hybrid approach combines both spatial and frequency tags to construct small tags that can be reliably tracked when moved and rotated. We show three applications demonstrating the above approaches on iPads and 3M's multi-touch displays. Neng-Hao Yu, Li-Wei Chan 0001, Seng-Yong Lau, Sung-Sheng Tsai, I-Chun Hsiao, Dian-Je Tsai, Fang-I Hsiao, Lung-Pan Cheng, Mike Y. Chen, Polly Huang, Yi-Ping Hung |
CHI | 11 |
| 2011 | Ordinal hyperplanes ranker with cost sensitivities for age estimationabstractIn this paper, we propose an ordinal hyperplane ranking algorithm called OHRank, which estimates human ages via facial images. The design of the algorithm is based on the relative order information among the age labels in a database. Each ordinal hyperplane separates all the facial images into two groups according to the relative order, and a cost-sensitive property is exploited to find better hyperplanes based on the classification costs. Human ages are inferred by aggregating a set of preferences from the ordinal hyperplanes with their cost sensitivities. Our experimental results demonstrate that the proposed approach outperforms conventional multiclass-based and regression-based approaches as well as recently developed ranking-based age estimation approaches. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
CVPR | 3 |
| 2011 | Novel projector calibration approaches of multi-resolution displayabstractThis paper proposes convenient and useful approaches to automatically calibrate the projectors of a multi-resolution display. The proposed approaches estimate both the keystone effect and misalignment of the projections with an assistance of a color camera. Structured light patterns are employed to construct the geometric relationship between projectors and the projection surface, and then pre-warp the images so that they appear undistorted as a result. Experimental results demonstrate that the proposed approaches successfully reduce the human-effort and lower the calibration time of multi-resolution display calibration task. Po-Hsun Chiu, Shih-Yao Lin 0001, Li-Wei Chan 0001, Neng-Hao Yu, Yi-Ping Hung |
ICME | 5 |
| 2011 | Egocentric View Transition for Video Monitoring in a Distributed Camera Network
Kuan-Wen Chen, Pei-Jyun Lee, Yi-Ping Hung |
MMM (1) | 3 |
| 2011 | i - m - Breath: The Effect of Multimedia Biofeedback on Learning Abdominal Breath
Meng-Chieh Yu, Jin-Shing Chen, King-Jen Chang, Su-Chu Hsu, Ming-Sui Lee, Yi-Ping Hung |
MMM (1) | 6 |
| 2011 | Clip-on gadgets: expanding multi-touch interaction area with unpowered tactile controlsabstractVirtual keyboards and controls, commonly used on mobile multi-touch devices, occlude content of interest and do not provide tactile feedback. Clip-on Gadgets solve these issues by extending the interaction area of multi-touch devices with physical controllers. Clip-on Gadgets use only conductive materials to map user input on the controllers to touch points on the edges of screens; therefore, they are battery-free, lightweight, and low-cost. In addition, they can be used in combination with multi-touch gestures. We present several hardware designs and a software toolkit, which enable users to simply attach Clip-on Gadgets to an edge of a device and start interacting with it. Neng-Hao Yu, Sung-Sheng Tsai, I-Chun Hsiao, Dian-Je Tsai, Meng-Han Lee, Mike Y. Chen, Yi-Ping Hung |
UIST | 7 |
| 2011 | Viewpoint-Independent Object Detection Based on Two-Dimensional Contours and Three-Dimensional SizesabstractWe propose a viewpoint-independent object-detection algorithm that detects objects in videos based on their 2-D and 3-D information. Object-specific quasi-3-D templates are proposed and applied to match objects' 2-D contours and to calculate their 3-D sizes. A quasi-3-D template is the contour and the 3-D bounding cube of an object viewed from a certain panning and tilting angle. Pedestrian templates amounting to 2660 and 1995 vehicle templates encompassing 19 tilting and 35 panning angles are used in this study. To detect objects, we first match the 2-D contours of object candidates with known objects' contours, and some object templates with large 2-D contour-matching scores are identified. In this step, we exploit some prior knowledge on the viewpoint on which the object is viewed to speed up the template matching, and the viewpoint likelihood for each contour-matched template is also assigned. Then, we calculate the 3-D widths, heights, and lengths of the contour-matched candidates, as well as the corresponding 3-D-size-matching scores. The overall matching score is obtained by combining the aforementioned likelihood and scores. The major contributions of this paper are to explore the joint use of 2-D and 3-D features in object detection. It shows that, by considering 2-D contours and 3-D sizes, one can achieve promising object detection rates. The proposed algorithms were evaluated on both pedestrian and vehicle sequences. It yielded significantly better detection results than the best results reported in PETS 2009, showing that our algorithm outperformed the state-of-the-art pedestrian-detection algorithms. Ping-Han Lee, Yen-Liang Lin, Shen-Chi Chen, Chia-Hsiang Wu, Cheng-Chih Tsai, Yi-Ping Hung |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2011 | Multi-Resolution Design for Large-Scale and High-Resolution MonitoringabstractLarge-scale and high-resolution monitoring systems are ideal for many visual surveillance applications. However, existing approaches have insufficient resolution and low frame rate per second, or have high complexity and cost. We take inspiration from the human visual system and propose a multi-resolution design, e-Fovea, which provides peripheral vision with a steerable fovea that is in higher resolution. In this paper, we firstly present two user studies, with a total of 36 participants, to compare e-Fovea to two existing multi-resolution visual monitoring designs. The user study results show that for visual monitoring tasks, our e-Fovea design with steerable focus is significantly faster than existing approaches and preferred by users. We then present our design and implementation of e-Fovea, which combines both multi-resolution camera input and multi-resolution steerable projector output. Finally, we present our deployment of e-Fovea in three installations to demonstrate its feasibility. Kuan-Wen Chen, Chih-Wei Lin 0004, Tzu-Hsuan Chiu, Mike Yen-Yang Chen, Yi-Ping Hung |
IEEE Trans. Multim. | 5 |
| 2011 | Adaptive Learning for Target Tracking and True Linking Discovering Across Multiple Non-Overlapping CamerasabstractTo track targets across networked cameras with disjoint views, one of the major problems is to learn the spatio-temporal relationship and the appearance relationship, where the appearance relationship is usually modeled as a brightness transfer function. Traditional methods learning the relationships by using either hand-labeled correspondence or batch-learning procedure are applicable when the environment remains unchanged. However, in many situations such as lighting changes, the environment varies seriously and hence traditional methods fail to work. In this paper, we propose an unsupervised method which learns adaptively and can be applied to long-term monitoring. Furthermore, we propose a method that can avoid weak links and discover the true valid links among the entry/exit zones of cameras from the correspondence. Experimental results demonstrate that our method outperforms existing methods in learning both the spatio-temporal and the appearance relationship, and can achieve high tracking accuracy in both indoor and outdoor environment. Kuan-Wen Chen, Chih-Chuan Lai, Pei-Jyun Lee, Chu-Song Chen, Yi-Ping Hung |
IEEE Trans. Multim. | 5 |
| 2011 | Editing by Viewing: Automatic Home Video Summarization by Viewing Behavior AnalysisabstractIn this paper, we propose the Interest Meter (IM), a system making the computer conscious of user's reactions to measure user's interest and thus use it to conduct video summarization. The IM takes account of users' spontaneous reactions when they view videos. To estimate user's viewing interest, quantitative interest measures are devised based on the perspectives of attention and emotion. For estimating attention states, variations of user's eye movement, blink, and head motion are considered. For estimating emotion states, facial expression is recognized as positive or neural emotion. By combining characteristics of attention and emotion by a fuzzy fusion scheme, we transform users' viewing behaviors into quantitative interest scores, determine interesting parts of videos, and finally concatenate them as video summaries. Experimental results show that the proposed concept “editing by viewing” works well and may provide a promising direction to consider the human factor in video summarization. Wei-Ting Peng, Wei-Ta Chu, Chia-Han Chang, Chien-Nan Chou, Wei-Jia Huang, Wen-Yan Chang, Yi-Ping Hung |
IEEE Trans. Multim. | 7 |
| 2010 | Touching the void: direct-touch interaction for intangible displaysabstractIn this paper, we explore the challenges in applying and investigate methodologies to improve direct-touch interaction on intangible displays. Direct-touch interaction simplifies object manipulation, because it combines the input and display into a single integrated interface. While traditional tangible display-based direct-touch technology is commonplace, similar direct-touch interaction within an intangible display paradigm presents many challenges. Given the lack of tactile feedback, direct-touch interaction on an intangible display may show poor performance even on the simplest of target acquisition tasks. In order to study this problem, we have created a prototype of an intangible display. In the initial study, we collected user discrepancy data corresponding to the interpretation of 3D location of targets shown on our intangible display. The result showed that participants performed poorly in determining the z-coordinate of the targets and were imprecise in their execution of screen touches within the system. Thirty percent of positioning operations showed errors larger than 30mm from the actual surface. This finding triggered our interest to design a second study, in which we quantified task time in the presence of visual and audio feedback. The pseudo-shadow visual feedback was shown to be helpful both in improving user performance and satisfaction. Li-Wei Chan 0001, HuiShan Kao, Mike Y. Chen, Ming-Sui Lee, Yung-Jen Hsu 0001, Yi-Ping Hung |
CHI | 6 |
| 2010 | Robust Face Recognition Using Probabilistic Facial Trait Code
Ping-Han Lee, Gee-Sern Hsu, Szu-Wei Wu, Yi-Ping Hung |
ECCV (1) | 4 |
| 2010 | Probabilistic Facial Trait Code for face recognitionabstractRecently, a new facial encoding scheme, namely Facial Trait Code (FTC), was proposed. FTC encoded human faces into a series of integers. Distances between codewords of different people were maximized during the code construction. FTC was applied to solve face recognition, and was reported with promising verification rates. However, due to several simplifications in the FTC encoding, its performance degraded considerably when there were only few images per individual available for enrollment in the gallery sets, or when the probe set included faces under large variations in illumination and expression. In this paper, we proposed the Probabilistic Facial Trait Code (PFTC) with a novel encoding scheme and a probabilistic codeword distance measure. The impact made by illumination and expression variations were also handled in the construction of PFTC. The proposed PFTC was evaluated and compared with state-of-the-art algorithms, including the FTC, the algorithm using sparse representation, and the one using Local Binary Pattern. PFTC out-performed the algorithms compared in this study in most scenarios. Ping-Han Lee, Szu-Wei Wu, Gee-Sern Hsu, Yi-Ping Hung |
ICIP | 4 |
| 2010 | Turning Rust into Gold: An ancient artifact as an interactive artworkabstractTurning Rust into Gold is inspired by a Chinese antique Mao-Kung Ting (cauldron) treasured by the National Palace Museum in Taiwan. Having a five-hundred-character inscription cast inside, and its weathered appearance made the Mao-Kung very unique. Motivated by revealing the great nature of the artifact and interpreting it into a meaningful narrative, we have proposed an interactive multimedia system that facilitates effective communication between museum audiences and the Mao-Kung Ting. Three technologies have been implemented to emphasize the weathered appearance of the bronze. De-/weathering simulation techniques have been deployed to revive the bronze to its original shiny gold color; while breath-based biofeedback and haptic technology have been utilized as user interfaces to trigger the de-weathering process of the Mao-Kung Ting. Also, the interactive scenarios have been designed with the Chinese cultural context and philosophy Qi, enabling users more easily fall into the Chinese civilization. The paper aims to present the development of the artwork Turing Rust into Gold, in order to further contribute to the feasibility of incorporating new media art in a historical museum context, and bring a new horizon in the museum sector. Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Liang-Chun Lin, I-Ling Liu, Meng-Chieh Yu, Chu-Song Chen, Jiaping Wang |
ICME | 3 |
| 2010 | A real-time user Interest Meter and its applications in home video summarizingabstractIn this paper, we propose the Interest Meter (IM), a system making computer conscious of user's reactions, to measure user's interest in real time. The Interest Meter takes account of users' spontaneous reactions when users interact with computers. In this work, we analyze variations of user's eye movement, blink, head motion, and facial expression. Furthermore, we propose an algorithm to combine those signals into interest score and determine important parts of video shots when people watch raw home videos. Experimental result shows that this new type of editing mechanism can effectively generate home video summaries. Wei-Ting Peng, Chia-Han Chang, Wei-Ta Chu, Wei-Jia Huang, Chien-Nan Chou, Wen-Yan Chang, Yi-Ping Hung |
ICME | 7 |
| 2010 | A Ranking Approach for Human Ages Estimation Based on Face ImagesabstractIn our daily life, it is much easier to distinguish which person is elder between two persons than how old a person is. When inferring a person's age, we may compare his or her face with many people whose ages are known, resulting in a series of comparative results, and then we conjecture the age based on the comparisons. This process involves numerous pairwise preferences information obtained by a series of queries, where each query compares the target person's face to those faces in a database. In this paper, we propose a ranking-based framework consisting of a set of binary queries. Each query collects a binary-classification-based comparison result. All the query results are then fused to predict the age. Experimental results show that our approach performs better than traditional multi-class-based and regression-based approaches for age estimation. Kuang-Yu Chang, Chu-Song Chen, Yi-Ping Hung |
ICPR | 3 |
| 2010 | Multi-Cue Integration for Multi-Camera TrackingabstractFor target tracking across multiple cameras with disjoint views, previous works usually employed multiple cues and focused on learning a better matching model of each cue, separately. However, none of them had discussed how to integrate these cues to improve performance, to our best knowledge. In this paper, we look into the multi-cue integration problem and propose an unsupervised learning method since a complicated training phase is not always viable. In the experiments, we evaluate several types of score fusion methods and show that our approach learns well and can be applied to large camera networks more easily. Kuan-Wen Chen, Yi-Ping Hung |
ICPR | 2 |
| 2010 | Automatic Gender Recognition Using Fusion of Facial StripsabstractWe propose a fully automatic system that detects and normalizes faces in images and recognizes their genders. To boost the recognition accuracy, we correct the in-plane and out-of-plane rotations of faces, and align faces based on estimated eye positions. To perform gender recognition, a face is first decomposed into several horizontal and vertical strips. Then, a regression function for each strip gives an estimation of the likelihood the strip sample belongs to a specific gender. The likelihoods from all strips are concatenated to form a new feature, based on which a gender classifier gives the final decision. The proposed approach achieved an accuracy of 88.1% in recognizing genders of faces in images collected from the World-Wide Web. For faces in the FERET dataset, our system achieved an accuracy of 98.8%, outperforming all the six state-of-the-art algorithms compared in this paper. Ping-Han Lee, Jui-Yu Hung, Yi-Ping Hung |
ICPR | 3 |
| 2010 | Real-Time 3D Model-Based Gesture Tracking for Multimedia ControlabstractThis paper presents a new 3D model-based gesture tracking system for controlling multimedia player in an intuitive way. The motivation of this paper is to make home appliance aware of user's intention. This 3D model-based gesture tracking system adopts a Bayesian framework to track the user's 3D hand position and to recognize meaning of these postures for controlling 3D player interactively. To avoid the high dimensionality of the whole 3D upper body model, which may complicate the gesture tracking problem, our system applies a novel hierarchical tracking algorithm to improve the system performance. Moreover, this system applies multiple cues for improving the accuracy of tracking results. Based on the above idea, we have implemented a 3D hand gesture interface for controlling multimedia players. Experimental results have shown that the proposed system robustly tracks the 3D position of the hand and has high potential for controlling the multimedia player. Shih-Yao Lin 0001, Yun-Chien Lai, Li-Wei Chan 0001, Yi-Ping Hung |
ICPR | 4 |
| 2010 | e-Fovea: a multi-resolution approach with steerable focus to large-scale and high-resolution monitoringabstractThis paper presents e-Fovea, a system that combines both multi-resolution camera input and multi-resolution steerable projector output to support large-scale and high-resolution visual monitoring. e-Fovea utilizes a design similar to the human eyes, which provides peripheral vision with a steerable fovea that is in higher resolution. e-Fovea is implemented using a steerable telephoto camera and a wide-angle camera. The telephoto image is displayed using a projector with a steerable mirror, and overlaid on the wide-angle image that is displayed using a second projector. Kuan-Wen Chen, Chih-Wei Lin 0004, Mike Y. Chen, Yi-Ping Hung |
ACM Multimedia | 4 |
| 2010 | Yongzheng emperor's interactive tabletop: seamless multimedia system in a museum contextabstractIn this paper, we propose the seamless multimedia system Yongzheng Emperor's interactive tabletop, which has been incorporated into the special exhibition "Harmony and Integrity: The Yongzheng Emperor and His Times" at the National Palace Museum in Taiwan. The multimedia system features the innovative use of physical artifacts - Yongzheng figurines and a model of Yongzheng-era calendar clock as tangible user interfaces, which have been used to activate on the Surface the emperor's life at court and chronological events of his times. Museum audiences can naturally and intuitively explore the Emperor's stories by the use of hand gestures. The system vividly connects the modern world of the audiences with the emperor's virtual world, engaging museum audiences in the most interactive and compelling way to learn about the emperor. The paper aims to present the development of the seamless tabletop system in a historical museum context, including design principles, implementation, applications, and effectiveness of the system. Our contribution in this project is to demonstrate a new exhibition display model for the museum sector. Chun-Ko Hsieh, I-Ling Liu, Neng-Hao Yu, Yueh-Hsuan Chiang, Hsiang-Tao Wu, Ying-Jui Chen, Yi-Ping Hung |
ACM Multimedia | 7 |
| 2010 | i-m-Space: interactive multimedia-enhanced space for rehabilitation of breast cancer patientsabstractThis paper presents i-m-Space, an interactive multimedia rehabilitation space that helps the post-surgery recovery of breast cancer patients. Our goal is to improve patients' physical therapy and psychological relaxation experience through careful applications of multimedia technology. i-m-Space consists of three types of breathing-based relaxation and three types for interactive exercise-based rehabilitation. Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Jin-Yao Lin, Szu-Wei Wu, Yi-Yu Chung, I-Ling Hu, Wei-Ting Peng, Shih-Yao Lin 0001, Chia-Han Chang, Pei-Hsuan Chou, King-Jen Chang, Mei-Lan Chang, Sue-Huei Chen, Jin-Shing Chen, Ming-Sui Lee, Mike Y. Chen, Yi-Ping Hung |
ACM Multimedia | 19 |
| 2010 | Multi-display map touring with tangible widgetabstractMany map systems are created to help the user finding a place or define a route to follow. Google Map extends the concept of "surfing the map" by adding a street view that allows the user to explore a place from real pictures, creating the same feeling of walking through the streets. The horizontal 2D map and vertical panoramic street view, however, cause usability problems, while operating with traditional computer mouse and keyboards and presenting by single vertical or horizontal display. This paper presents a new table system composed of a horizontal tabletop screen and a vertical screen. The map view and the street view are displayed on the horizontal and vertical displays of our system respectively. Users can place the tangible pawn on the 2D map to have direct access of the street view from the pawn's point of view. In the user study, we compare our system with a standard computer system in the navigation task. The results reported that our system improves the intuitiveness of use, efficiency of city exploring and ease of remembrance on spaces that are not familiar beforehand. We also discuss limitations of using tangible objects for map navigation. Marco Piovesana, Ying-Jui Chen, Neng-Hao Yu, Hsiang-Tao Wu, Li-Wei Chan 0001, Yi-Ping Hung |
ACM Multimedia | 6 |
| 2010 | Transformational Breathing between Present and Past: Virtual Exhibition System of the Mao-Kung Ting
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Szu-Wei Wu, Yi-Yu Chung, Liang-Chun Lin, Ming-Sui Lee, Chu-Song Chen, Jiaping Wang, Quo-Ping Lin, I-Ling Liu |
MMM | 3 |
| 2010 | Enabling beyond-surface interactions for interactive surface with an invisible projectionabstractThis paper presents a programmable infrared (IR) technique that utilizes invisible, programmable markers to support interaction beyond the surface of a diffused-illumination (DI) multi-touch system. We combine an IR projector and a standard color projector to simultaneously project visible content and invisible markers. Mobile devices outfitted with IR cameras can compute their 3D positions based on the markers perceived. Markers are selectively turned off to support multi-touch and direct on-surface tangible input. The proposed techniques enable a collaborative multi-display multi-touch tabletop system. We also present three interactive tools: i-m-View, i-m-Lamp, and i-m-Flashlight, which consist of a mobile tablet and projectors that users can freely interact with beyond the main display surface. Early user feedback shows that these interactive devices, combined with a large interactive display, allow more intuitive navigation and are reportedly enjoyable to use. Li-Wei Chan 0001, Hsiang-Tao Wu, HuiShan Kao, Ju-Chun Ko, Home-Ru Lin, Mike Y. Chen, Yung-Jen Hsu 0001, Yi-Ping Hung |
UIST | 8 |
| 2009 | Polymorphous Facial Trait Code
Ping-Han Lee, Gee-Sern Hsu, Yi-Ping Hung |
ACCV (3) | 3 |
| 2009 | Treasure Transformers: Novel Interpretative Installations for the National Palace Museum
Chun-Ko Hsieh, I-Ling Liu, Quo-Ping Lin, Li-Wei Chan 0001, Chuan-Heng Hsiao, Yi-Ping Hung |
ArtsIT | 6 |
| 2009 | To move or not to move: a comparison between steerable versus fixed focus region paradigms in multi-resolution tabletop display systemsabstractPrevious studies have outlined the advantages of multi-resolution large-area displays over their fixed-resolution counterparts, however the mobility of the focus region has up until the present time received little attention. To study this phenomenon further, we have developed a multi-resolution tabletop display system with a steerable high resolution focus region to compare the performance between steerable and fixed focus region systems under different working scenarios. We have classified these scenarios according to region of interest (ROI) with analogies to different eye movement types (fixed, saccadic, and pursuit ROI). Empirical data gathered during the course of a multi-faceted user study demonstrates that the steerable focus region system significantly outperforms the fixed focus region system. The former is shown to provide enhanced display manipulation and proves especially advantageous in cases where the user must maintain spatial awareness of the display content as is the case in which, within a single session, several regions of the display are to be visited. Chuan-Heng Hsiao, Li-Wei Chan 0001, Ting-Ting Hu, Mon-Chu Chen, Yung-Jen Hsu 0001, Yi-Ping Hung |
CHI | 6 |
| 2009 | Face verification and identification using Facial Trait CodeabstractWe propose the Facial Trait Code (FTC) to encode human facial images. The proposed FTC is motivated by the discovery of some basic patterns existing in certain local facial features. We call these basic patterns Distinctive Trait Patterns (DTP), which can be extracted from a large number of faces. We have also found that the fusion of these DTP's can accurately capture the appearance of a face. The extraction of DTP involves clustering and boosting for maximizing the discrimination between human faces. The extracted DTP's can be symbolized and used to make up the n-ary facial trait codes. A given face can be encoded at some prescribed facial traits to render an n-ary facial trait code with each symbol in its codeword corresponding to the closest DTP. We applied FTC to a face identification and verification problems with 3575 facial images from 840 people under different illumination conditions, and it yielded satisfactory results. Ping-Han Lee, Gee-Sern Hsu, Yi-Ping Hung |
CVPR | 3 |
| 2009 | Real-time pedestrian and vehicle detection in video using 3D cuesabstractExisting pedestrian and vehicle detection algorithms use 2D cues of objects, such as pixel values, color, texture, shape information or motion. The use of 3D cues in object detection, on the other hand, is not well studied in the literature. In this paper, we propose an efficient algorithm that detects pedestrian and vehicle using their 3D cues. The proposed algorithm first detects moving objects in a video frame using a background modeling technique. For each moving object, we extract its width and height in 3D space, with the aid of the intrinsic and extrinsic parameters of the camera monitoring the scene. To estimate the camera parameters, we apply a calibration-free method, which simply requires users to specify six vertices on a cuboid in the scene. Then based on its 3D cues, a object is verified whether it is a pedestrian(vehicle) or not by the class-specific Support Vector Machine (SVM). In our experiment, the proposed algorithm achieves a precision of 88.2% (89.1%) for pedestrian(vehicle) detection, at 32 frame-per-second on average upon five testing sequences. Ping-Han Lee, Tzu-Hsuan Chiu, Yen-Liang Lin, Yi-Ping Hung |
ICME | 4 |
| 2009 | Face synthesis using Facial Trait Code and its application to creating suspect's physical profilesabstractIn this work, we propose a novel approach to synthesize human faces. The proposed approach is based on the facial trait code, or FTC for short, which is originally proposed for solving the face recognition problem. The FTC decoding transforms codewords into faces. We find that the FTC code space is typically huge and only a very small subset of FTC codewords render smooth faces. We propose an algorithm which extracts codewords that render smooth faces efficiently, and we dub these codewords as legitimate codewords. The extraction of legitimate codewords is accompanied by a coarse-to-fine face synthesis scheme, which renders faces by mosaicking multiple patches of different facial parts. We implement a GUI to demonstrate that the proposed approach is able to help eyewitnesses to create suspect's physical profiles efficiently in the law enforcement environments, without the aid of face painters. Ping-Han Lee, Yi-Ping Hung |
ICME | 2 |
| 2009 | Generating pictorial-based representation of mental images for video monitoringabstractMulti-camera systems have been widely used in many video surveillance applications. When an event happens and is monitored across multiple cameras, it is easy for an expert to generate the corresponding spatial representation to comprehend the series of event. However, it is not trivial for users new to the environment. With support from psychological evidences, we propose an approach to mimic generating pictorial-based representation of mental images when a target is moving across the views of cameras. First we conduct a ball-rolling experiment to compare this approach with others. The empirical results demonstrate that the performance of users with this approach is significantly better than others. We suggest that it is because this approach is better for users to preserve spatial representation of the environment while transiting views between cameras. Then we propose a framework to realize this approach. The demonstrations in different situations indicate the validity of such framework. Chuan-Heng Hsiao, Wei-Chia Huang, Kuan-Wen Chen, Li-Wei Chang, Yi-Ping Hung |
IUI | 5 |
| 2009 | i-m-Tube: an interactive multi-resolution tubular displayabstractIn this paper we introduce a tubular interface, the i-m-Tube, which provides convenient user interaction with multimedia content by multi-touch input and multi-resolution display. With its tubular surface, the i-m-Tube is suitable for displaying panoramic image content like the Chinese scroll painting "Along the River During the Ch'ingming Festival" which has been regarded as a national treasure and widely known by its extraordinary width and its various details. The strength of our system is that it is not just suitable for displaying panoramic content, but also possible to create an intuitive and natural context for user interaction with multimedia content displayed on the i-m-Tube by multi-finger touch. It also provides a high resolution area on demand, which is projected by a steerable projector in a Full-HD resolution. We have developed our system as an interface which can be applied adaptively to different applications, having the features of multi-touch and multi-resolution interactivity on a tubular display. In this work, we have implemented several prototypes designed for different applications, like "Tetris360" aiming on co-op arcade gaming, and "Panoramic Cover Flow" for interactive digital signage, to show the capabilities and opportunities of the i-m-Tube display system. With our preliminary experiments on the i-m-Tube, it is exciting to see the potential of the i-m-Tube as a new interface for people to meet the interactive multimedia touch in this generation. Jin-Yao Lin, Ju-Chun Ko, HuiShan Kao, Tsun-Hung Tsai, Su-Chu Hsu, Yi-Ping Hung |
ACM Multimedia | 8 |
| 2009 | A User Experience Model for Home Video Summarization
Wei-Ting Peng, Wei-Jia Huang, Wei-Ta Chu, Chien-Nan Chou, Wen-Yan Chang, Chia-Han Chang, Yi-Ping Hung |
MMM | 7 |
| 2009 | Multicast communication in wormhole-routed 2D torus networks with hamiltonian cycle model
Neng-Chung Wang, Yi-Ping Hung |
J. Syst. Archit. | 2 |
| 2009 | Region-based image retrieval using color-size features of watershed regions
Cheng-Chieh Chiang, Yi-Ping Hung, Hsuan Yang, Greg C. Lee |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Tracking by Parts: A Bayesian Approach With Component CollaborationabstractInstead of using global-appearance information for visual tracking, as adopted by many methods, we propose a tracking-by-parts (TBP) approach that uses partial appearance information for the task. The proposed method considers the collaborations between parts and derives a probability propagation framework by encoding the spatial coherence in a Bayesian formulation. To resolve this formulation, a TBP particle-filtering method is introduced. Unlike existing methods that only use the spatial-coherence relationship for particle-weight estimation, our method further applies this relationship for state prediction based on system dynamics. Thus, the part-based information can be utilized efficiently, and the tracking performance can be improved. Experimental results show that our approach outperforms the factored-likelihood and particle reweight methods, which only use spatial coherence for weight estimation. Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | An adaptive learning method for target tracking across multiple camerasabstractThis paper proposes an adaptive learning method for tracking targets across multiple cameras with disjoint views. Two visual cues are usually employed for tracking targets across cameras: spatio-temporal cue and appearance cue. To learn the relationships among cameras, traditional methods used batch-learning procedures or hand-labeled correspondence, which can work well only within a short period of time. In this paper, we propose an unsupervised method which learns both spatio-temporal relationships and appearance relationships adaptively and can be applied to long-term monitoring. Our method performs target tracking across multiple cameras while also considering the environment changes, such as sudden lighting changes. Also, we improve the estimation of spatio-temporal relationships by using the prior knowledge of camera network topology. Kuan-Wen Chen, Chih-Chuan Lai, Yi-Ping Hung, Chu-Song Chen |
CVPR | 3 |
| 2008 | Localization and mapping of surveillance cameras in city mapabstractMany large cities have installed surveillance cameras to monitor human activities for security purposes. An important surveillance application is to track the motion of an object of interest, e.g., a car or a human, using one or more cameras, and plot the motion path in a city map. To achieve this goal, it is necessary to localize the cameras in the city map and to determine the correspondence mappings between the positions in the city map and the camera views. Since the view of the city map is roughly orthogonal to the camera views, there are very few common features between the two views for a computer vision algorithm to correctly identify corresponding points automatically. This paper proposes a method for camera localization and position mapping that requires minimum user inputs. Given approximate corresponding points between the city map and a camera view identified by a user, the method computes the orientation and position of the camera in the city map, and determines the mapping between the positions in the city map and the camera view. Both quantitative tests and practical application test have been performed. It can obtain the best-fit solutions even though the user-specified correspondence is inaccurate. The performance of the method is assessed in both quantitative tests and practical application. Quantitative test results show that the method is accurate and robust in camera localization and position mapping. Application test results are very encouraging, showing the usefulness of the method in real applications. Wee Kheng Leow, Cheng-Chieh Chiang, Yi-Ping Hung |
ACM Multimedia | 3 |
| 2008 | Interactive content presentation based on expressed emotion and physiological feedbackabstractIn this technical demonstration, we showcase an interactive content presentation (ICP) system that integrates media-expressed-emotion-based composition, user-perceived preference feedback, and interactive digital art creation. ICP harmonizes the browsing of multimedia contents by presenting them in the form of music videos (photos, blog articles with accompanied music) based on their expressed emotion similarity. ICP facilitates content browsing by automatically and dynamically selecting the media to be played next in real time, responding to user's preference feedback measured from physiological signals. In addition, ICP enhances the enjoyments of content browsing by incorporating interactive digital art creation. ICP achieves these goals by properly integrating recent researches on media-expressed emotion classification,cross-media composition, and physiological signal processing. Tien-Lin Wu, Hsuan-Kai Wang, Murphy Chien-Chang Ho, Yuan-Pin Lin, Ting-Ting Hu, Ming-Fang Weng, Li-Wei Chan 0001, Changhua Yang, Yi-Hsuan Yang, Yi-Ping Hung, Yung-Yu Chuang, Hsin-Hsi Chen, Homer H. Chen, Jyh-Horng Chen, Shyh-Kang Jeng |
ACM Multimedia | 10 |
| 2008 | Aesthetics-Based Automatic Home Video Skimming System
Wei-Ting Peng, Yueh-Hsuan Chiang, Wei-Ta Chu, Wei-Jia Huang, Wei-Lun Chang, Po-Chung Huang, Yi-Ping Hung |
MMM | 7 |
| 2007 | Analyzing Facial Expression by Fusing Manifolds
Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
ACCV (2) | 3 |
| 2007 | i-m-Top: An Interactive Multi-Resolution Tabletop Display SystemabstractThis article presents "i-m-Top", a rear-projection interactive multi-resolution tabletop display system. This table allows the user to perform regular and long-term works such as reading and writing of documents. The design of the tabletop system well meets the following concerns: (1) provide movable high-resolution perception of the tabletop, (2) avoid the shadowing effect by rear projection, (3) reserve space for the user to place their legs with comfort and (4) allow interactions on the tabletop by bare hands and a pen. Li-Wei Chan 0001, Yi-Wei Chia, Yi-Fan Chuang, Yi-Ping Hung |
ICME | 4 |
| 2007 | Cascading Multimodal Verification using Face, Voice and Iris InformationabstractIn this paper we propose a novel fusion strategy which fuses information from multiple physical traits via a cascading verification process. In the proposed system users are verified by each individual modules sequentially in turns of face, voice and iris, and would be accepted once he/she is verified by one of the modules without performing the rest of the verifications. Through adjusting thresholds for each module, the proposed approach exhibits different behavior with respect to security and user convenience. We provide a criterion to select thresholds for different requirements and we also design an user interface which helps users find the personalized thresholds intuitively. The proposed approach is verified with experiments on our in-house face-voice-iris database. The experimental results indicate that besides the flexibility between security and convenience, the proposed system also achieves better accuracy than its most accurate module. Ping-Han Lee, Lu-Jong Chu, Yi-Ping Hung, Sheng-Wen Shih, Chu-Song Chen, Hsin-Min Wang |
ICME | 3 |
| 2007 | Virtual Conduction System with Multi-Resolution Wall DisplayabstractThe virtual conduction system (VCS) allows a user to conduct a photo-realistic pseudo orchestra. The VCS includes four modules: (1) gesture recognition module, (2) audio rendering module, (3) video rendering module, and (4) multi-resolution display module. With gesture recognition and tempo adjustment, the user not only can change the playback rate of an audio and video recording, but also can control the volume from different portion of an orchestra in real time. With the video rendering module and multi-resolution display module, the VCS provides a new visual experience with an interactive multi-resolution wall-size display. Wei-Ting Peng, En-Wei Huang, Wei-Lun Chang, Po-Chung Huang, Jun-Ying Bai, Han-Ru Chen, Shao-Yi Chien, Shyh-Kang Jeng, Yi-Ping Hung, Li-Chen Fu, Lin-Shan Lee |
ICME | 9 |
| 2007 | An Image-Based Approach to Interactive 3D Virtual Exhibition
Yi-Ping Hung |
PSIVT | 1 |
| 2007 | Gesture-based interaction for a magic crystal ballabstractCrystal balls are generally considered as media to perform divination or fortune-telling. These imaginations are mainly from some fantasy films and fiction, in which an augur can see into the past, the present, or the future through a crystal ball. With the distinct impressions, crystal ball has revealed itself as a perfect interface for the users to access and to manipulate visual media in an intuitive, imaginative and playful manner. We developed an interactive visual display system named Magic Crystal Ball (MaC Ball). MaC Ball is a spherical display system, which allows the users to see a virtual object/scene appearing inside a transparent sphere, and to manipulate the displayed content with barehanded interactions. Interacting with MaC Ball makes the users feeling acting with magic power. With MaC Ball, user can manipulate the display with touch and hover interactions. For instance, the user waves hands above the ball, causing clouds blowing from bottom of the ball, or slides fingers on the ball to rotate the displayed object. In addition, the user can press single finger to select an object or to issue a button. MaC Ball takes advantages on the impressions of crystal balls, allowing the users acting with visual media following their imaginations. For applications, MaC Ball has high potential to be used for advertising and demonstration in museums, product launches, and other venues. Li-Wei Chan 0001, Yi-Fan Chuang, Meng-Chieh Yu, Yi-Liu Chao, Ming-Sui Lee, Yi-Ping Hung, Yung-Jen Hsu 0001 |
VRST | 6 |
| 2007 | Efficient hierarchical method for background subtraction
Chu-Song Chen, Chun-Rong Huang, Yi-Ping Hung |
Pattern Recognit. | 4 |
| 2007 | Fast and versatile algorithm for nearest neighbor search based on a lower bound tree
Yong-Sheng Chen, Yi-Ping Hung, Ting-Fang Yen, Chiou-Shann Fuh |
Pattern Recognit. | 2 |
| 2007 | Background Removal of Multiview Images by Learning Shape PriorsabstractImage-based rendering has been successfully used to display 3-D objects for many applications. A well-known example is the object movie, which is an image-based 3-D object composed of a collection of 2-D images taken from many different viewpoints of a 3-D object. In order to integrate image-based 3-D objects into a chosen scene (e.g., a panorama), one has to meet a hard challenge--to efficiently and effectively remove the background from the foreground object. This problem is referred to as multiview images (MVIs) segmentation. Another task requires MVI segmentation is image-based 3-D reconstruction using multiview images. In this paper, we propose a new method for segmenting MVI, which integrates some useful algorithms, including the well-known graph-cut image segmentation and volumetric graph-cut. The main idea is to incorporate the shape prior into the image segmentation process. The shape prior introduced into every image of the MVI is extracted from the 3-D model reconstructed by using the volumetric graph cuts algorithm. Here, the constraint obtained from the discrete medial axis is adopted to improve the reconstruction algorithm. The proposed MVI segmentation process requires only a small amount of user intervention, which is to select a subset of acceptable segmentations of the MVI after the initial segmentation process. According to our experiments, the proposed method can provide not only good MVI segmentation, but also provide acceptable 3-D reconstructed models for certain less-demanding applications. Yu-Pao Tsai, Cheng-Hung Ko, Yi-Ping Hung, Zen-Chung Shih |
IEEE Trans. Image Process. | 3 |
| 2006 | Augmented Stereo Panoramas
Chien-Wei Chen, Li-Wei Chan 0001, Yu-Pao Tsai, Yi-Ping Hung |
ACCV (1) | 4 |
| 2006 | A Method for Calibrating a Motorized Object Rig
Pang-Hung Huang, Yu-Pao Tsai, Wan-Yen Lo, Sheng-Wen Shih, Chu-Song Chen, Yi-Ping Hung |
ACCV (1) | 6 |
| 2006 | Integration of Background Modeling and Object TrackingabstractBackground model and tracking became critical components for many vision-based applications. Typically, background modeling and object tracking are mutually independent in many approaches. In this paper, we adopt a probabilistic framework that uses particle filtering to integrate these two approaches, and the observation model is measured by Bhattacharyya distance. Experimental results and quantitative evaluations show that the proposed integration framework is effective for moving object detection Chu-Song Chen, Yi-Ping Hung |
ICME | 3 |
| 2006 | Automatic Geometric and Photometric Calibration for Tiling Multiple Projectors with a Pan-Tilt-Zoom CameraabstractResearch on creating a large, high-resolution, low-cost display system has become increasingly important due to the growing desire in many fields for bigger and better displays. The goal of this research is to build a seamless large-scale display system by tiling multiple projectors with the help of a pan-tilt-zoom camera. In order to achieve this goal, we took two tasks into consideration: geometric calibration and photometric calibration. Compared to the previous work, our method for geometric calibration is more accurate, thanks to the much higher resolution images acquired by combining several zoom-in images and doing lens correction for the camera at first. For photometric calibration, we achieve color uniformity among projectors by utilizing the same PTZ camera, which is calibrated once with a colorimeter in a factory. Furthermore, we adopted the technique of producing a high dynamic range image from several images of different exposures to increase the measurement accuracy of the camera. In our experiments, the average error of photometric measurement achieved by our method is less than the mean perceptibility tolerance. Our method has great potential for many applications that require large and high-resolution displays Yu-Pao Tsai, Yen-nien Wu, Shou-Chun Liao, Zen-Chung Shih, Yi-Ping Hung |
ICME | 5 |
| 2006 | Real Time Large Motion Feature Tracking by Matching Characteristic Curves via Dynamic ProgrammingabstractThis paper addresses real time feature tracking when the frame-to-frame motion is large in the video. To solve this problem, we first estimate the locally piecewise, translational motion by matching the horizontal and vertical characteristic curves of the consecutive images. Then, we incorporate the motion estimates into the pyramidal Kanade-Lucas-Tomasi (KLT) feature tracker to accomplish the tracking task. To compute the motion estimates efficiently and effectively, we use dynamic programming to minimize the cost function. These motion estimates will serve as the coarse motion at the deepest pyramid level which makes the residual motion small enough such that the feature tracker can work well. In addition, we introduce a feature rejection method that improves the efficiency. Experiments show that our method can make the feature tracker suitable to track features with large motion. Cheng-Hung Ko, Yi-Ping Hung |
SMC | 2 |
| 2006 | Distinctive Personal Traits for Face Recognition Under OcclusionabstractExisting local feature methods for face recognition utilize visually salient regions around eye, nose, and mouth to model the characteristics of a person. The premise of such an approach is that there exists a set of features that are common in all human faces and yet distinct to tell one from the rest apart. In this paper we present an algorithm that selects the best set of features or templates for each individual, and uses these distinct personal traits to boost face recognition performance even when they are partially occluded. Borne out by numerous experiments and comparisons, we demonstrate that the proposed method is effective in recognizing faces with partial occlusion and variation in expression. Ping-Han Lee, Yun-Wen Wang, Ming-Hsuan Yang 0001, Gee-Sern Hsu, Yi-Ping Hung |
SMC | 5 |
| 2006 | A Flexible Display by Integrating a Wall-Size Display and Steerable Projectors
Li-Wei Chan 0001, Wei-Shian Ye, Shou-Chun Liao, Yu-Pao Tsai, Yung-Jen Hsu 0001, Yi-Ping Hung |
UIC | 6 |
| 2006 | Comparison between immersion-based and toboggan-based watershed image segmentationabstractWatershed segmentation has recently become a popular tool for image segmentation. There are two approaches to implementing watershed segmentation: immersion approach and toboggan simulation. Conceptually, the immersion approach can be viewed as an approach that starts from low altitude to high altitude and the toboggan approach as an approach that starts from high altitude to low altitude. The former seemed to be more popular recently (e.g., Vincent and Soille), but the latter had its own supporters (e.g., Mortensen and Barrett). It was not clear whether the two approaches could lead to exactly the same segmentation result and which approach was more efficient. In this paper, we present two "order-invariant" algorithms for watershed segmentation, one based on the immersion approach and the other on the toboggan approach. By introducing a special RIDGE label to achieve the property of order-invariance, we find that the two conceptually opposite approaches can indeed obtain the same segmentation result. When running on a Pentium-III PC, both of our algorithms require only less than 1/30 s for a 256 x 256 image and 1/5 s for a 512 x 512 image, on average. What is more surprising is that the toboggan algorithm, which is less well known in the computer vision community, turns out to run faster than the immersion algorithm for almost all the test images we have used, especially when the image is large, say, 512 x 512 or larger. This paper also gives some explanation as to why the toboggan algorithm can be more efficient in most cases. Yung-Chieh Lin, Yu-Pao Tsai, Yi-Ping Hung, Zen-Chung Shih |
IEEE Trans. Image Process. | 3 |
| 2005 | Appearance-Guided Particle Filtering for Articulated Hand TrackingabstractWe propose a model-based tracking method, called appearance-guided particle filtering (AGPF), which integrates both sequential motion transition information and appearance information. A probability propagation model is derived from a Bayesian formulation for this framework, and a sequential Monte Carlo method is introduced for its realization. We apply the proposed method to articulated hand tracking, and show that it performs better than methods that only use either sequential motion transition information or only use appearance information. Wen-Yan Chang, Chu-Song Chen, Yi-Ping Hung |
CVPR (1) | 3 |
| 2005 | An Efficient Approach to Multimodal Person Identity Verification by Fusing Face and Voice InformationabstractThis paper presents an effective method to combine speech recognition, speaker verification and face verification for biometric authentication. Our method provides a light-weight enrollment process and an easy-to-use verification interface. A multi-face/single-sentence strategy is used to combine voice and face verification modules, and support vector machine is employed for information fusion. Experimental results show that our method can achieve high verification accuracies Hsien-Ting Cheng, Yi-Hsiang Chao, Shih-Liang Yeh, Chu-Song Chen, Hsin-Min Wang, Yi-Ping Hung |
ICME | 6 |
| 2005 | Visualization for High-Dimensional Data: VisHDabstractThis paper presents a visualization tool, VisHD, that can visualize the spatial distribution of vector points in high dimensional feature space. It is important to handle high dimensional information in many areas of computer science. VisHD provides several methods for dimension reduction in order to map the data from high dimensional space to low dimensional one. Next, this system builds intuitive visualization for observing the characteristics of the data set, whether these data are pre-defined labels or not. In addition, some useful functions have been implemented to facilitate the information visualization. This paper, finally, gives some experiments and discussions for showing the abilities of VisHD for visualizing high-dimensional data. Cheng-Chih Yang, Cheng-Chieh Chiang, Yi-Ping Hung, Greg C. Lee |
IV | 3 |
| 2005 | A Bayesian approach to video object segmentation via merging 3-D watershed volumes
Yu-Pao Tsai, Chih-Chuan Lai, Yi-Ping Hung, Zen-Chung Shih |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2003 | A real-time robust eye tracking system for autostereoscopic displays using stereo camerasabstractAutostereoscopic display systems can provide users a natural 3D visualization environment by projecting stereo video onto the user's eyes. Eye position localization is a central module in this kind of display system when users are allowed to move freely. This paper presents robust 3D eye tracking techniques that can provide accurate eye positions in real time. Technical batteries comprise: (1) robust face detection based on eigenspace method; (2) real-time face tracking; and (3) eye detection in the obtained face region. According to our implementation on a PC with a Pentium IV 1.2 GHz CPU, the frame rate of the eye tracking process can achieve 25 Hz. Chan-Hung Su, Yong-Sheng Chen, Yi-Ping Hung, Chu-Song Chen, Jiun-Hung Chen |
ICRA | 3 |
| 2002 | A New Approach to Automatic Reconstruction of a 3-D World Using Active Stereo Vision
Sheng-Wen Shih, Yi-Ping Hung, Gregory Y. Tang |
Comput. Vis. Image Underst. | 3 |
| 2002 | Augmenting panoramas with object movies by generating novel views with disparity-based view morphingabstractAbstract Our goal is to augment a panorama with object movies in a visually 3D‐consistent way. Notice that a panorama is recorded as one single 2D image and an object movie (OM) is composed of a set of 2D images taken around a 3D object. The challenge is how to integrate the above two sources of 2D images in a 3D‐consistent way so that the user can easily manipulate object movies in a panorama. To solve this problem, we adopt a purely image‐based approach that does not have to reconstruct the geometric models of the 3D objects to be inserted in the panorama. A critical issue of this method is how to generate the novel views required for showing an OM in different places of a panorama, and we have proposed a view morphing technique, called t‐DBVM, to solve this problem. Our experiments have shown that this purely image‐based approach can effectively generate visually convincing OM‐augmented panoramas. This method has great potential for many applications that require integration of panoramas and object movies, such as virtual malls, virtual museum, and interior design. Copyright © 2002 John Wiley & Sons, Ltd. Yi-Ping Hung, Chu-Song Chen, Yu-Pao Tsai, Szu-Wei Lin |
Comput. Animat. Virtual Worlds | 1 |
| 2001 | Fast Algorithm for Nearest Neighbor Search Based on a Lower Bound Tree
Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICCV | 2 |
| 2001 | Simple and efficient method of calibrating a motorized zoom lens
Yong-Sheng Chen, Sheng-Wen Shih, Yi-Ping Hung, Chiou-Shann Fuh |
Image Vis. Comput. | 3 |
| 2001 | Three-dimensional ego-motion estimation from motion fields observed with multiple cameras
Yong-Sheng Chen, Lin-Gwo Liou, Yi-Ping Hung, Chiou-Shann Fuh |
Pattern Recognit. | 3 |
| 2001 | Fast block matching algorithm based on the winner-update strategyabstractBlock matching is a widely used method for stereo vision, visual tracking, and video compression. Many fast algorithms for block matching have been proposed in the past, but most of them do not guarantee that the match found is the globally optimal match in a search range. This paper presents a new fast algorithm based on the winner-update strategy which utilizes an ascending lower bound list of the matching error to determine the temporary winner. Two lower bound lists derived by using partial distance and by using Minkowski's inequality are described. The basic idea of the winner-update strategy is to avoid, at each search position, the costly computation of the matching error when there exists a lower bound larger than the global minimum matching error. The proposed algorithm can significantly speed up the computation of the block matching because: 1) computational cost of the lower bound we use is less than that of the matching error itself; 2) an element in the ascending lower bound list will be calculated only when its preceding element has already been smaller than the minimum matching error computed so far; 3) for many search positions, only the first several lower bounds in the list need to be calculated. Our experiments have shown that, when applying to motion vector estimation for several widely-used test videos, 92% to 98% of operations can be saved while still guaranteeing the global optimality. Moreover, the proposed algorithm can be easily modified either to meet the limited time requirement or to provide an ordered list of best candidate matches. Our source codes of the proposed algorithm are available at http://smart.iis.sinica.edu.tw/html/winup.html. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
IEEE Trans. Image Process. | 2 |
| 2000 | Winner-Update Algorithm for Nearest Neighbor SearchabstractThis paper presents an algorithm, called the winner-update algorithm, for accelerating the nearest neighbor search. By constructing a hierarchical structure for each feature point in the l/sub p/ metric space, this algorithm can save a large amount of computation at the expense of moderate preprocessing and twice the memory storage. Given a query point, the cost for computing the distances from this point to all the sample points can be reduced by using a lower bound list of the distance established from Minkowski's inequality. Our experiments have shown that the proposed algorithm can save a large amount of computation, especially when the distance between the query point and its nearest neighbor is relatively small. With slight modification, the winner-update algorithm can also speed up the search for k nearest neighbors, neighbors within a specified distance threshold, and neighbors close to the nearest neighbor. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICPR | 2 |
| 2000 | Camera Calibration with a Motorized Zoom LensabstractThis paper presents a simple and efficient method of calibrating the intrinsic camera parameters for all the lens settings of a motorized zoom lens. We fix the aperture setting and perform the camera calibration, adaptively, over the ranges of the room and focus settings. Bilinear interpolation is used to provide the values of the intrinsic camera parameters for those lens settings where no observations are taken. Our experiments show that the proposed method can provide accurate intrinsic camera parameters for all the lens settings, even though camera calibration is performed only for a small number of sampled lens settings. A calibration object suitable for zoom lens calibration is also presented. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh, Sheng-Wen Shih |
ICPR | 2 |
| 2000 | Automatic Detection and Tracking of Human Heads Using an Active Stereo Vision SystemabstractA new head tracking algorithm for automatically detecting and tracking human heads in complex backgrounds is proposed. By using an elliptical model for the human head, our Maximum Likelihood (ML) head detector can reliably locate human heads in images having complex backgrounds and is relatively insensitive to illumination and rotation of the human heads. Our head detector consists of two channels: the horizontal and the vertical channels. Each channel is implemented by multiscale template matching. Using a hierarchical structure in implementing our head detector, the execution time for detecting the human heads in a 512×512 image is about 0.02 second in a Sparc 20 workstation (not including the time for image acquisition). Based on the ellipse-based ML head detector, we have developed a head tracking method that can monitor the entrance of a person, detect and track the person's head, and then control the stereo cameras to focus their gaze on this person's head. In this method, the ML head detector and the mutually-supported constraint are used to extract the corresponding ellipses in a stereo image pair. To implement a practical and reliable face detection and tracking system, further verification using facial features, such as eyes, mouth and nostrils, may be essential. The 3D position computed from the centers of the two corresponding ellipses is then used for fixation. An active stereo head has been used to perform the experiments and has demonstrated that the proposed approach is feasible and promising for practical uses. Cheng-Yuan Tang, Zen Chen, Yi-Ping Hung |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 1999 | New Calibration-free Approach for Augmented Reality based on Parameterized Cuboid StructureabstractA new method called PCS (parameterized cuboid structure) is presented for augmented reality. In particular, our method can insert animated virtual objects into a static scene, with geometric consistency, and also allow the user to interactively position and rotate the virtual objects with respect to a world coordinate system in a physically (or intuitively) meaningful way. Such capability cannot be achieved by using the existing calibration-free methods. To achieve this goal, we develop a new method for estimating camera parameters, which uses a cuboid structure (or more generally, a parallelepiped structure) as the reference object. The reference cuboid structure can be either explicit or implicit-implicit in the sense that the cuboid structure can be inferred by human perception even though it does not appear explicitly in the image. This method can determine the sizes of the cuboid (or parallelepiped) as well as the intrinsic and extrinsic parameters of the camera. To insert a virtual object into a single uncalibrated image, some human interaction is unavoidable. We have implemented an AR authoring system based on the proposed PCS method which provides an auxiliary line and a refinement criterion to assist human interaction. Experimental results have demonstrated that our method can successfully insert virtual objects into both static and dynamic scenes with highly convincing geometric consistency. Chu-Song Chen, Chi-Kuo Yu, Yi-Ping Hung |
ICCV | 3 |
| 1999 | RANSAC-Based DARCES: A New Approach to Fast Automatic Registration of Partially Overlapping Range ImagesabstractIn this paper, we propose a new method, the RANSAC-based DARCES method (data-aligned rigidity-constrained exhaustive search based on random sample consensus), which can solve the partially overlapping 3D registration problem without any initial estimation. For the noiseless case, the basic algorithm of our method can guarantee that the solution it finds is the true one, and its time complexity can be shown to be relatively low. An extra characteristic is that our method can be used even for the case that there are no local features in the 3D data sets. Chu-Song Chen, Yi-Ping Hung, Jen-Bo Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Automatic Detection and Tracking of Human Heads Using an Active Stereo Vision System
Cheng-Yuan Tang, Yi-Ping Hung, Zen Chen |
ACCV (1) | 2 |
| 1998 | A Fast and Robust Approach for Registration of Partially Overlapping Range ImagesabstractA popular approach for 3D registration of partially-overlapping range images is the ICP (iterative closest point) method and many of its variations. The major drawback of this type of iterative approaches is that they require a good initial estimate to guarantee that the correct solution can always be found. In this paper, we propose a new method, the RANSAC-based DARCES (data-aligned rigidity-constrained exhaustive search) method, which can solve the partially-overlapping 3D registration problem efficiently and reliably without any initial estimation. Another important characteristic of our method is that it requires no local features in the 3D data set. An extra characteristic is that, for the noiseless case, the basic algorithm of our DARCES method can guarantee that the solution it finds is the true one, due to its exhaustive-search nature. Even with the nature of exhaustive search, its time complexity can be shown to be relatively low. Experiments have demonstrated that our method is efficient and reliable for registering partially-overlapping range images. Chu-Song Chen, Yi-Ping Hung, Jen-Bo Cheng |
ICCV | 2 |
| 1998 | Free-hand pointer by use of an active stereo vision systemabstractWe developed a system of free-hand pointer by tracking the finger of the speaker with an active stereo vision system and then computing the projection direction in either of the following two modes: the finger-orientation mode and the eye-to-fingertip mode. We prefer the latter for its robustness. In order to allow the speaker to move around in a wider 3D space without reducing the pointing resolution, we utilize a well-calibrated active stereo vision system which has a relatively small field of view but can control its stereo cameras to fixate at the moving finger. Our experiments have successfully demonstrated the feasibility of developing a free-handpointer using an active stereo vision system. Yi-Ping Hung, Yao-Strong Yang, Yong-Sheng Chen, Ing-Bor Hsieh, Chiou-Shann Fuh |
ICPR | 1 |
| 1998 | Toward automatic reconstruction of 3D environment with an active binocular headabstractThis paper presents a new automatic approach to reconstructing a model for the 3D environment by use of an active binocular head (the IIS head). To efficiently store and access the depth estimates, we propose the use of the inverse polar octree. The depth estimates are computed by using the asymptotic Bayesian estimation method. The path of the local motion required by the asymptotic Bayesian method is determined online automatically to reduce the ambiguity of stereo matching. Some rules for checking the consistency between the new observation and the previous observations have been developed. Our experimental results have shown that the proposed approach is very promising for automatic generation of 3D models which can be used for rendering a 3D scene in a virtual reality system. Sheng-Wen Shih, Yi-Ping Hung |
ICPR | 3 |
| 1998 | Integrating virtual objects into real images for augmented realityabstractA systematic approach for integrating virtual objects into real images is developed in this paper.We propose the P3P-IC.method to solve the camera pose estimation problem.A robust tracking method is developed via the combination of the LMedS technique and the P3P-ICP method.With the 3D models and the robust tracking methods, we can determine the camera poses associated with each frame in the image sequence.Knowing the camera poses for each image frame, we can then integrate virtual objects into a video segment. 1.1 Chu-Song Chen, Yi-Ping Hung, Sheng-Wen Shih, Chen-Chiung Hsieh, Chen-Yuan Tang, Chih-Guo Yu, You-Chung Cheng |
VRST | 2 |
| 1998 | Disparity-based view morphing - a new technique for image-based renderingabstractThis paper proposes a new powerful technique for image-based rendering, i.e., the disparity morphing technique. The disparity morphing depends on the estimated disparity map for each image pair based on the epipolar geometry. The epipolar constraints can not only speed up the searching of correspondences for disparity estimation, but also improve the accuracy of the estimated disparity maps. Then the forward and backward morphing functions are defined based on the estimated disparity maps. Notice that the forward and backward morphing functions are usually not linear and not similar to the traditional image morphing functions. Finally, the interpolation function is defined to deal with the object occlusion problems correctly. The resulted transition views are almost the same as what viewers can see in the real world. For object movie systems, the disparity morphing can generate interpolated views between two or more viewing directions. For panoramic imaging systems, the disparity morph... Ho Chao Huang, Shung-Hua Nain, Yi-Ping Hung, Tse Cheng |
VRST | 3 |
| 1998 | Panoramic Stereo Imaging System with Automatic Disparity Warping and Seaming
Ho Chao Huang, Yi-Ping Hung |
Graph. Model. Image Process. | 2 |
| 1998 | Multipass hierarchical stereo matching for generation of digital terrain models from aerial images
Yi-Ping Hung, Chu-Song Chen, Kuan-Chung Hung, Yong-Sheng Chen, Chiou-Shann Fuh |
Mach. Vis. Appl. | 1 |
| 1998 | Calibration of an active binocular headabstractIn this paper, we show how an active binocular head, the IIS head, can be easily calibrated with very high accuracy. Our calibration method can also be applied to many other binocular heads. In addition to the proposal and demonstration of a four-stage calibration process, there are three major contributions in this paper. First, we propose a motorized-focus lens (MFL) camera model which assumes constant nominal extrinsic parameters. The advantage of having constant extrinsic parameters is to having a simple head/eye relation. Second, a calibration method for the MFL camera model is proposed in this paper, which separates estimation of the image center and effective focal length from estimation of the camera orientation and position. This separation has been proved to be crucial; otherwise, estimates of camera parameters would be very noise-sensitive. Thirdly, we show that, once the parameters of the MFL camera model is calibrated, a nonlinear recursive least-square estimator can be used to refine all the 35 kinematic parameters. Real experiments have shown that the proposed method can achieve accuracy of one pixel prediction error and 0.2 pixel epipolar error, even when all the joints, including the left and right focus motors, are moved simultaneously. This accuracy is good enough for many 3D vision applications, such as navigation, object tracking and reconstruction. Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 1998 | New closed-form solution for kinematic parameter identification of a binocular head using point measurementsabstractThis paper proposes a new closed-form solution for identifying the kinematic parameters of an active binocular head having four revolute joints and two prismatic joints by using three-dimensional (3-D) point (position) measurements of a calibration point. Since this binocular head is composed of off-the-shelf components, its kinematic parameters are unknown. Therefore, we can not directly apply those existing nonlinear optimization methods. Even if we want to use the nonlinear optimization methods, a closed-form solution can be first applied to obtain accurate enough initial values. Hence, this paper considers only methods that provide closed-form solutions, i.e., those requiring no initial estimates. Notice that most existing closed-form solutions require pose (i.e., both position and orientation) measurements. However, as far as we know, there is no inexpensive technique which can provide accurate pose measurements. Therefore, existing closed-form solutions based on pose measurements can not give us the required accuracy. As a result, we have developed a new method that does not require orientation measurements and can use only the position measurements of a calibration point to obtain highly accurate estimates of kinematic parameters using closed-form solutions. The proposed method is based on the complete and parametrically continuous (CPC) kinematic model, and can be applied to any kind of kinematic parameter identification problems with or without multiple end-effecters, providing that the links are rigid, the joints are either revolute or prismatic and no closed-loop kinematic chain is included. Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1997 | Ego-Motion Estimation Using Optical Flow Fields Observed from Multiple CamerasabstractIn this paper, we consider a multi-camera vision system mounted on a moving object in a static three-dimensional environment. By using the motion flow fields seen by all of the cameras, an algorithm which does not need to solve the point-correspondence problem among the cameras is proposed to estimate the 3D ego-motion parameters of the moving object. Our experiments have shown that using multiple optical flow fields obtained from different cameras can be very helpful for ego-motion estimation. An-Ting Tsao, Chiou-Shann Fuh, Yi-Ping Hung, Yong-Sheng Chen |
CVPR | 3 |
| 1997 | Adaptive Early Jump-Out Technique for Fast Motion Estimation in Video Coding
Ho Chao Huang, Yi-Ping Hung |
CVGIP Graph. Model. Image Process. | 2 |
| 1997 | Image Registration Using a New Edge-Based Approach
Jun-Wei Hsieh, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Tat Ko, Yi-Ping Hung |
Comput. Vis. Image Underst. | 5 |
| 1997 | Range data acquisition using color structured lighting and stereo vision
Chu-Song Chen, Yi-Ping Hung, Chiann-Chu Chiang, Ja-Ling Wu |
Image Vis. Comput. | 2 |
| 1996 | Model-based object recognition using range images by combining morphological feature extraction and geometric hashingabstractThis paper proposes a new approach for model-based object recognition with range images by combining morphological feature extraction and geometric hashing. In low-level processing, range images are segmented into 3D-connected surface patches. In middle-level processing: each connected component is processed by using morphological operations to extract the skeletons of high-variation regions. These skeleton points can be viewed as invariant salient feature primitives. In high-level processing, geometric hashing is used to recognize objects. To reduce the number of spurious hypotheses, we propose a basis-similarity constraint. Experimental results have shown that the proposed method is effective and has great potential for model-based object recognition using range images. Chu-Song Chen, Yi-Ping Hung, Ja-Ling Wu |
ICPR | 2 |
| 1996 | Adaptive early jump-out technique for fast motion estimation in video codingabstractAn adaptive early jump-out technique for speeding up the block-based motion estimation is proposed. By using the new technique, we can speed up the full range search several times without losing the picture quality significantly. The proposed technique can also be embedded into almost all the existing fast motion estimation algorithms to speed up the computation further. Since the proposed technique can be embedded into the existing motion estimation algorithms, it can be applied to almost all the standard video codecs, such as the MPEG coder, and improve the coding speed of such codecs significantly. Our technique has been tested on the H.261 and the MPEG-I codecs, and the coding speed improves significantly. Ho Chao Huang, Yi-Ping Hung, Wen-Liang Hwang |
ICPR | 2 |
| 1996 | Accuracy analysis on the estimation of camera parameters for active vision systemsabstractIn camera calibration, due to the correlations between certain camera parameters, an estimate of a set of camera parameters which minimizes a given criterion does not guarantee that the estimates of the physical camera parameter are themselves accurate. This problem has not drawn much attention from our computer vision society because conventional computer vision applications require only accurate 3D measurements and do not care much about the values of the physical parameters as long as their composite effect is satisfactory. However, when dealing with an active vision system, accuracy of the physical parameters is very critical because we need that accuracy to establish the relation between the motor positions and the camera parameters (both intrinsic and extrinsic). The contribution of this work is mainly in error analysis of camera calibration, especially in the accuracy of the physical camera parameters themselves, for four different types of calibration problems. With our error analysis, the most suitable camera calibration technique and calibration configuration for providing accurate camera parameters can be determined. Based on this error analysis, we have developed a method for calibrating our eight-degree-of-freedom binocular system. Real experiments have shown that this method can achieve accuracy of one pixel prediction error and 0.2 pixel epipolar error, even when all the eight joints, including the focus motors, are moved simultaneously. This high accuracy greatly alleviates the difficulties encountered in applying the active vision approach to automatic reconstruction of 3D objects and environments, which is our current endeavor. Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
ICPR | 2 |
| 1995 | Kinematic parameter Identification of a Binocular Head Using Stereo Measurements of Single Calibration PointabstractThis paper proposes a new closed-form solution for identifying the kinematic parameters of an active binocular head by using a single calibration point. This method is based on the complete and parametrically continuous (CPC) kinematic model, and can be applied to any kind of kinematic parameter identification problems with or without multiple end-effectors, providing that the links are rigid, the joints are either revolute or prismatic and no closed-loop kinematic chain is included. As a practical example, this paper focuses on the calibration of a binocular head having four revolute joints and two prismatic joints. Simulation and real experiments have shown that the proposed method of using point measurements can achieve much higher accuracy than that of using pose measurements. Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
ICRA | 2 |
| 1995 | When should we consider lens distortion in camera calibration
Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
Pattern Recognit. | 2 |
| 1995 | Comments on "A linear solution to the kinematic parameter identification of robot manipulator" and some modificationsabstractThe paper by Zhuang and Roth (ibid. vol.9, p.174-85 (1993)) presents a linear solution to kinematic parameter identification of robot manipulators. With their method, the orientation parameters for all revolute joints have to be solved first before solving for the translation parameters for all revolute joints and the orientation parameters for all prismatic joints simultaneously. Our major modification here is to decompose the kinematic parameters estimation problem into many subproblems of a single joint axis such that the complexity is reduced and easier implementation is derived. In addition, this gives a unified solution for any serial manipulator with an arbitrary combination of prismatic and revolute joints.> Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
IEEE Trans. Robotics Autom. | 2 |
| 1993 | Comparison between asymptotic Bayesian approach and Kalman filter-based technique for 3D reconstruction using an image sequenceabstractTwo statistical approaches for 3-D reconstruction from an image sequence are compared: the asymptotic Bayesian surface reconstruction and the Kalman filter-based depth estimation. Both techniques are recursive algorithms where relevant information contained in previously taken images is summarized in a prior term (prior to the taking of the next image). This means that the reconstruction results are based upon information from all images but the storage and computation required do not grow dramatically. Experiments with both real images and computer generated images demonstrate that the asymptotic Bayesian approach achieves better results than the Kalman filter-based approach, largely due to better problem formulation.> Chun-Jen Tsai, Yi-Ping Hung, Sheun-Ching Hsu |
CVPR | 2 |
| 1992 | Accuracy assessment on camera calibration method not considering lens distortionabstractThe authors investigate the effect of neglecting lens distortion, and present a theoretical analysis of the calibration accuracy. The error bound derived is a function of a few factors, including the number of calibration points, the observation error of 2-D image points, the radial lens distortion coefficient, and the image size and resolution. This error bound provides a guideline for selecting both a proper camera calibration configuration and an appropriate camera model while providing the desired accuracy. Experimental results from both computer simulations and real experiments are given.> Sheng-Wen Shih, Yi-Ping Hung, Wei-Song Lin |
CVPR | 2 |
| 1991 | Asymptotic bayesian surface estimation using an image sequence
Yi-Ping Hung, David B. Cooper, Bruno Cernuschi-Frías |
Int. J. Comput. Vis. | 1 |
| 1990 | Model-based segmentation and estimation of 3D surfaces from two or more intensity images using Markov random fieldsabstractAn approach and algorithm for 3D primitive model recognition, parameter estimation, and segmentation from a sequence of images taken by one or more calibrated cameras are presented. Though the approach and algorithm are applicable to more general models, the experiments described are for primitive objects that are 3D planes. Given two or more images taken by one or more calibrated cameras, the algorithm simultaneously segments the images and 3D space into regions, each region associated with a single planar patch, and estimates the parameters of the 3D plane associated with each segmented region. The algorithm is suitable for parallel processing and should function at close to the best possible accuracy. Markov random fields are used to provide very coarse prior knowledge of the regions occupied by the planar patches, resulting in markedly enhanced accuracy.> Jayashree Subrahmonia, Yi-Ping Hung, David B. Cooper |
ICPR (1) | 2 |
| 1989 | Toward a Model-Based Bayesian Theory for Estimating and Recognizing Parameterized 3-D Objects Using Two or More Images Taken from Different PositionsabstractA parametric modeling and statistical estimation approach is proposed and simulation data are shown for estimating 3-D object surfaces from images taken by calibrated cameras in two positions. The parameter estimation suggested is gradient descent, though other search strategies are also possible. Processing image data in blocks (windows) is central to the approach. After objects are modeled as patches of spheres, cylinders, planes and general quadrics-primitive objects, the estimation proceeds by searching in parameter space to simultaneously determine and use the appropriate pair of image regions, one from each image, and to use these for estimating a 3-D surface patch. The expression for the joint likelihood of the two images is derived and it is shown that the algorithm is a maximum-likelihood parameter estimator. A concept arising in the maximum likelihood estimation of 3-D surfaces is modeled and estimated. Cramer-Rao lower bounds are derived for the covariance matrices for the errors in estimating the a priori unknown object surface shape parameters.> Bruno Cernuschi-Frías, David B. Cooper, Yi-Ping Hung, Peter N. Belhumeur |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1988 | A New Model-based Stereo Approach For 3D Surface Reconstruction Using Contours On The Surface Pattern
David B. Cooper, Yi-Ping Hung, Gabriel Taubin |
ICCV | 2 |
| 1988 | Bayesian estimation of 3D surfaces from a sequence of imagesabstractAn approach is introduced to estimating object surfaces in 3D space from a sequence of images. A 3D surface of interest is modeled as a function known up to the values of a few parameters. Surface estimation is then treated as the general problem of maximum-likelihood parameter estimation based on two or more functionally related data sets, which constitute a sequence of images taken at different locations and orientations. Experiments are run to illustrate the various advantages of using as many images as possible in the estimation and of distributing camera positions from first to last over as large a baseline as possible. The authors introduce the use of asymptotic Bayesian approximations to summarize the useful information in a sequence of images, thereby drastically reducing both storage and processing. This results in a Bayesian estimator for the surface parameters. All the usual tools of statistical signal analysis can be brought to bear, the information extraction appears to be robust and computationally reasonable, the concepts are geometric and simple, and essentially optimal accuracy should result.> Yi-Ping Hung, David B. Cooper, Bruno Cernuschi-Frías |
ICRA | 1 |