Kuan-Wen Chen

dblp:59/968 · DBLP profile ↗
← Back
41ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 2 since 2021Systems, architecture and hardware · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A deep feature distillation framework based on multilayer stacked decomposition and aggregation for high-volatility sequence prediction
Chi-Jie Lu, Ting-Jen Chang, Kuan-Wen Chen
Neurocomputing3
2026 Fully test-time rPPG estimation via synthetic signal-guided feature learning
Pei-Kai Huang, Tzu-Hsien Chen, Ya-Ting Chan, Kuan-Wen Chen, Shih-Yu Yang, Yen-Chun Chou, Chiou-Ting Hsu
Pattern Recognit.4
2026 Ultra-short rPPG estimation via periodicity guidance and signal reconstruction
Pei-Kai Huang, Ya-Ting Chan, Kuan-Wen Chen, Chiou-Ting Hsu, Xiaoding Wang, Mohammad Jalil Piran
Pattern Recognit.3
2026 Lifelong Learner: Discovering Versatile Neural Solvers for Vehicle Routing Problems
abstract
Deep learning has been extensively explored to solve vehicle routing problems (VRPs), which yields a range of data-driven neural solvers with promising outcomes. However, most neural solvers are trained to tackle VRP instances in a relatively monotonous context, e.g., simplifying VRPs by using Euclidean distance between nodes and adhering to a single problem size, which harms their off-the-shelf application in different scenarios. To enhance their versatility, this paper presents a novel lifelong learning framework that incrementally trains a neural solver to manage VRPs in distinct contexts. Specifically, we propose a lifelong learner (LL), exploiting a Transformer network as the backbone, to solve a series of VRPs. The inter-context self-attention mechanism is proposed within LL to transfer the knowledge obtained from solving preceding VRPs into the succeeding ones. On top of that, we develop a dynamic context scheduler (DCS), employing the cross-context experience replay to further facilitate LL looking back on the attained policies of solving preceding VRPs. Extensive results on synthetic and benchmark instances (problem sizes up to 18k) show that our LL is capable of discovering effective policies for tackling generic VRPs in varying contexts, which outperforms other neural solvers and achieves the best performance for most VRPs.
Shaodi Feng, Zhuoyi Lin, Jianan Zhou 0002, Kuan-Wen Chen, J. Senthilnath 0001, Yew-Soon Ong
IEEE Trans. Intell. Transp. Syst.6
2025 TC-NeRF:Temporal Consistent Neural Radiance Fields with Cross-View Complementation for Occluded Object Removal
abstract
In this paper, a novel object removal method for Neural Radiance Fields (NeRF) is proposed to address inconsistencies encountered by current object removal methods. Existing methods typically inpaint masked images independently before integrating them into NeRF, often leading to cross-view inconsistencies. Our approach focuses on inpainting only regions that are entirely unseen or visible in very few views, thereby improving temporal consistency. We leverage the cross-view complementation property of NeRF to determine visible and fully occluded regions. By projecting masks across different views and counting pixel visibility, we identify regions that require inpainting. This selective inpainting strategy ensures coherent reconstructions and higher rendering quality.
Zicheng Wu, Li-Hsuan Chang, Kuan-Wen Chen
ICME3
2025 Efficient and Precise Drone Rephotography for Video Sequences
abstract
Precise drone rephotography technology aims to recover camera poses from a reference sequence and obtain well-aligned image sequences, playing a crucial role in autonomous drone inspection tasks. However, existing rephotography methods rely on static image inputs, resulting in low efficiency and limited applicability in real-world scenarios. This paper presents a novel video-based precise drone rephotography system leveraging video sequences. To the best of our knowledge, this is the first work to extend precise drone rephotography from still images to videos while significantly reducing rephotography time. The proposed approach integrates advanced visual SLAM techniques with a dense flow prediction model to continuously refine the drone’s pose, enabling robust and precise rephotography tasks. To further quantify system performance, we introduce a trajectory-based visual similarity evaluation standard—Dynamic Frame Alignment Error (DFAE), which assesses the visual similarity of drone-captured videos of varying durations. We conducted multiple experiments with drones in real-world scenarios. Experimental results demonstrate that the proposed system achieves efficient and precise rephotography across multiple indoor and outdoor trials. Specifically, the average rephotography error is only 7.956 pixels indoors and 9.800 pixels outdoors. More importantly, the rephotography time is only half of the baseline.
Hao-Liang Xu, Chu-Chun Chi, Kuan-Wen Chen
IROS3
2025 DD-rPPGNet: De-Interfering and Descriptive Feature Learning for Unsupervised rPPG Estimation
Pei-Kai Huang, Tzu-Hsien Chen, Ya-Ting Chan, Kuan-Wen Chen, Chiou-Ting Hsu
IEEE Trans. Inf. Forensics Secur.4
2024 LCCRAFT: LiDAR and Camera Calibration Using Recurrent All-Pairs Field Transforms Without Precise Initial Guess
abstract
LiDAR-camera fusion plays a pivotal role in 3D reconstruction for self-driving applications. A fundamental prerequisite for effective fusion is the precise calibration between LiDAR and camera systems. Many existing calibration methods are constrained by predefined mis-calibration ranges in the training data, essentially tying the network to a specific data distribution. However, if the range of evaluation data differs from what the network has been trained on, the resulting estimates may not meet expectations. Moreover, most methods require a precise initial guess for calibration to succeed. In this paper, we introduce LCCRAFT, an online calibration network designed for LiDAR and camera systems. Leveraging the 4D correlation volume and correlation lookup techniques inherited from RAFT, we apply them to correlate RGB images and depth maps derived from the projection of point clouds. Through weight sharing between update iterations and by enabling the update operator to learn from data with varying degrees of error, LCCRAFT demonstrates adaptability to diverse miscalibration scenarios. This includes cases where the initial mis- calibration is even more severe than what the system encountered during training, demonstrating the robustness of the model. The calibration process executes in 93ms on a single GPU, meeting real-time requirements. Despite the modest 9M model parameters, LCCRAFT achieves competitive performance as compared to the state-of-the-art method, which entails 69M parameters.
Yu-Chen Lee, Kuan-Wen Chen
ICRA2
2024 Photometric Consistency for Precise Drone Rephotography
abstract
This paper proposes a precise drone rephotography system for fixed-domain patrolling scenarios. The proposed system integrates computer-vision-based localization and fine-tuned pixel-level dense flow prediction to achieve consistent and precise rephotography images with viewpoints that closely align with those of target images. The proposed Keypoints Alignment Through Dense Flow Prediction (KADFP) model effectively handles challenges arising from lighting variations and background differences. Moreover, a novel flight procedure is implemented in the proposed system. This procedure involves using an Interleaved Drone Controller to alternate between translation and rotation adjustments to ensure smooth flight dynamics during rephotography. Experiments indicated that the proposed system provided considerably more precise rephotography results (error of 4.72 pixels indoors) than did an existing localization approach (error of 35.56 pixels).
Hsaun-Jui Chang, Tzu-Chun Huang, Hao-Liang Xu, Kuan-Wen Chen
IROS4
2024 CollabLoc: Collaborative Information Sharing for Real-Time Multiuser Visual Localization System
abstract
This paper presents CollabLoc, a novel approach for real-time multi-user visual localization. Typically, localization systems employ a client-server design for locating cameras. In these systems, lightweight simultaneous localization and mapping computations are performed on the client side, while the server handles intensive localization tasks. This approach harnesses the complementary capabilities of the client and server, resulting in accurate, real-time localization results. However, existing architectures primarily operate on a one-to-one client-server structure, limiting their scalability and multi-user capabilities. Therefore, CollabLoc is designed to accommodate multiple clients through collaborative information sharing to considerably reduce computational overhead and enhance overall efficiency and accuracy. We propose a tracking confidence module that evaluates the tracking quality of individual clients and plays a pivotal role in prioritizing client requests by the server-side algorithm. On the server, we utilize fused poses to accelerate image retrieval. Moreover, we enhance the efficiency of optical flow estimation by employing a simplified feature extraction module and leveraging spatial similarities among neighboring clients to improve its performance. Finally, via the Pose Fusion Module, the server can periodically adjust fused poses to mitigate accumulated errors. Experimental results indicate that compared with a baseline method, CollabLoc improves positioning efficiency by nearly twice and achieves higher accuracy in multi-user scenarios.
Teng-Te Yu, Yo-Chung Lau, Kai-Li Wang, Kuan-Wen Chen
IROS4
2024 Neural Architecture Search Using Covariance Matrix Adaptation Evolution Strategy
abstract
Evolution-based neural architecture search methods have shown promising results, but they require high computational resources because these methods involve training each candidate architecture from scratch and then evaluating its fitness, which results in long search time. Covariance Matrix Adaptation Evolution Strategy (CMA-ES) has shown promising results in tuning hyperparameters of neural networks but has not been used for neural architecture search. In this work, we propose a framework called CMANAS which applies the faster convergence property of CMA-ES to the deep neural architecture search problem. Instead of training each individual architecture seperately, we used the accuracy of a trained one shot model (OSM) on the validation data as a prediction of the fitness of the architecture, resulting in reduced search time. We also used an architecture-fitness table (AF table) for keeping a record of the already evaluated architecture, thus further reducing the search time. The architectures are modeled using a normal distribution, which is updated using CMA-ES based on the fitness of the sampled population. Experimentally, CMANAS achieves better results than previous evolution-based methods while reducing the search time significantly. The effectiveness of CMANAS is shown on two different search spaces using four datasets: CIFAR-10, CIFAR-100, ImageNet, and ImageNet16-120. All the results show that CMANAS is a viable alternative to previous evolution-based methods and extends the application of CMA-ES to the deep neural architecture search field.
Nilotpal Sinha, Kuan-Wen Chen
Evol. Comput.2
2023 LGCNet: Feature Enhancement and Consistency Learning Based on Local and Global Coherence Network for Correspondence Selection
abstract
Correspondence selection, a crucial step in many computer vision tasks, aims to distinguish between inliers and outliers from putative correspondences. The coherence of correspondences is often used for predicting inlier probability, but it is difficult for neural networks to extract coherence contexts based only on quadruple coordinates. To overcome this difficulty, we propose enhancing the preliminary features using local and global handcrafted coherent characteristics before model learning, which strengthens the discrimination of each correspondence and guides the model to prune obvious outliers. Furthermore, to fully utilize local information, neighbors are searched in coordinate space as well as feature space. These two kinds of neighbors provide complementary and plentiful contexts for inlier probability prediction. Finally, a novel neighbor representation and a fusion architecture are proposed to retain detailed features. Experiments demonstrate that our method achieves state-of-the-art performance on relative camera pose estimation and correspondence selection metrics on the outdoor YFCC100M [1] and the indoor SUN3D [2] datasets.
Tzu-Han Wu, Kuan-Wen Chen
ICRA2
2023 Object-Level Unknown Obstacle Detection
abstract
This paper presents a novel method for object-level unknown obstacle detection in driving scenes that reduces false positives. The proposed method combines existing anomaly detectors, depth estimation, and object detection techniques to achieve object-level predictions. Our method can predict anomalies as bound-box instance detections. These bounding boxes can then be used to refine anomaly detection by suppressing false positives outside of the bounding boxes. The proposed method has several advantages, including object-level detections that are more practical than pixel-level detections, and the ability to find and refine region proposals for obstacle detection. The paper provides a detailed explanation of all components of the system and includes an ablation study on the usage of depth estimation, as well as execution time averages on different hardware. The proposed method is evaluated using different metrics and benchmarks, demonstrating the effectiveness and relevance of the existing proposed methods. Overall, our proposed method has the potential to significantly improve object-level anomaly detection making it suitable for real-world applications.
Chuan-Yuan Huang, Cheng-Tsung Chen, Yu-An Chen, Kuan-Wen Chen
IROS4
2023 MUFeat: Multi-Level CNN and Unsupervised Learning for Local Feature Detection and Description
abstract
Local feature detection and description are two essential steps in many visual applications. Most learned local feature methods require high-quality labeled data to achieve superior performance, but such labels are often expensive. To address this problem, we propose MUFeat, an unsupervised learning framework of jointly learning local feature detector and descriptor without requirement of ground-truth correspondences. MUFeat trains the network based on the putative matches from the pretrained model and two proposed unsupervised loss functions. Furthermore, the MUFeat framework includes a pyramidal feature hierarchy network to obtain keypoints and descriptors from feature maps. Experiments indicate that MUFeat outperforms most state-of-the-art supervised learning methods on image matching, medical image registration and visual localization tasks.
Sheng-Hung Kuo, Tzu-Han Wu, Zheng-Yan Chen, Kuan-Wen Chen
IROS4
2023 Enhance Local Feature Consistency with Structure Similarity Loss for 3D Semantic Segmentation
abstract
Recently, many research studies have been carried out on using deep learning methods for 3D point cloud understanding. However, there is still no remarkable result on 3D point cloud semantic segmentation compared to those of 2D research. One important reason is that 3D data has higher dimensionality but lacks large datasets, which means that the deep learning model is difficult to optimize and easy to overfit. To overcome this, an essential method is to provide more priors to the learning of deep models. In this paper, we focus on semantic segmentation for point clouds in the real world. To provide priors to the model, we propose a novel loss function called Linearity and Planarity to enhance local feature consistency in the regions with similar local structure. Experiments show that the proposed method improves baseline performance on both indoor and outdoor datasets e.g. S3DIS and Semantic3D.
Cheng-Wei Lin, Fang-Yu Syu, Yi-Ju Pan, Kuan-Wen Chen
IROS4
2022 Predicting Opportune Moments to Deliver Notifications in Virtual Reality
abstract
Virtual reality (VR) has increasingly been used in many areas, and the need to deliver notifications in VR is also expected to increase accordingly. However, untimely interruptions could largely impact the experience in VR. Identifying opportune times to deliver notifications to users allows for notifications to be scheduled in a way that minimizes disruption. We conducted a study to investigate the use of sensor data available on an off-the-shelf VR device and additional contextual information, including current activity and engagement of users, to predict opportune moments for sending notifications using deep learning models. Our analysis shows that using mainly sensor features could achieve 72% recall, 71% precision and 0.86 area under receiver operating characteristic (AUROC); performance can be further improved to 81% recall, 82% precision, and 0.93 AUROC if information about activity and summarized user engagement is included.
Kuan-Wen Chen, Yung-Ju Chang, Li-Wei Chan 0001
CHI1
2022 Neural architecture search using progressive evolution
abstract
Vanilla neural architecture search using evolutionary algorithms (EA) involves evaluating each architecture by training it from scratch, which is extremely time-consuming. This can be reduced by using a supernet to estimate the fitness of every architecture in the search space due to its weight sharing nature. However, the estimated fitness is very noisy due to the co-adaptation of the operations in the supernet. In this work, we propose a method called pEvoNAS wherein the whole neural architecture search space is progressively reduced to smaller search space regions with good architectures. This is achieved by using a trained supernet for architecture evaluation during the architecture search using genetic algorithm to find search space regions with good architectures. Upon reaching the final reduced search space, the supernet is then used to search for the best architecture in that search space using evolution. The search is also enhanced by using weight inheritance wherein the supernet for the smaller search space inherits its weights from previous trained supernet for the bigger search space. Experimentally, pEvoNAS gives better results on CIFAR-10 and CIFAR-100 while using significantly less computational resources as compared to previous EA-based methods. The code for our paper can be found here.
Nilotpal Sinha, Kuan-Wen Chen
GECCO2
2022 Temporally-Aggregating Multiple-Discontinuous-Image Saliency Prediction with Transformer-Based Attention
abstract
In this paper, we aim to apply deep saliency prediction to automatic drone exploration, which should consider not only one single image, but multiple images from different view angles or localizations in order to determine the exploration direction. However, little attention has been paid to such saliency prediction problem over multiple-discontinuous-image and none of existing methods take temporal information into consideration, which may mean that the current predicted saliency map is not consistent with the previous predicted results. For this purpose, we propose a method named Temporally-Aggregating Multiple-Discontinuous-Image Saliency Prediction Network (TA-MSNet). It utilizes a transformer-based attention module to correlate relative saliency information among multiple discontinuous images and, furthermore, applies the ConvLSTM module to capture the temporal information. Experiments show that the proposed TA-MSNet can estimate better and more consistent results than previous works for time series data.
Pin-Jie Huang, Chi-An Lu, Kuan-Wen Chen
ICRA3
2022 DroneTalk: An Internet-of-Things-Based Drone System for Last-Mile Drone Delivery
abstract
At present, drone delivery systems usually function only in global positioning system (GPS)-friendly environments and thus cannot deliver goods into customers’ houses, which is an element of urban last-mile delivery. In this study, we investigated solutions for enabling drones to fly autonomously in mixed indoor–outdoor environments. We propose a novel Internet of Things (IoT)-based drone delivery system, namely DroneTalk, for mail delivery. DroneTalk combines a GPS, an inertial measurement unit, and visual information to achieve mixed indoor–outdoor autopilot operation. Furthermore, DroneTalk integrates an autonomous drone control system with an IoT device management platform to enable a drone to achieve automatic online weather awareness. Most importantly, DroneTalk is a novel work considering communication, processing, and control latencies simultaneously during autopilot operation to avoid problems related to drone collisions. These problems have always existed but become more serious when drones are flying close to buildings. Simulation results indicate that the proposed system can attain a high user-defined flight success rate (i.e., 99% in our simulations) without any collisions; thus, the proposed system can be feasibly used in real-world environments.
Kuan-Wen Chen, Ming-Ru Xie, Yu-Min Chen, Ting-Tsan Chu, Yi-Bing Lin
IEEE Trans. Intell. Transp. Syst.1
2021 Evolving neural architecture using one shot model
abstract
Previous evolution based architecture search require high computational resources resulting in large search time. In this work, we propose a novel way of applying a simple genetic algorithm to the neural architecture search problem called EvNAS (Evolving Neural Architecture using One Shot Model) which reduces the search time significantly while still achieving better result than previous evolution based methods. The architectures are represented by architecture parameter of one shot model which results in the weight sharing among the given population of architectures and also weight inheritance from one generation to the next generation of architectures. We use the accuracy of partially trained architecture on validation data as a prediction of its fitness to reduce the search time. We also propose a decoding technique for the architecture parameter which is used to divert majority of the gradient information towards the given architecture and is also used for improving the fitness prediction of the given architecture from the one shot model during the search process. EvNAS searches for architecture on CIFAR-10 for 3.83 GPU day on a single GPU with top-1 test error 2.47%, which is then transferred to CIFAR-100 and ImageNet achieving top-1 error 16.37% and top-5 error 7.4% respectively.
Nilotpal Sinha, Kuan-Wen Chen
GECCO2
2021 Collaborative Learning of Multiple-Discontinuous-Image Saliency Prediction for Drone Exploration
abstract
Most of the existing saliency prediction research focuses on either single images or videos (or more precisely multiple images in sequence). However, to apply saliency prediction to drone exploration that has to consider multiple images from different view angles or localizations to determine the direction to explore, saliency prediction over multiple discontinuous images is required. In this paper, we propose a deep relative saliency model (MS-Net) for such an application. MS-Net starts with a single-image saliency feature extraction network for each image separately and then integrate these images by using a GCN-based mechanism called multi-image saliency fusion that learns relative saliency information among all the images. Finally, it predicts the saliency of each image by considering the relative information. Because there are no existing saliency prediction datasets with such multiple discontinuous images, we randomly cropped a large number of sub-images from 360° images of the existing 360° image saliency datasets to build our own dataset for both training and evaluation. Experimental results showed that the proposed MSNet considerably outperformed both single-image and video saliency prediction methods and could achieve comparative performance to that of 360° image saliency prediction even with only limited field-of-views, i.e., five sub-images, considered.
Ting-Tsan Chu, Po-Heng Chen, Pin-Jie Huang, Kuan-Wen Chen
ICRA4
2021 Finding Robust 2D-to-3D Correspondence with LSTM Score Estimation for Camera Localization
abstract
2D-to-3D correspondence estimation is the key step of 3D model-based image localization, and most of the existing research in this field focuses on improving the feature matching performance. Even with the best feature matching method, there are still some outliers, and thus, almost all the methods simply apply the RANSAC algorithm to select the inliers and estimate the camera pose afterwards. However, the reliability of RANSAC depends considerably on the inlier ratio. Once the inlier ratio decreases, for example a challenging scenario occurs, it will be unable to select the inliers well and lead to a worse camera pose. In this study, we attempted to build a neural network to learn the geometric relationship between 2D images and the 3D model to select the correct correspondence from the initial 2D-to-3D matching results to improve the performance of camera localization. Because the number of inputs, i.e., the number of 2D-to-3D correspondences, is unknown and different for each image, we propose a PointNet-based Geometric Consistency Network (GCC-Net) for the correct correspondence estimation and an LSTM-based Hypothesis Rating Network (HR-Net) to enhance GCC-Net with the camera localization loss. Experimental results showed that the proposed method outperforms RANSAC considerably on the camera pose estimation, particularly when the inlier ratio of the initial correspondence was low.
Tsu-Kuan Huang, Po-Heng Chen, Li-Yang Wang, Kuan-Wen Chen
IROS4
2021 V-Eye: A Vision-Based Navigation System for the Visually Impaired
abstract
Numerous systems for helping visually impaired people navigate in unfamiliar places have been proposed. However, few can detect and warn about moving obstacles, provide correct orientation in real time, or support navigation between indoor and outdoor spaces. Accordingly, this paper proposes V-Eye, which fulfills these needs by utilizing a novel global localization method (VB-GPS) and image-segmentation techniques to achieve better scene understanding with a single camera. Our experiments establish that the proposed system can reliably provide precise locations and orientation information (with a median error of approximately 0.27 m and 0.95°); detect unpredictable obstacles; and support navigating both within and between indoor and outdoor environments. The results of a user-experience study of V-eye further indicate that it helped the participants not only with navigation, but also improved their awareness of obstacles, enhanced their spatial awareness more generally, and led them to feel more secure and independent while walking.
Ping-Jung Duh, Yu-Cheng Sung, Liang-Yu Fan Chiang, Yung-Ju Chang, Kuan-Wen Chen
IEEE Trans. Multim.5
2020 PA-FlowNet: Pose-Auxiliary Optical Flow Network for Spacecraft Relative Pose Estimation
abstract
During the process of space travelling and space landing, the spacecraft attitude estimation is the indispensable work for navigation. Since there are not enough satellites for GPS-like localization in space, the computer vision technique is adopted to address the issue. The most crucial task for localization is the extraction of correspondences. In computer vision, optical flow estimation is often used for finding correspondences between images. As the deep neural network being more popular in recent years, FlowNet2 has played a vital role which achieves great success. In this paper, we present PA-FlowNet, an end-to-end pose-auxiliary optical flow network which can use the predicted relative camera pose to improve the performance of optical flow. PA-FlowNet is composed of two sub-networks, the foreground-attention flow network and the pose regression network. The foreground-attention flow network is constructed by FlowNet2 model and modified with the proposed foreground-attention approach. We introduced this approach with the concept of curriculum learning for foreground-background segmentation to avoid backgrounds from resulting in flow prediction error. The pose regression network is used to regress the relative camera pose as an auxiliary for increasing the accuracy of the flow estimation. In addition, to simulate the test environment for spacecraft pose estimation, we construct a 64K moon model and simulate aerial photography with various attitudes to generate Moon64K dataset in this paper. PA-FlowNet significantly outperforms all existing methods on the proposed Moon64K dataset. Furthermore, we also predict the relative camera pose via proposed PA-FlowNet and accomplish the remarkable performance.
Zhi-Yu Chen, Po-Heng Chen, Kuan-Wen Chen, Chen-Yu Chan
ICPR3
2020 IF-Net: An Illumination-invariant Feature Network
abstract
Feature descriptor matching is a critical step is many computer vision applications such as image stitching, image retrieval and visual localization. However, it is often affected by many practical factors which will degrade its performance. Among these factors, illumination variations are the most influential one, and especially no previous descriptor learning works focus on dealing with this problem. In this paper, we propose IF-Net, aimed to generate a robust and generic descriptor under crucial illumination changes conditions. We find out not only the kind of training data important but also the order it is presented. To this end, we investigate several dataset scheduling methods and propose a separation training scheme to improve the matching accuracy. Further, we propose a ROI loss and hard-positive mining strategy along with the training scheme, which can strengthen the ability of generated descriptor dealing with large illumination change conditions. We evaluate our approach on public patch matching benchmark and achieve the best results compared with several state-of-the-arts methods. To show the practicality, we further evaluate IF-Net on the task of visual localization under large illumination changes scenes, and achieves the best localization accuracy.
Po-Heng Chen, Zhao-Xu Luo, Zu-Kuan Huang, Kuan-Wen Chen
ICRA5
2020 MVSNet++: Learning Depth-Based Attention Pyramid Features for Multi-View Stereo
abstract
The goal of Multi-View Stereo (MVS) is to reconstruct 3D point-cloud model from multiple views. On the basis of the considerable progress of deep learning, an increasing amount of research has moved from traditional MVS methods to learning-based ones. However, two issues remain unsolved in the existing state-of-the-art methods: (1) only high-level information is considered for depth estimation. This may reduce the localization accuracy of 3D points as the learned model lacks spatial information; and (2) most of the methods require additional post-processing or network refinement to generate a smooth 3D model. This significantly increases the number of model parameters or the computational complexity. To this end, we propose MVSNet++, an end-to-end trainable network for dense depth estimation. Such an estimated depth map can further be applied to 3D model reconstruction. Different from previous methods, in the proposed method, we first adopt feature pyramid structures for both feature extraction and cost volume regularization. This can lead to accurate 3D point localization by fusing multi-level information. To generate smooth depth map, we then carefully integrate instance normalization into MVSNet++ without increasing model parameters and computational burden. Furthermore, we additionally design three loss functions and integrate Curriculum Learning framework into the training process, which can lead to an accurate reconstruction of 3D model. MVSNet++ is evaluated on DTU and Tanks & Temples benchmarks with comprehensive ablation studies. Experimental results demonstrate that our proposed method performs favorably against previous state-of-the-art methods, showing the accuracy and effectiveness of the proposed MVSNet++.
Po-Heng Chen, Hsiao-Chien Yang, Kuan-Wen Chen, Yong-Sheng Chen
IEEE Trans. Image Process.3
2020 FADE: Feature Aggregation for Depth Estimation With Multi-View Stereo
abstract
Both structural and contextual information is essential and widely used in image analysis. However, current multi-view stereo (MVS) approaches usually use a single common pre-trained model as pixel descriptor to extract features, which mix structural and contextual information together and thus increase the difficulty of matching correspondence. In this paper, we propose FADE (feature aggregation for depth estimation), which treats spatial and context information separately and focuses on aggregating features for efficient learning of the MVS problem. Spatial information includes image details such as edges and corners, whereas context information comprises object features such as shapes and traits. To aggregate these multi-level features, we use an attention mechanism to select important features for matching. We then build a plane sweep volume by using a homography backward warping method to generate match candidates. Furthermore, we propose a novel cost volume regularization network aims to minimize the noise in the matching candidates. Finally, we take advantage of 3D stacked hourglass and regression to produces high-quality depth maps. With these well-aggregated features, FADE can efficiently perform dense depth reconstruction, achieving state-of-the-art performance in terms of accuracy and requiring the least amount of model parameters.
Hsiao-Chien Yang, Po-Heng Chen, Kuan-Wen Chen, Chen-Yi Lee, Yong-Sheng Chen
IEEE Trans. Image Process.3
2017 A novel egocentric pointing system based on smart glasses
abstract
In this paper, we propose a novel, egocentric pointing system based on Google Glass which is equipped with an optical head-mounted display (OHMD) and a near-eye camera, with the eye-pointing line passing through the lower left corner of the display. For a pointed target, the pointing (or ranging) algorithm is based on a distance-pixel curve established from the camera-eye (epipolar) geometry. Additional pointing algorithms for estimating gazing point on a planar surface are also developed by establishing another distance-pixel curve along the same epipolar line. Experiments show that less than 0.32° angular error in the egocentric pointing can be achieved for object distance ranging from 80cm to 178cm by the best estimation scheme, with slightly less accurate results (i.e. 0.58°) achievable by simpler estimation schemes.
Yi-Yu Hsieh, Yu-Han Wei, Kuan-Wen Chen, Jen-Hui Chuang
VCIP3
2017 Vision-Based Positioning for Internet-of-Vehicles
abstract
This paper presents an algorithm for ego-positioning by using a low-cost monocular camera for systems based on the Internet-of-Vehicles. To reduce the computational and memory requirements, as well as the communication load, we tackle the model compression task as a weighted k-cover problem for better preserving the critical structures. For real-world vision-based positioning applications, we consider the issue of large scene changes and introduce a model update algorithm to address this problem. A large positioning data set containing data collected for more than a month, 106 sessions, and 14275 images is constructed. Extensive experimental results show that submeter accuracy can be achieved by the proposed ego-positioning algorithm, which outperforms existing vision-based approaches.
Kuan-Wen Chen, Chun-Hsin Wang, Qiao Liang 0002, Chu-Song Chen, Ming-Hsuan Yang 0001, Yi-Ping Hung
IEEE Trans. Intell. Transp. Syst.1
2014 A sleep monitoring system based on audio, video and depth information for detecting sleep events
abstract
The purpose of this study is to develop a non-invasive sleep monitoring system to distinguish sleep disturbances based on multiple sensors. Unlike clinical sleep monitoring which records biological information such as EEG, EOG, and EMG, in this study, we aim to identify occurrences of events from a sleep environment. A device with an infrared depth sensor, a RGB camera, and a four-microphone array is used to detect three types of events: motion events, lighting events, and sound events. Given streams of depth signals and color images, we build two background models to detect movements and lighting effects, and audio signals are scored simultaneously. Moreover, we classify events by an epoch approach algorithm and provide a graphical sleep diagram for browsing corresponding video clips. Experimental results in sleep condition show the efficiency and reliability of our system, and it is convenient and cost-effective to be used in home context.
Lyn Chao-ling Chen, Kuan-Wen Chen, Yi-Ping Hung
ICME2
2014 Appearance-Based Gaze Tracking with Free Head Movement
abstract
In this work, we develop an appearance-based gaze tracking system allowing user to move their head freely. The main difficulty of the appearance-based gaze tracking method is that the eye appearance is sensitive to head orientation. To overcome the difficulty, we propose a 3-D gaze tracking method combining head pose tracking and appearance-based gaze estimation. We use a random forest approach to model the neighbor structure of the joint head pose and eye appearance space, and efficiently select neighbors from the collected high dimensional data set. L1-optimization is then used to seek for the best solution for regression from the selected neighboring samples. Experiment results shows that it can provide robust binocular gaze tracking results with less constraints but still provides moderate estimation accuracy of gaze estimation.
Chih-Chuan Lai, Kuan-Wen Chen, Shen-Chi Chen, Sheng-Wen Shih, Yi-Ping Hung
ICPR3
2014 Large-Area, Multilayered, and High-Resolution Visual Monitoring Using a Dual-Camera System
abstract
Large-area, high-resolution visual monitoring systems are indispensable in surveillance applications. To construct such systems, high-quality image capture and display devices are required. Whereas high-quality displays have rapidly developed, as exemplified by the announcement of the 85-inch 4K ultrahigh-definition TV by Samsung at the 2013 Consumer Electronics Show (CES), high-resolution surveillance cameras have progressed slowly and remain not widely used compared with displays. In this study, we designed an innovative framework, using a dual-camera system comprising a wide-angle fixed camera and a high-resolution pan-tilt-zoom (PTZ) camera to construct a large-area, multilayered, and high-resolution visual monitoring system that features multiresolution monitoring of moving objects. First, we developed a novel calibration approach to estimate the relationship between the two cameras and calibrate the PTZ camera. The PTZ camera was calibrated based on the consistent property of distinct pan-tilt angle at various zooming factors, accelerating the calibration process without affecting accuracy; this calibration process has not been reported previously. After calibrating the dual-camera system, we used the PTZ camera and synthesized a large-area and high-resolution background image. When foreground targets were detected in the images captured by the wide-angle camera, the PTZ camera was controlled to continuously track the user-selected target. Last, we integrated preconstructed high-resolution background and low-resolution foreground images captured using the wide-angle camera and the high-resolution foreground image captured using the PTZ camera to generate a large-area, multilayered, and high-resolution view of the scene.
Chih-Wei Lin 0004, Kuan-Wen Chen, Shen-Chi Chen, Cheng-Wu Chen, Yi-Ping Hung
ACM Trans. Multim. Comput. Commun. Appl.2
2013 Target-driven video summarization in a camera network
abstract
Nowadays, ever expanding camera network makes it difficult to find the suspect from lengthy video records. This paper proposes a target-driven video summarization framework which provides two-step Filtered Summarized Video (FSV) for tracing suspects. Before the target is identified, users can find the target efficiently using the firststep FSV of any arbitrary camera. The first-step FSV filters all the attributes of the target including the time information and the target's categories. After identifying the target, the second-step FSV with additional spatio-temporal & appearance cues are triggered in the neighbor cameras. To enhance the accuracy of the object classification for FSV, we propose a Perspective Dependent Model (PDM) which consists of many grid-based models. Finally, the experimental results show that grid-based model is more robust than general detectors and the user study demonstrates better performance for target finding and tracking in camera network for surveillance.
Shen-Chi Chen, Shih-Yao Lin 0001, Kuan-Wen Chen, Chih-Wei Lin 0004, Chu-Song Chen, Yi-Ping Hung
ICIP4
2013 Real-time camera tampering detection using two-stage scene matching
abstract
We propose a tampering detection method using two-stage scene matching for real application with high efficiency and low false alarm rate. In the first stage, we use the intensity of edges as the main cue to detect the camera tampering events. Instead of using the entire edge points of the images, we sample the most significant edge points to represent the scene. Analyzing the edge variation with only the sample points, we discover that the events of camera tampering can be detected with low computation cost. Whenever the first stage detects the tampering event, the second stage is triggered to reduce false alarms. In the second stage, we propose an illumination change detector which can check the consistency of the scene structure using cell-based matching method. The experimental results demonstrate that our system can detect the camera tampering precisely and minimize false alarm even when the illumination changes dramatically or large crowds passing through the scene.
Chao-Ching Shih, Shen-Chi Chen, Cheng-Feng Hung, Kuan-Wen Chen, Shih-Yao Lin 0001, Chih-Wei Lin 0004, Yi-Ping Hung
ICME4
2011 Egocentric View Transition for Video Monitoring in a Distributed Camera Network
Kuan-Wen Chen, Pei-Jyun Lee, Yi-Ping Hung
MMM (1)1
2011 Multi-Resolution Design for Large-Scale and High-Resolution Monitoring
abstract
Large-scale and high-resolution monitoring systems are ideal for many visual surveillance applications. However, existing approaches have insufficient resolution and low frame rate per second, or have high complexity and cost. We take inspiration from the human visual system and propose a multi-resolution design, e-Fovea, which provides peripheral vision with a steerable fovea that is in higher resolution. In this paper, we firstly present two user studies, with a total of 36 participants, to compare e-Fovea to two existing multi-resolution visual monitoring designs. The user study results show that for visual monitoring tasks, our e-Fovea design with steerable focus is significantly faster than existing approaches and preferred by users. We then present our design and implementation of e-Fovea, which combines both multi-resolution camera input and multi-resolution steerable projector output. Finally, we present our deployment of e-Fovea in three installations to demonstrate its feasibility.
Kuan-Wen Chen, Chih-Wei Lin 0004, Tzu-Hsuan Chiu, Mike Yen-Yang Chen, Yi-Ping Hung
IEEE Trans. Multim.1
2011 Adaptive Learning for Target Tracking and True Linking Discovering Across Multiple Non-Overlapping Cameras
abstract
To track targets across networked cameras with disjoint views, one of the major problems is to learn the spatio-temporal relationship and the appearance relationship, where the appearance relationship is usually modeled as a brightness transfer function. Traditional methods learning the relationships by using either hand-labeled correspondence or batch-learning procedure are applicable when the environment remains unchanged. However, in many situations such as lighting changes, the environment varies seriously and hence traditional methods fail to work. In this paper, we propose an unsupervised method which learns adaptively and can be applied to long-term monitoring. Furthermore, we propose a method that can avoid weak links and discover the true valid links among the entry/exit zones of cameras from the correspondence. Experimental results demonstrate that our method outperforms existing methods in learning both the spatio-temporal and the appearance relationship, and can achieve high tracking accuracy in both indoor and outdoor environment.
Kuan-Wen Chen, Chih-Chuan Lai, Pei-Jyun Lee, Chu-Song Chen, Yi-Ping Hung
IEEE Trans. Multim.1
2010 Multi-Cue Integration for Multi-Camera Tracking
abstract
For target tracking across multiple cameras with disjoint views, previous works usually employed multiple cues and focused on learning a better matching model of each cue, separately. However, none of them had discussed how to integrate these cues to improve performance, to our best knowledge. In this paper, we look into the multi-cue integration problem and propose an unsupervised learning method since a complicated training phase is not always viable. In the experiments, we evaluate several types of score fusion methods and show that our approach learns well and can be applied to large camera networks more easily.
Kuan-Wen Chen, Yi-Ping Hung
ICPR1
2010 e-Fovea: a multi-resolution approach with steerable focus to large-scale and high-resolution monitoring
abstract
This paper presents e-Fovea, a system that combines both multi-resolution camera input and multi-resolution steerable projector output to support large-scale and high-resolution visual monitoring. e-Fovea utilizes a design similar to the human eyes, which provides peripheral vision with a steerable fovea that is in higher resolution. e-Fovea is implemented using a steerable telephoto camera and a wide-angle camera. The telephoto image is displayed using a projector with a steerable mirror, and overlaid on the wide-angle image that is displayed using a second projector.
Kuan-Wen Chen, Chih-Wei Lin 0004, Mike Y. Chen, Yi-Ping Hung
ACM Multimedia1
2009 Generating pictorial-based representation of mental images for video monitoring
abstract
Multi-camera systems have been widely used in many video surveillance applications. When an event happens and is monitored across multiple cameras, it is easy for an expert to generate the corresponding spatial representation to comprehend the series of event. However, it is not trivial for users new to the environment. With support from psychological evidences, we propose an approach to mimic generating pictorial-based representation of mental images when a target is moving across the views of cameras. First we conduct a ball-rolling experiment to compare this approach with others. The empirical results demonstrate that the performance of users with this approach is significantly better than others. We suggest that it is because this approach is better for users to preserve spatial representation of the environment while transiting views between cameras. Then we propose a framework to realize this approach. The demonstrations in different situations indicate the validity of such framework.
Chuan-Heng Hsiao, Wei-Chia Huang, Kuan-Wen Chen, Li-Wei Chang, Yi-Ping Hung
IUI3
2008 An adaptive learning method for target tracking across multiple cameras
abstract
This paper proposes an adaptive learning method for tracking targets across multiple cameras with disjoint views. Two visual cues are usually employed for tracking targets across cameras: spatio-temporal cue and appearance cue. To learn the relationships among cameras, traditional methods used batch-learning procedures or hand-labeled correspondence, which can work well only within a short period of time. In this paper, we propose an unsupervised method which learns both spatio-temporal relationships and appearance relationships adaptively and can be applied to long-term monitoring. Our method performs target tracking across multiple cameras while also considering the environment changes, such as sudden lighting changes. Also, we improve the estimation of spatio-temporal relationships by using the prior knowledge of camera network topology.
Kuan-Wen Chen, Chih-Chuan Lai, Yi-Ping Hung, Chu-Song Chen
CVPR1