EDBT 2026 Demo / reviewers in the wild / expert
Yong-Sheng Chen
dblp:32/2144
· DBLP profile ↗
47ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-5581-850XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | US3Net: Ultralightweight Self-Supervised Stereo Matching Network using Depth-Aware Geometric Soft OcclusionabstractAbstract An ultralightweight self-supervised stereo matching network, called US $$^3$$ 3 Net, which requires only 12K parameters, is designed for efficient and accurate depth estimation using resource-constrained devices. US $$^3$$ 3 Net incorporates two key innovations: a low-complexity feature extraction module and a soft occlusion detection approach for performance improvement. First, we design a low-complexity feature extraction module to reduce the computational burden while preserving structural details necessary for stereo matching. By refining the encoder backbone and aggregation module, our design ensures a better balance between model complexity and accuracy. Second, to address occlusion-related errors in disparity estimation, we propose a novel occlusion detection method, called Depth-Aware Geometric Soft Occlusion (DAGSO), to adaptively define the occlusion confidence scores based on depth information. DAGSO can effectively mitigate false occlusions in distant regions and can improve the accuracy of disparity estimation. Experimental results using KITTI datasets demonstrate that US $$^3$$ 3 Net achieves state-of-the-art performance in terms of model complexity and depth estimation accuracy. It outperforms previous self-supervised stereo matching methods and monocular depth estimation methods in metrics such as AbsRel, SqRel, RMSE, and RMSElog at a reduction of parameter size by 47% compared with ES $$^3$$ 3 Net (23K). This makes US $$^3$$ 3 Net a practical solution for real-time depth estimation on edge devices such as drones and autonomous systems. Code is available at: https://github.com/g830319ag/US3Net. Po-Chung Jen, Tzu-Chi Liu, I-Sheng Fang, Hsiao-Chieh Wen, Chia-Lun Hsu, Ping-Yang Chen, Chang-Hsing Lee, Yong-Sheng Chen |
Int. J. Comput. Vis. | 8 |
| 2025 | Stands on Shoulders of Giants: Learning to Lift 2D Detection to 3D with Geometry-Driven Objectivesabstract3D detection of vehicles is an essential component for autonomous driving applications. Nevertheless, collecting the supervised training data for learning 3D vehicle detectors would be costly (e.g. utilization of expensive LiDAR sensors) and labor-intensive (for human annotation). In comparison to 3D detection, 2D object detection has achieved a welldeveloped status, boosting stable and robust performance with widespread application in numerous fields, thanks to the large scale (i.e. amount of samples) of existing training datasets of 2D object detection. Hence, in our work, we propose to realize 3D detection via leveraging the robustness of 2D detectors and developing a network that lifts 2D detections to 3D. With the flexibility of building upon various backbone models (e.g. the models which take image regions detected by 2D detector as inputs to predict their corresponding 3D bounding boxes, or the existing monocular 3D detection models which have the intermediate output of 2 D bounding boxes), we propose several geometry-driven objectives, including projection consistency loss, geometry depth loss, and opposite bin loss, to improve the training upon 2D-to-3D lifting. Our extensive experimental results demonstrate that our proposed geometrydriven objectives not only contribute to the superior results of 3D detection but also provide better generalizability across datasets. Jhih-Rong Chen, Che-Yuan Chang, Szu-Han Tseng, Chih-Sheng Huang, Yong-Sheng Chen, Walon Wei-Chen Chiu |
ICRA | 5 |
| 2025 | Weighted Stratification in Multi-label Contrastive Learning for Long-Tailed Medical Image Classification
Ying-Chih Lin, Yong-Sheng Chen |
MICCAI (13) | 2 |
| 2024 | Best of Both Sides: Integration of Absolute and Relative Depth Sensing Modalities Based on iToF and RGB Cameras
I-Sheng Fang, Walon Wei-Chen Chiu, Yong-Sheng Chen |
ICPR (16) | 3 |
| 2024 | IIOF: Intra- and Inter-feature orthogonal fusion of local and global features for music emotion recognition
Pei-Chun Chang, Yong-Sheng Chen, Chang-Hsing Lee |
Pattern Recognit. | 2 |
| 2023 | Fisheye Multiple Object Tracking by Learning Distortions Without DewarpingabstractWe develop a new Multiple Object Tracking (MOT) scheme for fisheye cameras that can directly perform vehicle detection, re-identification, and tracking under fisheye distortions without explicit dewarping. Fisheye cameras provide omnidirectional coverage that is wider than traditional cameras, reducing fewer need of cameras to monitor road intersections. However, the problem of distorted views introduces new challenges for fisheye MOT. In this paper, we propose a Fish-Eye Multiple Object Tracking (FEMOT) approach with two novelties. We develop the Distorted Fisheye Image Augmentation (DFIA) method to improve object detection and re-identification on fisheye cameras, where fisheye model training can be performed on existing datasets of traditional cameras via fisheye data synthesis and augmentation. We also develop the Hybrid Data Association (HDA) method to perform tracking directly on fisheye views, without the need of de-warping. The developed FEMOT framework provides practical design and advancement that enables large-scale use of fisheye cameras in smart city and surveillance applications. Ping-Yang Chen, Jun-Wei Hsieh, Ming-Ching Chang, Munkhjargal Gochoo, Fang-Pang Lin, Yong-Sheng Chen |
ICIP | 6 |
| 2023 | Distilling knowledge for occlusion robust monocular 3D face reconstruction
Hitika Tiwari, Vinod K. Kurmi, K. S. Venkatesh, Yong-Sheng Chen |
Image Vis. Comput. | 4 |
| 2023 | Real-time self-supervised achromatic face colorization
Hitika Tiwari, K. S. Venkatesh, Yong-Sheng Chen |
Vis. Comput. | 3 |
| 2022 | Self-Supervised Robustifying Guidance for Monocular 3D Face Reconstruction
Hitika Tiwari, Min-Hung Chen, Yi-Min Tsai, Hsien-Kai Kuo, Hung-Jen Chen 0004, Kevin Jou, K. S. Venkatesh, Yong-Sheng Chen |
BMVC | 8 |
| 2022 | Facial Image Reconstruction from Functional Magnetic Resonance Imaging via GAN Inversion with Improved Attribute ConsistencyabstractNeuroscience studies have revealed that the brain encodes visual content and embeds information in neural activity. Recently, deep learning techniques have facilitated attempts to address visual reconstructions by mapping brain activity to image stimuli using generative adversarial networks (GANs). However, none of these studies have considered the semantic meaning of latent code in image space. Omitting semantic information could potentially limit the performance. In this study, we propose a new framework to reconstruct facial images from functional Magnetic Resonance Imaging (fMRI) data. With this framework, the GAN inversion is first applied to train an image encoder to extract latent codes in image space, which are then bridged to fMRI data using linear transformation. Following the attributes identified from fMRI data using an attribute classifier, the direction in which to manipulate attributes is decided and the attribute manipulator adjusts the latent code to improve the consistency between the seen image and the reconstructed image. Our experimental results suggest that the proposed framework accomplishes two goals: (1) reconstructing clear facial images from fMRI data and (2) maintaining the consistency of semantic characteristics. Pei-Chun Chang, Yan-Yu Tien, Chia-Lin Chen, Li-Fen Chen, Yong-Sheng Chen, Hui-Ling Chan |
IJCNN | 5 |
| 2022 | Occlusion Resistant Network for 3D Face Reconstructionabstract3D face reconstruction from a monocular face image is a mathematically ill-posed problem. Recently, we observed a surge of interest in deep learning-based approaches to address the issue. These methods possess extreme sensitivity towards occlusions. Thus, in this paper, we present a novel context-learning-based distillation approach to tackle the occlusions in the face images. Our training pipeline focuses on distilling the knowledge from a pre-trained occlusion-sensitive deep network. The proposed model learns the context of the target occluded face image. Hence our approach uses a weak model (unsuitable for occluded face images) to train a highly robust network towards partially and fully-occluded face images. We obtain a landmark accuracy of 0.77 against 5.84 of recent state-of-the-art-method for real-life challenging facial occlusions. Also, we propose a novel end-to-end training pipeline to reconstruct 3D faces from multiple variations of the target image per identity to emphasize the significance of visible facial features during learning. For this purpose, we leverage a novel composite multi-occlusion loss function. Our multi-occlusion per identity model shows a dip in the landmark error by a large margin of 6.67 in comparison to a recent state-of-the-art method. We deploy the occluded variations of the CelebA validation dataset and AFLW2000-3D face dataset: naturally-occluded and artificially occluded, for the comparisons. We comprehensively compare our results with the other approaches concerning the accuracy of the reconstructed 3D face mesh for occluded face images. Hitika Tiwari, Vinod K. Kurmi, K. S. Venkatesh, Yong-Sheng Chen |
WACV | 4 |
| 2022 | Mixed Stage Partial Network and Background Data Augmentation for Surveillance Object DetectionabstractState-of-the-art (SoTA) object detection models and their accuracy have been improved by a large margin via CNNs (Convolutional Neural Networks); however, these models still perform poorly for small road objects. Moreover, the SoTA models are mainly trained on public benchmark datasets such as MS COCO, which include more complicated backgrounds and thus make them robust for object detection. However, for surveillance or road videos, their monotone backgrounds make these SoTA detectors background-over-fitted. In applications such as autonomous driving or traffic flow estimation, the background-over-fitting problem will increase various challenges and lead to accuracy degradation in object detection. One novelty of this paper is to propose an MBA (Mixed Background Augmentation) method to improve detection accuracy without adding new labeling efforts and any pre-training processes. During the inference stage, only one input image is needed for vehicle detection without involving background subtraction. Another novelty of this paper is the design of an efficient MSP (Mixed Stage Partial) network to detect objects more accurately and efficiently from surveillance videos. Extensive experiments on KITTI and UA-DETRAC benchmarks show that the proposed method achieves the SoTA results for highly accurate and efficient vehicle detection. The detection accuracy is improved from 78.53% to 83.59% with 25.7$fps$on the UA-DETRAC data set. The implementation code is available athttps://github.com/pingyang1117/MSPNet. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Decoding Neural Representations of Rhythmic Sounds From MagnetoencephalographyabstractNeuroscience studies have revealed neural processes involving rhythm perception, suggesting that brain encodes rhythmic sounds and embeds information in neural activity. In this work, we investigate how to extract rhythmic information embedded in the brain responses and to decode the original audio waveforms from the extracted information. A spatiotemporal convolutional neural network is adopted to extract compact rhythm-related representations from the noninvasively measured magnetoencephalographic (MEG) signals evoked by listening to rhythmic sounds. These learned MEG representations are then used to condition an audio generator network for the synthesis of the original rhythmic sounds. In the experiments, we evaluated the proposed method by using the MEG signals recorded from eight participants and demonstrated that the generated rhythms are highly related to those evoking the MEG signals. Interestingly, we found that the auditory-related MEG channels reveal high importance in encoding rhythmic representations, the distribution of these representations relate to the timing of beats, and the behavior performance is consistent with the performance of neural decoding. These results suggest that the proposed method can synthesize rhythms by decoding neural representations from MEG. Pei-Chun Chang, Jia-Ren Chang, Li-Kai Cheng, Jen-Chuen Hsieh, Hsin-Yen Yu, Li-Fen Chen, Yong-Sheng Chen |
ICASSP | 8 |
| 2021 | Learning Facial Representations from the Cycle-consistency of FaceabstractFaces manifest large variations in many aspects, such as identity, expression, pose, and face styling. Therefore, it is a great challenge to disentangle and extract these characteristics from facial images, especially in an unsupervised manner. In this work, we introduce cycle-consistency in facial characteristics as free supervisory signal to learn facial representations from unlabeled facial images. The learning is realized by superimposing the facial motion cycle-consistency and identity cycle-consistency constraints. The main idea of the facial motion cycle-consistency is that, given a face with expression, we can perform de-expression to a neutral face via the removal of facial motion and further perform re-expression to reconstruct back to the original face. The main idea of the identity cycle-consistency is to exploit both de-identity into mean face by depriving the given neutral face of its identity via feature re-normalization and re-identity into neutral face by adding the personal attributes to the mean face. At training time, our model learns to disentangle two distinct facial representations to be useful for performing cycle-consistent face reconstruction. At test time, we use the linear protocol scheme for evaluating facial representations on various tasks, including facial expression recognition and head pose regression. We also can directly apply the learnt facial representations to person recognition, frontalization and image-to-image translation. Our experiments show that the results of our approach is competitive with those of existing methods, demonstrating the rich and unique information embedded in the dis-entangled representations. Code is available at https://github.com/JiaRenChang/FaceCycle. Jia-Ren Chang, Yong-Sheng Chen, Walon Wei-Chen Chiu |
ICCV | 2 |
| 2021 | Light-Weight Mixed Stage Partial Network for Surveillance Object Detection with Background Data AugmentationabstractState-of-the-art (SoTA) models have improved object detection accuracy with a large margin via convolutional neural networks, however still with an inferior performance for small objects. Moreover, these models are trained mainly based on the COCO dataset, and its backgrounds are more complicated than road environments, and thus degrade the accuracy of small road object detection. Compared with the COCO dataset, the background of a surveillance video is relatively stable and can be used to enhance the accuracy of road object detection. This paper designs a computationally efficient mixed stage partial (MSP) network to detect road objects. Another novelty of this paper is to propose a mixed background data augmentation method to enhance the detection accuracy without adding new labelling efforts. During inference, only the input image is used to detect road objects without further using any subtraction information. Extensive experiments on KITTI and UA-DETRAC benchmarks show the proposed method achieves the SoTA results for highly-accurate and efficient road object detection. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
ICIP | 4 |
| 2021 | Demystifying T1-MRI to FDG18-PET Image Translation via Representational Similarity
Chia-Hsiang Kao, Yong-Sheng Chen, Li-Fen Chen, Walon Wei-Chen Chiu |
MICCAI (3) | 2 |
| 2021 | MS-SincResNet: Joint Learning of 1D and 2D Kernels Using Multi-scale SincNet and ResNet for Music Genre ClassificationabstractIn this study, we proposed a new end-to-end convolutional neural network, called MS-SincResNet, for music genre classification. MS-SincResNet appends 1D multi-scale SincNet (MS-SincNet) to 2D ResNet as the first convolutional layer in an attempt to jointly learn 1D kernels and 2D kernels during the training stage. First, an input music signal is divided into a number of fixed-duration (3 seconds in this study) music clips, and the raw waveform of each music clip is fed into 1D MS-SincNet filter learning module to obtain three-channel 2D representations. The learned representations carry rich timbral, harmonic, and percussive characteristics comparing with spectrograms, harmonic spectrograms, percussive spectrograms and Mel-spectrograms. ResNet is then used to extract discriminative embeddings from these 2D representations. The spatial pyramid pooling (SPP) module is further used to enhance the feature discriminability, in terms of both time and frequency aspects, to obtain the classification label of each music clip. Finally, the voting strategy is applied to summarize the classification results from all 3-second music clips. In our experimental results, we demonstrate that the proposed MS-SincResNet outperforms the baseline SincNet and many well-known hand-crafted features. Considering individual 2D representation, MS-SincResNet also yields competitive results with the state-of-the-art methods on the GTZAN dataset and the ISMIR2004 dataset. The code is available at https://github.com/PeiChunChang/MS-SincResNet. Pei-Chun Chang, Yong-Sheng Chen, Chang-Hsing Lee |
ICMR | 2 |
| 2021 | Exploiting Spatial Relation for Reducing Distortion in Style TransferabstractThe power of convolutional neural networks in arbitrary style transfer has been amply demonstrated; however, existing stylization methods tend to generate spatially inconsistent results with noticeable artifacts. One solution to this problem involves the application of a segmentation mask or affinity-based image matting to preserve spatial information related to image content. The main idea of this work is to model spatial relation between content image pixels and thus to maintain this relationship in stylization for reducing artifacts. The proposed network architecture is called spatial relation-augmented VGG (SRVGG), in which long-range spatial dependency is modeled by a spatial relation module. Based on this spatial information extracted from SRVGG, we design a novel relation loss which can minimize the difference of spatial dependency between content images and stylizations. We evaluate the proposed framework on both optimization-based and feedforward-based style transfer methods. The effectiveness of SRVGG in stylization is demonstrated by generating stylized images of high quality and spatial consistency without the need for segmentation masks or affinity-based image matting. The quantitative evaluation also suggests that the proposed framework achieve better performance compared with other methods. Jia-Ren Chang, Yong-Sheng Chen |
WACV | 2 |
| 2021 | Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object DetectionabstractThis paper proposes the Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN) for fast and accurate single-shot object detection. Feature Pyramid (FP) is widely used in recent visual detection, however the top-down pathway of FP cannot preserve accurate localization due to pooling shifting. The advantage of FP is weakened as deeper backbones with more layers are used. In addition, it cannot keep up accurate detection of both small and large objects at the same time. To address these issues, we propose a new parallel FP structure with bi-directional (top-down and bottom-up) fusion and associated improvements to retain high-quality features for accurate localization. We provide the following design improvements: (1) A parallel bifusion FP structure with a bottom-up fusion module (BFM) to detect both small and large objects at once with high accuracy. (2) A concatenation and re-organization (CORE) module provides a bottom-up pathway for feature fusion, which leads to the bi-directional fusion FP that can recover lost information from lower-layer feature maps. (3) The CORE feature is further purified to retain richer contextual information. Such CORE purification in both top-down and bottom-up pathways can be finished in only a few iterations. (4) The adding of a residual design to CORE leads to a new Re-CORE module that enables easy training and integration with a wide range of deeper or lighter backbones. The proposed network achieves state-of-the-art performance on the UAVDT17 and MS COCO datasets. Code is available at https://github.com/pingyang1117/PRBNet_PyTorch. Ping-Yang Chen, Ming-Ching Chang, Jun-Wei Hsieh, Yong-Sheng Chen |
IEEE Trans. Image Process. | 4 |
| 2020 | Attention-Aware Feature Aggregation for Real-Time Stereo Matching on Edge Devices
Jia-Ren Chang, Pei-Chun Chang, Yong-Sheng Chen |
ACCV (1) | 3 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 6 |
| 2020 | Deep Real-time Hand Detectoin Using CFPN on Embedded SystemsabstractReal-time HI (Human Interface) systems need accurate and efficient hand detection models to meet the limited resources in budget, dimension, memory, computing, and electric power. In recent years, object detection became a less challenging task with the latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide the desired efficiency and accuracy for HI systems on embedded devices due to their complex time-consuming architecture. In addition, the detection of small hands ( pixels) is still a challenging task for all the above existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) to provide above mentioned performance for small hand detection. The superiority of CFPN is confirmed on a HandFlow dataset with mAP:0.5 of 95.6 and FPS of 33 on Nvidia TX2. The COCO dataset is also used to compare with other state-of-the-art method and shows the highest efficiency and accuracy with the proposed CFPN model. Thus we conclude that the proposed model is useful for real-life small hand detection on embedded devices. Pirdiansyah Hendri, Jun-Wei Hsieh, Ping-Yang Chen, Munkhjargal Gochoo, Yong-Sheng Chen |
ICPR | 5 |
| 2020 | MVSNet++: Learning Depth-Based Attention Pyramid Features for Multi-View StereoabstractThe goal of Multi-View Stereo (MVS) is to reconstruct 3D point-cloud model from multiple views. On the basis of the considerable progress of deep learning, an increasing amount of research has moved from traditional MVS methods to learning-based ones. However, two issues remain unsolved in the existing state-of-the-art methods: (1) only high-level information is considered for depth estimation. This may reduce the localization accuracy of 3D points as the learned model lacks spatial information; and (2) most of the methods require additional post-processing or network refinement to generate a smooth 3D model. This significantly increases the number of model parameters or the computational complexity. To this end, we propose MVSNet++, an end-to-end trainable network for dense depth estimation. Such an estimated depth map can further be applied to 3D model reconstruction. Different from previous methods, in the proposed method, we first adopt feature pyramid structures for both feature extraction and cost volume regularization. This can lead to accurate 3D point localization by fusing multi-level information. To generate smooth depth map, we then carefully integrate instance normalization into MVSNet++ without increasing model parameters and computational burden. Furthermore, we additionally design three loss functions and integrate Curriculum Learning framework into the training process, which can lead to an accurate reconstruction of 3D model. MVSNet++ is evaluated on DTU and Tanks & Temples benchmarks with comprehensive ablation studies. Experimental results demonstrate that our proposed method performs favorably against previous state-of-the-art methods, showing the accuracy and effectiveness of the proposed MVSNet++. Po-Heng Chen, Hsiao-Chien Yang, Kuan-Wen Chen, Yong-Sheng Chen |
IEEE Trans. Image Process. | 4 |
| 2020 | FADE: Feature Aggregation for Depth Estimation With Multi-View StereoabstractBoth structural and contextual information is essential and widely used in image analysis. However, current multi-view stereo (MVS) approaches usually use a single common pre-trained model as pixel descriptor to extract features, which mix structural and contextual information together and thus increase the difficulty of matching correspondence. In this paper, we propose FADE (feature aggregation for depth estimation), which treats spatial and context information separately and focuses on aggregating features for efficient learning of the MVS problem. Spatial information includes image details such as edges and corners, whereas context information comprises object features such as shapes and traits. To aggregate these multi-level features, we use an attention mechanism to select important features for matching. We then build a plane sweep volume by using a homography backward warping method to generate match candidates. Furthermore, we propose a novel cost volume regularization network aims to minimize the noise in the matching candidates. Finally, we take advantage of 3D stacked hourglass and regression to produces high-quality depth maps. With these well-aggregated features, FADE can efficiently perform dense depth reconstruction, achieving state-of-the-art performance in terms of accuracy and requiring the least amount of model parameters. Hsiao-Chien Yang, Po-Heng Chen, Kuan-Wen Chen, Chen-Yi Lee, Yong-Sheng Chen |
IEEE Trans. Image Process. | 5 |
| 2019 | Optimal Charging Scheduling for Electric Vehicle in Parking Lot with Renewable Energy SystemabstractAs the popularity of electric vehicles grows, it is an important issue to effectively schedule the charging plan of electric vehicles in order to supply a large number of electric vehicles while charging and reducing the impact on the grid. The charging station managers obtain the time of use pricing provided by the power company and minimize the charging cost. Owing to electric vehicle charging schedule belongs to combinatorial optimization problems, the popular intelligent control algorithm may have been generated too more infeasible solutions. In this paper, the elitism simulated annealing is proposed for the electric vehicle charging schedule, and the storage batteries and photovoltaic system are included in the case study, so that the charging demand of each electric vehicle can be satisfied and the minimum electricity cost can be attained. Chao-Rong Chen, Yong-Sheng Chen, Tzu-Chiao Lin |
SMC | 2 |
| 2018 | Pyramid Stereo Matching NetworkabstractRecent work has shown that depth estimation from a stereo pair of images can be formulated as a supervised learning task to be resolved with convolutional neural networks (CNNs). However, current architectures rely on patch-based Siamese networks, lacking the means to exploit context information for finding correspondence in ill-posed regions. To tackle this problem, we propose PSMNet, a pyramid stereo matching network consisting of two main modules: spatial pyramid pooling and 3D CNN. The spatial pyramid pooling module takes advantage of the capacity of global context information by aggregating context in different scales and locations to form a cost volume. The 3D CNN learns to regularize cost volume using stacked multiple hourglass networks in conjunction with intermediate supervision. The proposed approach was evaluated on several benchmark datasets. Our method ranked first in the KITTI 2012 and 2015 leaderboards before March 18, 2018. The codes of PSMNet are available at: https://github.com/JiaRenChang/PSMNet. Jia-Ren Chang, Yong-Sheng Chen |
CVPR | 2 |
| 2017 | Deep Competitive Pathway NetworksabstractIn the design of deep neural architectures, recent studies have demonstrated the benefits of grouping subnetworks into a larger network. For examples, the Inception architecture integrates multi-scale subnetworks and the residual network can be regarded that a residual unit combines a residual subnetwork with an identity shortcut. In this work, we embrace this observation and propose the Competitive Pathway Network (CoPaNet). The CoPaNet comprises a stack of competitive pathway units and each unit contains multiple parallel residual-type subnetworks followed by a max operation for feature competition. This mechanism enhances the model capability by learning a variety of features in subnetworks. The proposed strategy explicitly shows that the features propagate through pathways in various routing patterns, which is referred to as pathway encoding of category information. Moreover, the cross-block shortcut can be added to the CoPaNet to encourage feature reuse. We evaluated the proposed CoPaNet on four object recognition benchmarks: CIFAR-10, CIFAR-100, SVHN, and ImageNet. CoPaNet obtained the state-of-the-art or comparable results using similar amounts of parameters. The code of CoPaNet is available at: \urlhttps://github.com/JiaRenChang/CoPaNet. Jia-Ren Chang, Yong-Sheng Chen |
ACML | 2 |
| 2015 | Efficient calibration for multi-plane homography using a laser levelabstractAn efficient calibration method for multi-plane homography is proposed in this paper. Two laser levels are used to cast laser lines to construct virtual poles in the environment without deploying real objects. HSV color model, Hough transform, and least squares method are applied to locate the laser lines in the captured images. Using the features of vanishing line and the view-invariant cross-ratio model, the 3-D coordinate of the camera can be estimated. The multi-plane homography between camera image and the world space can be derived based on two layers homography. The first layer homography relates the ground and the image, whereas the second layer homography can be efficiently obtained using the first layer homography and the virtual poles. Experimental results show that the virtual poles and the derived homography are both accurately estimated. Yen-Chou Tai, Chin-Wei Liu, Yong-Sheng Chen, Jen-Hui Chuang |
ICIP | 3 |
| 2011 | Person Identification Using Electroencephalographic Signals Evoked by Visual Stimuli
Jia-Ping Lin, Yong-Sheng Chen, Li-Fen Chen |
ICONIP (1) | 2 |
| 2011 | Human Object Inpainting Using Manifold Learning-Based Posture Sequence EstimationabstractWe propose a human object inpainting scheme that divides the process into three steps: 1) human posture synthesis; 2) graphical model construction; and 3) posture sequence estimation. Human posture synthesis is used to enrich the number of postures in the database, after which all the postures are used to build a graphical model that can estimate the motion tendency of an object. We also introduce two constraints to confine the motion continuity property. The first constraint limits the maximum search distance if a trajectory in the graphical model is discontinuous, and the second confines the search direction in order to maintain the tendency of an object's motion. We perform both forward and backward predictions to derive local optimal solutions. Then, to compute an overall best solution, we apply the Markov random field model and take the potential trajectory with the maximum total probability as the final result. The proposed posture sequence estimation model can help identify a set of suitable postures from the posture database to restore damaged/missing postures. It can also make a reconstructed motion sequence look continuous. Chih-Hung Ling, Yu-Ming Liang, Chia-Wen Lin, Yong-Sheng Chen, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 4 |
| 2011 | Virtual Contour Guided Video Object Inpainting Using Posture Mapping and RetrievalabstractThis paper presents a novel framework for object completion in a video. To complete an occluded object, our method first samples a 3-D volume of the video into directional spatio-temporal slices, and performs patch-based image inpainting to complete the partially damaged object trajectories in the 2-D slices. The completed slices are then combined to obtain a sequence of virtual contours of the damaged object. Next, a posture sequence retrieval technique is applied to the virtual contours to retrieve the most similar sequence of object postures in the available non-occluded postures. Key-posture selection and indexing are used to reduce the complexity of posture sequence retrieval. We also propose a synthetic posture generation scheme that enriches the collection of postures so as to reduce the effect of insufficient postures. Our experiment results demonstrate that the proposed method can maintain the spatial consistency and temporal motion continuity of an object simultaneously. Chih-Hung Ling, Chia-Wen Lin, Chih-Wen Su, Yong-Sheng Chen, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 4 |
| 2010 | Video object inpainting using manifold-based action predictionabstractThis paper presents a novel scheme for object completion in a video. The framework includes three steps: posture synthesis, graphical model construction, and action prediction. In the very beginning, a posture synthesis method is adopted to enrich the number of postures. Then, all postures are used to build a graphical model of object action which can provide possible motion tendency. We define two constraints to confine the motion continuity property. With the two constraints, possible candidates between every two consecutive postures are significantly reduced. Finally, we apply the Markov Random Field model to perform global matching. The proposed approach can effectively maintain the temporal continuity of the reconstructed motion. The advantage of this action prediction strategy is that it can handle the cases such as non-periodic motion or complete occlusion. Chih-Hung Ling, Yu-Ming Liang, Chia-Wen Lin, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 4 |
| 2009 | Beamformer-Based Spatiotemporal Imaging of Correlated Brain ActivitiesabstractThe past findings have suggested that temporal correlation may relate the communications between the distributed areas. There are some studies in Magnetoencephalography and electroencephalography to analyze the functional connectivity between cortical areas with the oscillations feature of neuronal activity. However, it is also important to observe the functional connectivity through temporal correlation between cortical areas. We proposed a beamformer-based approach which exploits a maximum correlation criterion to maximize the significance level of correlation between brain activities. This criterion leads to a closed-form solution of the dipole orientation. Experiments with simulation data clearly demonstrate the effectiveness, necessity, and accuracy of the proposed method. I-Tzu Chen, Yong-Sheng Chen, Li-Fen Chen |
BIBE | 2 |
| 2009 | Automated Sleep Staging Using Single EEG Channel for REM Sleep DeprivationabstractIn medical literatures, it has been reported that the increased REM (rapid eye movement) density is one of the characters of depressed sleep. Some experiments were conducted to confirm that REM sleep deprivation (REM-SD) for a period of time is therapeutic for endogenous depressed patients. However, because of its high complexity and intensive labor requirement, this therapy has not yet been proved validity by a sufficient amount of depressed patients. Therefore, we propose to develop an automated sleep staging system using only single EEG channel to achieve on-line detection for REM state during sleep. For classifier design, we use a dataset of 25 subjects and the staging accuracy can achieve 80%. Once the REM state is detected by the system, the system will alarm the subject to deprive the REM sleep. The effect of REM sleep deprivation can be examined by hypnogram and the proposed system will be applied for clinical trials of depression therapy. Yu-Hsun Lee, Yong-Sheng Chen, Li-Fen Chen |
BIBE | 2 |
| 2009 | Video object inpainting using posture mappingabstractThis paper presents a novel framework for object-based video inpainting. To complete an occluded object, our method first samples a 3-D volume of the video into directional spatio-temporal slices, and then performs patch-based image inpainting to repair the partially damaged object trajectories in the 2-D slices. The completed slices are subsequently combined to obtain a sequence of virtual contours of the damaged object. The virtual contours and a posture sequence retrieval technique are then used to retrieve the most similar sequence of object postures in the available non-occluded postures. Key-posture selection and indexing are performed to reduce the complexity of posture sequence retrieval. We also propose a synthetic posture generation scheme that enriches the collection of key-postures so as to reduce the effect of insufficient key-postures. Our experimental results demonstrate that the proposed method can maintain the spatial consistency and temporal motion continuity of an object simultaneously. Chih-Hung Ling, Chia-Wen Lin, Chih-Wen Su, Hong-Yuan Mark Liao, Yong-Sheng Chen |
ICIP | 5 |
| 2009 | Lead Field Space Projection for Spatiotemporal Imaging of Independent Brain Activities
Huiling Chan, Yong-Sheng Chen, Li-Fen Chen, Tzu-Hua Chen, I-Tzu Chen |
ISNN (3) | 2 |
| 2007 | Fast and versatile algorithm for nearest neighbor search based on a lower bound tree
Yong-Sheng Chen, Yi-Ping Hung, Ting-Fang Yen, Chiou-Shann Fuh |
Pattern Recognit. | 1 |
| 2003 | A real-time robust eye tracking system for autostereoscopic displays using stereo camerasabstractAutostereoscopic display systems can provide users a natural 3D visualization environment by projecting stereo video onto the user's eyes. Eye position localization is a central module in this kind of display system when users are allowed to move freely. This paper presents robust 3D eye tracking techniques that can provide accurate eye positions in real time. Technical batteries comprise: (1) robust face detection based on eigenspace method; (2) real-time face tracking; and (3) eye detection in the obtained face region. According to our implementation on a PC with a Pentium IV 1.2 GHz CPU, the frame rate of the eye tracking process can achieve 25 Hz. Chan-Hung Su, Yong-Sheng Chen, Yi-Ping Hung, Chu-Song Chen, Jiun-Hung Chen |
ICRA | 2 |
| 2001 | Fast Algorithm for Nearest Neighbor Search Based on a Lower Bound Tree
Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICCV | 1 |
| 2001 | Simple and efficient method of calibrating a motorized zoom lens
Yong-Sheng Chen, Sheng-Wen Shih, Yi-Ping Hung, Chiou-Shann Fuh |
Image Vis. Comput. | 1 |
| 2001 | Three-dimensional ego-motion estimation from motion fields observed with multiple cameras
Yong-Sheng Chen, Lin-Gwo Liou, Yi-Ping Hung, Chiou-Shann Fuh |
Pattern Recognit. | 1 |
| 2001 | Fast block matching algorithm based on the winner-update strategyabstractBlock matching is a widely used method for stereo vision, visual tracking, and video compression. Many fast algorithms for block matching have been proposed in the past, but most of them do not guarantee that the match found is the globally optimal match in a search range. This paper presents a new fast algorithm based on the winner-update strategy which utilizes an ascending lower bound list of the matching error to determine the temporary winner. Two lower bound lists derived by using partial distance and by using Minkowski's inequality are described. The basic idea of the winner-update strategy is to avoid, at each search position, the costly computation of the matching error when there exists a lower bound larger than the global minimum matching error. The proposed algorithm can significantly speed up the computation of the block matching because: 1) computational cost of the lower bound we use is less than that of the matching error itself; 2) an element in the ascending lower bound list will be calculated only when its preceding element has already been smaller than the minimum matching error computed so far; 3) for many search positions, only the first several lower bounds in the list need to be calculated. Our experiments have shown that, when applying to motion vector estimation for several widely-used test videos, 92% to 98% of operations can be saved while still guaranteeing the global optimality. Moreover, the proposed algorithm can be easily modified either to meet the limited time requirement or to provide an ordered list of best candidate matches. Our source codes of the proposed algorithm are available at http://smart.iis.sinica.edu.tw/html/winup.html. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
IEEE Trans. Image Process. | 1 |
| 2000 | Winner-Update Algorithm for Nearest Neighbor SearchabstractThis paper presents an algorithm, called the winner-update algorithm, for accelerating the nearest neighbor search. By constructing a hierarchical structure for each feature point in the l/sub p/ metric space, this algorithm can save a large amount of computation at the expense of moderate preprocessing and twice the memory storage. Given a query point, the cost for computing the distances from this point to all the sample points can be reduced by using a lower bound list of the distance established from Minkowski's inequality. Our experiments have shown that the proposed algorithm can save a large amount of computation, especially when the distance between the query point and its nearest neighbor is relatively small. With slight modification, the winner-update algorithm can also speed up the search for k nearest neighbors, neighbors within a specified distance threshold, and neighbors close to the nearest neighbor. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh |
ICPR | 1 |
| 2000 | Camera Calibration with a Motorized Zoom LensabstractThis paper presents a simple and efficient method of calibrating the intrinsic camera parameters for all the lens settings of a motorized zoom lens. We fix the aperture setting and perform the camera calibration, adaptively, over the ranges of the room and focus settings. Bilinear interpolation is used to provide the values of the intrinsic camera parameters for those lens settings where no observations are taken. Our experiments show that the proposed method can provide accurate intrinsic camera parameters for all the lens settings, even though camera calibration is performed only for a small number of sampled lens settings. A calibration object suitable for zoom lens calibration is also presented. Yong-Sheng Chen, Yi-Ping Hung, Chiou-Shann Fuh, Sheng-Wen Shih |
ICPR | 1 |
| 1998 | Free-hand pointer by use of an active stereo vision systemabstractWe developed a system of free-hand pointer by tracking the finger of the speaker with an active stereo vision system and then computing the projection direction in either of the following two modes: the finger-orientation mode and the eye-to-fingertip mode. We prefer the latter for its robustness. In order to allow the speaker to move around in a wider 3D space without reducing the pointing resolution, we utilize a well-calibrated active stereo vision system which has a relatively small field of view but can control its stereo cameras to fixate at the moving finger. Our experiments have successfully demonstrated the feasibility of developing a free-handpointer using an active stereo vision system. Yi-Ping Hung, Yao-Strong Yang, Yong-Sheng Chen, Ing-Bor Hsieh, Chiou-Shann Fuh |
ICPR | 3 |
| 1998 | Multipass hierarchical stereo matching for generation of digital terrain models from aerial images
Yi-Ping Hung, Chu-Song Chen, Kuan-Chung Hung, Yong-Sheng Chen, Chiou-Shann Fuh |
Mach. Vis. Appl. | 4 |
| 1997 | Ego-Motion Estimation Using Optical Flow Fields Observed from Multiple CamerasabstractIn this paper, we consider a multi-camera vision system mounted on a moving object in a static three-dimensional environment. By using the motion flow fields seen by all of the cameras, an algorithm which does not need to solve the point-correspondence problem among the cameras is proposed to estimate the 3D ego-motion parameters of the moving object. Our experiments have shown that using multiple optical flow fields obtained from different cameras can be very helpful for ego-motion estimation. An-Ting Tsao, Chiou-Shann Fuh, Yi-Ping Hung, Yong-Sheng Chen |
CVPR | 4 |