Ping-Yang Chen

dblp:211/3420 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-9834-3671ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 US3Net: Ultralightweight Self-Supervised Stereo Matching Network using Depth-Aware Geometric Soft Occlusion
abstract
Abstract An ultralightweight self-supervised stereo matching network, called US $$^3$$ 3 Net, which requires only 12K parameters, is designed for efficient and accurate depth estimation using resource-constrained devices. US $$^3$$ 3 Net incorporates two key innovations: a low-complexity feature extraction module and a soft occlusion detection approach for performance improvement. First, we design a low-complexity feature extraction module to reduce the computational burden while preserving structural details necessary for stereo matching. By refining the encoder backbone and aggregation module, our design ensures a better balance between model complexity and accuracy. Second, to address occlusion-related errors in disparity estimation, we propose a novel occlusion detection method, called Depth-Aware Geometric Soft Occlusion (DAGSO), to adaptively define the occlusion confidence scores based on depth information. DAGSO can effectively mitigate false occlusions in distant regions and can improve the accuracy of disparity estimation. Experimental results using KITTI datasets demonstrate that US $$^3$$ 3 Net achieves state-of-the-art performance in terms of model complexity and depth estimation accuracy. It outperforms previous self-supervised stereo matching methods and monocular depth estimation methods in metrics such as AbsRel, SqRel, RMSE, and RMSElog at a reduction of parameter size by 47% compared with ES $$^3$$ 3 Net (23K). This makes US $$^3$$ 3 Net a practical solution for real-time depth estimation on edge devices such as drones and autonomous systems. Code is available at: https://github.com/g830319ag/US3Net.
Po-Chung Jen, Tzu-Chi Liu, I-Sheng Fang, Hsiao-Chieh Wen, Chia-Lun Hsu, Ping-Yang Chen, Chang-Hsing Lee, Yong-Sheng Chen
Int. J. Comput. Vis.6
2024 SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object Tracking
abstract
Despite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper introduces SMILEtrack, an innovative object tracker that effectively addresses these challenges by integrating an efficient object detector with a Siamese network-based Similarity Learning Module (SLM). The technical contributions of SMILETrack are twofold. First, we propose an SLM that calculates the appearance similarity between two objects, overcoming the limitations of feature descriptors in Separate Detection and Embedding (SDE) models. The SLM incorporates a Patch Self-Attention (PSA) block inspired by the vision Transformer, which generates reliable features for accurate similarity matching. Second, we develop a Similarity Matching Cascade (SMC) module with a novel GATE function for robust object matching across consecutive video frames, further enhancing MOT performance. Together, these innovations help SMILETrack achieve an improved trade-off between the cost (e.g., running speed) and performance (e.g., tracking accuracy) over several existing state-of-the-art benchmarks, including the popular BYTETrack method. SMILETrack outperforms BYTETrack by 0.4-0.8 MOTA and 2.1-2.2 HOTA points on MOT17 and MOT20 datasets. Code is available at http://github.com/pingyang1117/SMILEtrack_official.
Yu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Hung-Hin So, Xin Li 0005
AAAI3
2023 SARAS-Net: Scale and Relation Aware Siamese Network for Change Detection
abstract
Change detection (CD) aims to find the difference between two images at different times and output a change map to represent whether the region has changed or not. To achieve a better result in generating the change map, many State-of-The-Art (SoTA) methods design a deep learning model that has a powerful discriminative ability. However, these methods still get lower performance because they ignore spatial information and scaling changes between objects, giving rise to blurry boundaries. In addition to these, they also neglect the interactive information of two different images. To alleviate these problems, we propose our network, the Scale and Relation-Aware Siamese Network (SARAS-Net) to deal with this issue. In this paper, three modules are proposed that include relation-aware, scale-aware, and cross-transformer to tackle the problem of scene change detection more effectively. To verify our model, we tested three public datasets, including LEVIR-CD, WHU-CD, and DSFIN, and obtained SoTA accuracy. Our code is available at https://github.com/f64051041/SARAS-Net.
Chao-Peng Chen, Jun-Wei Hsieh, Ping-Yang Chen, Yi-Kuan Hsieh, Bor-Shiun Wang
AAAI3
2023 Fisheye Multiple Object Tracking by Learning Distortions Without Dewarping
abstract
We develop a new Multiple Object Tracking (MOT) scheme for fisheye cameras that can directly perform vehicle detection, re-identification, and tracking under fisheye distortions without explicit dewarping. Fisheye cameras provide omnidirectional coverage that is wider than traditional cameras, reducing fewer need of cameras to monitor road intersections. However, the problem of distorted views introduces new challenges for fisheye MOT. In this paper, we propose a Fish-Eye Multiple Object Tracking (FEMOT) approach with two novelties. We develop the Distorted Fisheye Image Augmentation (DFIA) method to improve object detection and re-identification on fisheye cameras, where fisheye model training can be performed on existing datasets of traditional cameras via fisheye data synthesis and augmentation. We also develop the Hybrid Data Association (HDA) method to perform tracking directly on fisheye views, without the need of de-warping. The developed FEMOT framework provides practical design and advancement that enables large-scale use of fisheye cameras in smart city and surveillance applications.
Ping-Yang Chen, Jun-Wei Hsieh, Ming-Ching Chang, Munkhjargal Gochoo, Fang-Pang Lin, Yong-Sheng Chen
ICIP1
2022 COFENet: Co-Feature Neural Network Model for Fine-Grained Image Classification
abstract
It is challenging to classify patterns with small inter-class variations but large intra-class variations especially for textured objects with relatively small sizes and blurry boundaries. We propose the Co-Feature Network (COFENet), a novel deep learning network for fine-grained texture-based image classification. State-of-the-art (SoTA) methods on this mostly rely on feature concatenation by merging convolutional features into fully connected layers. Some existing work explored the variation between pair-wise features during learning, they only considered the relations in the feature channels, and did not explore the spatial or structural relations among the image regions where the features are extracted from. We propose to leverage such information among the features and their relative spatial layouts to capture richer pairwise, orientationwise, and distancewise relations among feature channels for end-to-end learning of intra-class and inter-class variations.
Bor-Shiun Wang, Jun-Wei Hsieh, Yi-Kuan Hsieh, Ping-Yang Chen
ICIP4
2022 Mixed Stage Partial Network and Background Data Augmentation for Surveillance Object Detection
abstract
State-of-the-art (SoTA) object detection models and their accuracy have been improved by a large margin via CNNs (Convolutional Neural Networks); however, these models still perform poorly for small road objects. Moreover, the SoTA models are mainly trained on public benchmark datasets such as MS COCO, which include more complicated backgrounds and thus make them robust for object detection. However, for surveillance or road videos, their monotone backgrounds make these SoTA detectors background-over-fitted. In applications such as autonomous driving or traffic flow estimation, the background-over-fitting problem will increase various challenges and lead to accuracy degradation in object detection. One novelty of this paper is to propose an MBA (Mixed Background Augmentation) method to improve detection accuracy without adding new labeling efforts and any pre-training processes. During the inference stage, only one input image is needed for vehicle detection without involving background subtraction. Another novelty of this paper is the design of an efficient MSP (Mixed Stage Partial) network to detect objects more accurately and efficiently from surveillance videos. Extensive experiments on KITTI and UA-DETRAC benchmarks show that the proposed method achieves the SoTA results for highly accurate and efficient vehicle detection. The detection accuracy is improved from 78.53% to 83.59% with 25.7$fps$on the UA-DETRAC data set. The implementation code is available athttps://github.com/pingyang1117/MSPNet.
Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen
IEEE Trans. Intell. Transp. Syst.1
2021 Learnable Discrete Wavelet Pooling (LDW-Pooling) for Convolutional Networks
Bor-Shiun Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Lipeng Ke, Siwei Lyu
BMVC3
2021 Light-Weight Mixed Stage Partial Network for Surveillance Object Detection with Background Data Augmentation
abstract
State-of-the-art (SoTA) models have improved object detection accuracy with a large margin via convolutional neural networks, however still with an inferior performance for small objects. Moreover, these models are trained mainly based on the COCO dataset, and its backgrounds are more complicated than road environments, and thus degrade the accuracy of small road object detection. Compared with the COCO dataset, the background of a surveillance video is relatively stable and can be used to enhance the accuracy of road object detection. This paper designs a computationally efficient mixed stage partial (MSP) network to detect road objects. Another novelty of this paper is to propose a mixed background data augmentation method to enhance the detection accuracy without adding new labelling efforts. During inference, only the input image is used to detect road objects without further using any subtraction information. Extensive experiments on KITTI and UA-DETRAC benchmarks show the proposed method achieves the SoTA results for highly-accurate and efficient road object detection.
Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen
ICIP1
2021 Towards Deep Learning-Based Sarcopenia Screening with Body Joint Composition Analysis
abstract
Sarcopenia, a newly recognized geriatric syndrome, now prevalent in the rapidly aging region of Asia, is characterized by the age-related decline of skeletal muscle mass plus relatively low muscle strength and/or physical performance. Doctors screen for sarcopenia by observing patients’ habitual gait features without quantification and the performance of gait disturbances differ in various people that are considered to be sarcopenic, which is an important basis along with reduced physical functioning for the diagnosis of sarcopenia. Such a subjective diagnosis has been seen as a problem because diagnostic results may differ among doctors and factors such as fatigue may affect diagnosis. To strengthen and aid the use of these observations, we built a novel automatic deep learning model based on random forest for real-time human body joint detection coupled with a modified Long Short-Term Memory (LSTM) to recognize gait features for further clinical analysis. Aligned with the Asian Working Group for Sarcopenia (AWGS) [1] aims, our goal is to facilitate the implementation of standardized sarcopenia diagnosis in clinical practice by providing an automatic gait analysis system. Our model is recorded from geriatric patients for whole gait understanding. Experimental results demonstrate that our proposed model improves gait recognition performance compared to baseline methods. We believe, the quantitative evaluation provided by our method will assist the clinical diagnosis of sarcopenia and the experimental results on our gait datasets verify the feasibility and effectiveness of the proposed method.
Yung-Chih Chen, Jun-Wei Hsieh, Yao-Hong Yang, Chien-Hung Lee, Pei-Yi Yu, Ping-Yang Chen, Arpita Samanta Santa
ICIP6
2021 Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object Detection
abstract
This paper proposes the Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN) for fast and accurate single-shot object detection. Feature Pyramid (FP) is widely used in recent visual detection, however the top-down pathway of FP cannot preserve accurate localization due to pooling shifting. The advantage of FP is weakened as deeper backbones with more layers are used. In addition, it cannot keep up accurate detection of both small and large objects at the same time. To address these issues, we propose a new parallel FP structure with bi-directional (top-down and bottom-up) fusion and associated improvements to retain high-quality features for accurate localization. We provide the following design improvements: (1) A parallel bifusion FP structure with a bottom-up fusion module (BFM) to detect both small and large objects at once with high accuracy. (2) A concatenation and re-organization (CORE) module provides a bottom-up pathway for feature fusion, which leads to the bi-directional fusion FP that can recover lost information from lower-layer feature maps. (3) The CORE feature is further purified to retain richer contextual information. Such CORE purification in both top-down and bottom-up pathways can be finished in only a few iterations. (4) The adding of a residual design to CORE leads to a new Re-CORE module that enables easy training and integration with a wide range of deeper or lighter backbones. The proposed network achieves state-of-the-art performance on the UAVDT17 and MS COCO datasets. Code is available at https://github.com/pingyang1117/PRBNet_PyTorch.
Ping-Yang Chen, Ming-Ching Chang, Jun-Wei Hsieh, Yong-Sheng Chen
IEEE Trans. Image Process.1
2020 Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at Intersections
abstract
Drones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency.
Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao
ICIP1
2020 Lownet: Privacy Preserved Ultra-Low Resolution Posture Image Classification
abstract
Indoor posture recognition is vital for monitoring/detecting exercises, activities of daily living, accidental falls, unusual behavior, etc. However, high-resolution image based systems have a high accuracy, they are considered as intrusive and most of the current state-of-the-art image classifiers (VGG, ImageNet, ResNext) are not applicable for ultra-low resolution (<; 32 pixels in extent) image classification due to their downsizing feature extraction architecture. Thus, we propose a shallow LowNet model for classifying privacy preserved 16x16 posture images with its feature preserving architecture, variable ReLU slopes, and a custom loss function. LowNet outperformed, with an Accuracy of 98.94% and F1-score of 79.86%, the existing models (LeNet, ResNet1, ResNet-2) which can run on our Ultra lowresolution Thermal Posture Image (UTPI38) dataset (offered here) with 38 classes (4374 samples) collected from 23 volunteers. More experimental results are discussed on the custom loss, and variable ReLU slopes which gave 8.2% performance increase. Thus, we conclude that LowNet is useful in a multiclass ultra-low-resolution thermal posture image classification task.
Munkhjargal Gochoo, Tan-Hsu Tan, Fady Shibata-Alnajjar, Jun-Wei Hsieh, Ping-Yang Chen
ICIP5
2020 Deep Real-time Hand Detectoin Using CFPN on Embedded Systems
abstract
Real-time HI (Human Interface) systems need accurate and efficient hand detection models to meet the limited resources in budget, dimension, memory, computing, and electric power. In recent years, object detection became a less challenging task with the latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide the desired efficiency and accuracy for HI systems on embedded devices due to their complex time-consuming architecture. In addition, the detection of small hands ( pixels) is still a challenging task for all the above existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) to provide above mentioned performance for small hand detection. The superiority of CFPN is confirmed on a HandFlow dataset with mAP:0.5 of 95.6 and FPS of 33 on Nvidia TX2. The COCO dataset is also used to compare with other state-of-the-art method and shows the highest efficiency and accuracy with the proposed CFPN model. Thus we conclude that the proposed model is useful for real-life small hand detection on embedded devices.
Pirdiansyah Hendri, Jun-Wei Hsieh, Ping-Yang Chen, Munkhjargal Gochoo, Yong-Sheng Chen
ICPR3
2019 Real-Time Video-Based Person Re-Identification Surveillance with Light-Weight Deep Convolutional Networks
abstract
Today's person re-ID system mostly focuses on accuracy and ignores efficiency. But in most real-world surveillance systems, efficiency is often considered the most important focus of research and development. Therefore, for a person re-ID system, the ability to perform real-time identification is the most important consideration. In this study, we implemented a real-time multiple camera video-based person re-ID system using the NVIDIA Jetson TX2 platform. This system can be used in a field that requires high privacy and immediate monitoring. This system uses YOLOv3-tiny based light-weight strategies and person re-ID technology, thus reducing 46% of computation, cutting down 39.9% of model size, and accelerating 21% of computing speed. The system also effectively upgrades the pedestrian detection accuracy. In addition, the proposed person re-ID example mining and training method improves the model's performance and enhances the robustness of cross-domain data. Our system also supports the pipeline formed by connecting multiple edge computing devices in series. The system can operate at a speed up to 18 fps at 1920×1080 surveillance video stream. The demo of our developed systems can be found at https://sites.google.com/g.ncu.edu.tw/video-based-person-re-id/.
Chien-Yao Wang, Ping-Yang Chen, Ming-Chiao Chen, Jun-Wei Hsieh, Hong-Yuan Mark Liao
AVSS2
2019 Smaller Object Detection for Real-Time Embedded Traffic Flow Estimation Using Fish-Eye Cameras
abstract
Real-time embedded traffic flow estimation (RETFE) systems need accurate and efficient vehicle detection models to meet limited resources in budget, dimension, memory, and computing power. In recent years, object detection became a less challenging task with latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide desired performance for RETFE systems due to their complex time-consuming architecture. In addition, small object (<; 30×30 pixels) detection is still a challenging task for existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) that inspired from YOLOv3 to provide above mentioned performance for the smaller object detection. Main contribution is a proposed concatenated block (CB) which has reduced number of convolutional layers and concatenations instead of time-consuming algebraic operations. The superiority of CFPN is confirmed on the COCO and an in-house CarFlow datasets on Nvidia TX2. Thus we conclude that CFPN is useful for real-time embedded smaller object detection task.
Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Chien-Yao Wang, Hong-Yuan Mark Liao
ICIP1