VLDB 2026 Research / reviewers in the wild / expert
Jun-Wei Hsieh
dblp:83/5722
· DBLP profile ↗
105ranked-venue papers
31as first author
31since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 17 first-author · 24 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 3 since 2021Systems, architecture and hardware · 8 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorComputer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RepSFNet : A Single Fusion Network with Structural Reparameterization for Crowd CountingabstractCrowd counting remains challenging in variable-density scenes due to scale variations, occlusions, and the high computational cost of existing models. To address this, we propose RepSFNet (Reparameterized Single Fusion Network), a lightweight architecture designed for accurate and real-time crowd estimation. RepSFNet combines large-kernel convolutional power with a efficient, suitable for low-power edge computing. The architecture includes three components: (i) a RepLK-ViT backbone using large reparameterized kernels for efficient multi-scale feature extraction; (ii) a Feature Fusion module that integrates ASPP and CAN for robust, density adaptive context modeling; and (iii) a Concatenate Fusion module to preserve spatial resolution and produce high-quality density maps. By avoiding attention mechanisms and multi-branch designs, RepSFNet reduces both parameters and FLOPs, enhancing runtime efficiency. The loss function combines Mean Squared Error (MSE) and Optimal Transport (OT), further improving count accuracy. Experiments on ShanghaiTech, NWPU, and UCF-QNRF show that RepSFNet delivers competitive accuracy with up to 34% lower inference latency compared to P2PNet, M-SFANet, M-SegNet, STEERER, and Gramformer, making it more efficient and suitable for low-power edge computing. Mas Nurul Achmadiah, Chi-Chia Sun, Wen-Kai Kuo, Jun-Wei Hsieh |
AVSS | 4 |
| 2025 | ISSR-UNet: Intrinsic Supervision Shuffle Residual UNet for Underwater Image RestorationabstractUnderwater image restoration is a challenging task due to color distortion and hazing effects caused by light absorption and scattering. We propose a novel image restoration architecture, the Intrinsic Supervision Shuffle Residual UNet (ISSR-UNet), based on the UNet framework. ISSR-UNet leverages Adaptive Selective Intrinsic Supervised Features and a cross-scale feature shuffling mechanism to enhance image restoration performance and preserve more image details. It incorporates four pivotal components: the Selective Residual (SR) block, the Cross-scale Feature Shuffling (CFS) module, the Multi-Degradation Supervision (MDS) module, and the Adaptive Selective Intrinsic Supervised Feature Extraction (ASISFE) module. Extensive experiments validate that ISSR-UNet achieves superior performance, surpassing state-of-the-art methods on underwater image restoration benchmarks. Chia-Cheng Chang, Jun-Wei Hsieh, Chiao-Ching Chou, Yu-Hong Lee, Pin-Wei Lin, Ling-Yun Chu |
AVSS | 2 |
| 2025 | PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and EvaluationabstractThis paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance. Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee |
AVSS | 24 |
| 2025 | FaceLiVT: Face Recognition Using Linear Vision Transformer with Structural Reparameterization for Mobile DeviceabstractThis paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovative and lightweight Multi-Head Linear Attention (MHLA) mechanism. By combining MHLA alongside a reparameterized token mixer, FaceLiVT effectively reduces computational complexity while preserving competitive accuracy. Extensive evaluations on challenging benchmarks—including LFW, CFP-FP, AgeDB-30, IJB-B, and IJB-C—highlight its superior performance compared to state-of-the-art lightweight models. MHLA notably improves inference speed, allowing FaceLiVT to deliver high accuracy with lower latency on mobile devices. Specifically, FaceLiVT is 8.6× faster than EdgeFace, a recent hybrid CNN-Transformer model optimized for edge devices, and 21.2× faster than a pure ViT-Based model. With its balanced design, FaceLiVT offers an efficient and practical solution for real-time face recognition on resource-constrained platforms. Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jun-Wei Hsieh |
ICIP | 5 |
| 2025 | BF-YOLOv7: Enhancing Helmet Rule Violation Detection
Chun-Ming Tsai, Jun-Wei Hsieh, Ming-Ching Chang |
IEA/AIE (2) | 2 |
| 2025 | MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge DeviceabstractThe Vision Transformer (ViT) has demonstrated state-of-the-art performance in various computer vision tasks, but its high computational demands make it impractical for edge devices with limited resources. This paper presents MicroViT, a lightweight Vision Transformer architecture optimized for edge devices by significantly reducing computational complexity while maintaining high accuracy. The core of MicroViT is the Efficient Single Head Attention (ESHA) mechanism, which utilizes group convolution to reduce feature redundancy and processes only a fraction of the channels, thus lowering the burden of the self-attention mechanism. MicroViT is designed using a multi-stage MetaFormer architecture, stacking multiple MicroViT encoders to enhance efficiency and performance. Comprehensive experiments on the ImageNet-1K and COCO datasets demonstrate that MicroViT achieves competitive accuracy while significantly improving 3.6× faster inference speed and reducing energy consumption with 40% higher efficiency than the MobileViT series, making it suitable for deployment in resource-constrained environments such as mobile and edge devices. Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jun-Wei Hsieh |
ISCAS | 5 |
| 2025 | Scale-Aware Crowd Counting Network With Annotation Error ModelingabstractTraditional crowd-counting networks suffer from information loss when feature maps are reduced by pooling layers, leading to inaccuracies in counting crowds at a distance. Existing methods often assume correct annotations during training, disregarding the impact of noisy annotations, especially in crowded scenes. Furthermore, using a fixed Gaussian density model does not account for the varying pixel distribution of the camera distance. To overcome these challenges, we propose a Scale-Aware Crowd Counting Network (SACC-Net) that introduces a scale-aware loss function with error-compensation capabilities of noisy annotations. For the first time, we simultaneously model labeling errors (mean) and scale variations (variance) by spatially varying Gaussian distributions to produce fine-grained density maps for crowd counting. Furthermore, the proposed scale-aware Gaussian density model can be dynamically approximated with a low-rank approximation, leading to improved convergence efficiency with comparable accuracy. To create a smoother scale-aware feature space, this paper proposes a novel Synthetic Fusion Module (SFM) and an Intra-block Fusion Module (IFM) to generate fine-grained heat maps for better crowd counting. The lightweight version of our model, named SACC-LW, enhances the computational efficiency while retaining accuracy. The superiority and generalization properties of scale-aware loss function are extensively evaluated for different backbone architectures and performance metrics on six public datasets: UCF-QNRF, UCF CC 50, NWPU, ShanghaiTech A, ShanghaiTech B, and JHU. Experimental results also demonstrate that SACC-Net outperforms all state-of-the-art methods, validating its effectiveness in achieving superior crowd-counting accuracy. The source code is available at https://github.com/Naughty725. Yi-Kuan Hsieh, Jun-Wei Hsieh, Xin Li 0005, Yu-Ming Zhang, Yu-Chee Tseng, Ming-Ching Chang |
IEEE Trans. Image Process. | 2 |
| 2025 | Cross-Scale Overlapping Patch-Based Attention Network for Road Crack DetectionabstractCracks on road surfaces pose serious risks to both pedestrians and drivers. Traditional manual crack detection methods are not only slow but also pose safety risks. Automating this process has the potential to greatly enhance detection efficiency and consequently improve driving safety. Although previous methods have shown promise in road crack detection, they often neglect interactions between multiple scales, causing smaller cracks to be overlooked in later stages of detection. This paper introduces the Cross-scale Overlapping Patch-based attention Network (COP-Net), which incorporates two critical components: the Scale-aware Channel Attention (SCA) module and the Patch-based Cross-scale Attention (PCA) module for crack detection. These innovations enable dynamic inference on multiple scales, resulting in a significant improvement in crack detection and segmentation. Notably, our approach excels at detecting both small and large cracks simultaneously. To validate the effectiveness of our approach, we conducted evaluations on three open datasets: CRACK500, CFD, and AEL. These evaluation results demonstrate that COP-Net surpasses eleven comparison methods, including HED, DeepCrack, UHDN, SSGNet, MFANet, FPHBN, DeepCrack, PBNet, PAFNet, CarNet, and SegFormer. Our model achieves new State-of-The-Art (SoTA) performance levels in terms of segmentation metrics such as AIU, ODS, and OIS. Jun-Wei Hsieh, Yi-Kuan Hsieh, Chuan-Wang Chang, Deng-Yuan Huang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Lightweight Computation Single-Image Fog Removal Based on a New Improved Adaptive Dark Channel PriorabstractThis paper addresses the challenge of image quality degradation in intelligent vehicle systems caused by adverse weather conditions, particularly fog, on embedded platforms. It proposes an optimized dark channel prior (DCP)-based fog removal algorithm, optimized for real-time processing on resource-constrained platforms, such as those equipped with a single neural processing unit (NPU). The algorithm introduces an adaptive mechanism to address overexposure issues commonly found in regions with sky or sun, ensuring robust performance across diverse hazy conditions and exposure levels. Tailored for intelligent vehicle applications, the method supports real-time processing of high-definition video, delivering a speed 151.14 times faster than state-of-The-Art (SoTA) learning-based approaches like FFA-Net and DehazeNet, while achieving a structural similarity index (SSIM) of up to 0.914 compared to reference fog-free images. Moreover, it significantly enhances critical vision-based tasks, such as YOLOv4 object detection, improving object detection accuracy by up to 18%. This lightweight and efficient solution makes it particularly suitable for deployment in advanced driver assistance systems (ADAS) and autonomous navigation platforms, contributing to improved safety and reliability in foggy conditions. Chi-Chia Sun, Nguyen Hoang Hai Pham, Achmad Arif Bryantono, Jun-Wei Hsieh |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Pushing the Limit of Fine-Tuning for Few-Shot Learning: Where Feature Reusing Meets Cross-Scale AttentionabstractDue to the scarcity of training samples, Few-Shot Learning (FSL) poses a significant challenge to capture discriminative object features effectively. The combination of transfer learning and meta-learning has recently been explored by pre-training the backbone features using labeled base data and subsequently fine-tuning the model with target data. However, existing meta-learning methods, which use embedding networks, suffer from scaling limitations when dealing with a few labeled samples, resulting in suboptimal results. Inspired by the latest advances in FSL, we further advance the approach of fine-tuning a pre-trained architecture by a strengthened hierarchical feature representation. The technical contributions of this work include: 1) a hybrid design named Intra-Block Fusion (IBF) to strengthen the extracted features within each convolution block; and 2) a novel Cross-Scale Attention (CSA) module to mitigate the scaling inconsistencies arising from the limited training samples, especially for cross-domain tasks. We conducted comprehensive evaluations on standard benchmarks, including three in-domain tasks (miniImageNet, CIFAR-FS, and FC100), as well as two cross-domain tasks (CDFSL and Meta-Dataset). The results have improved significantly over existing state-of-the-art approaches on all benchmark datasets. In particular, the FSL performance on the in-domain FC100 dataset is more than three points better than the latest PMF (Hu et al. 2022). Ying-Yu Chen, Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang |
AAAI | 2 |
| 2024 | SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object TrackingabstractDespite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper introduces SMILEtrack, an innovative object tracker that effectively addresses these challenges by integrating an efficient object detector with a Siamese network-based Similarity Learning Module (SLM). The technical contributions of SMILETrack are twofold. First, we propose an SLM that calculates the appearance similarity between two objects, overcoming the limitations of feature descriptors in Separate Detection and Embedding (SDE) models. The SLM incorporates a Patch Self-Attention (PSA) block inspired by the vision Transformer, which generates reliable features for accurate similarity matching. Second, we develop a Similarity Matching Cascade (SMC) module with a novel GATE function for robust object matching across consecutive video frames, further enhancing MOT performance. Together, these innovations help SMILETrack achieve an improved trade-off between the cost (e.g., running speed) and performance (e.g., tracking accuracy) over several existing state-of-the-art benchmarks, including the popular BYTETrack method. SMILETrack outperforms BYTETrack by 0.4-0.8 MOTA and 2.1-2.2 HOTA points on MOT17 and MOT20 datasets. Code is available at http://github.com/pingyang1117/SMILEtrack_official. Yu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Hung-Hin So, Xin Li 0005 |
AAAI | 2 |
| 2024 | Strengthening 3D Point Cloud Classification through Self-Attention and Plane FeaturesabstractTo address the unique attributes of three-dimensional point cloud data, this paper introduces an innovative architecture for precise 3D point cloud classification. Expanding on the PointMLP framework, we incorporate an embedding module that elevates the point cloud to higher-dimensional feature representations which are followed by geometric feature mapping and extraction modules to capture point cloud characteristics. To estimate local geometric structures, we use plane features to determine planes associated with nearby points. Additionally, we integrate self-attention mechanisms to capture intricate local geometric features. Moreover, MLP modules with residual connections are employed for efficient feature extraction. The derived features are then reduced in size using Max Pooling layers. For classification purposes, we utilize fully connected layers, batch normalization, activation functions, and random weight dropping techniques to enhance generalization and ensure robustness on unseen data. By adopting these architectural decisions, our proposed model achieves significant progress in accurately classifying 3D point clouds. Jin-Cheng Liu, Jun-Wei Hsieh, Yu-Ming Zhang, Chun-Chieh Lee, Kuo-Chin Fan |
AVSS | 2 |
| 2024 | Image Manipulation Detection with Implicit Neural Representation and Limited Supervision
Zhenfei Zhang, Mingyang Li 0007, Xin Li 0005, Ming-Ching Chang, Jun-Wei Hsieh |
ECCV (88) | 5 |
| 2024 | Class-Specific Channel Attention For Few Shot LearningabstractFew-Shot Learning (FSL) has attracted growing attention in computer vision due to its capability in model training without the need for excessive data. FSL is challenging because the training and testing categories (the base vs. novel sets) can be largely diversified. Conventional transfer-based solutions that aim to transfer knowledge learned from large labeled training sets to target testing sets are limited, as critical adverse impacts of the shift in task distribution are not adequately addressed. In this paper, we extend the solution of transfer-based methods by incorporating the concept of metric-learning and channel attention. To better exploit the feature representations extracted by the feature backbone, we propose Class-Specific Channel Attention (CSCA) module, which learns to highlight the discriminative channels in each class by assigning each class one CSCA weight vector. Unlike general attention modules designed to learn globalclass features, the CSCA module aims to learn local and class-specific features with very effective computation. We evaluated the performance of the CSCA module on standard benchmarks including miniImagenet, Tiered-ImageNet, CIFAR-FS, and CUB-200-2011. Experiments are performed in inductive and in/cross-domain settings. We achieve new state-of-the-art results. Yi-Kuan Hsieh, Jun-Wei Hsieh, Ying-Yu Chen |
ICIP | 2 |
| 2024 | Set-Nas: Sample-Efficient Training For Neural Architecture Search With Strong Predictor And Stratified SamplingabstractSample-efficient neural architecture search (NAS) techniques have advanced rapidly. Two lines of methods, namely neural predictor and sequential search, have shown promising performance in improving the sample efficiency of NAS. However, as far as we know, little attention has been paid to the middle ground between these two lines. Inspired by the analogy between NAS and evolutionary optimization, we propose a new Sample-Efficient Training for NAS (SETNAS) based on strategies that improve fitness scores and sampling mechanisms. We develop a strong neural predictor called the Fully Bidirectional Graph Convolutional Network evolutionary (Fully-BiGCN) that significantly enhances the predictor capability of the features in each layer. The developed predictor is embedded into an iterative stratified sampling process to retain only a subset of best-fit architectures using the same training budget. SET-NAS achieves remarkable results compared to the state-of-the-art in predictor-based NAS. Using NASBench-201 as the benchmark, SET-NAS takes only $27.1 \%$ (CIFAR-10), $49.0 \%$ (CIFAR-100), and $51.75 \%$ (ImageNet-16) of training cost of other state-of-the-art predictor-based methods to find the promising network architecture. Yu-Ming Zhang, Jun-Wei Hsieh, Yu-Hsiu Chang, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan |
ICIP | 2 |
| 2024 | Patch-Based Prototypical Cross-Scale Attention Network for Anomaly Detection
Tung-Lin Wang, Jun-Wei Hsieh, Yi-Kuan Hsieh |
ICPR (2) | 2 |
| 2024 | MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture SearchabstractNeural Architecture Search (NAS) methods seek effective optimization toward performance metrics regarding model accuracy and generalization while facing challenges regarding search costs and GPU resources. Recent Neural Tangent Kernel (NTK) NAS methods achieve remarkable search efficiency based on a training-free model estimate; however, they overlook the non-convex nature of the DNNs in the search process. In this paper, we develop Multi-Objective Training-based Estimate (MOTE) for efficient NAS, retaining search effectiveness and achieving the new state-of-the-art in the accuracy and cost trade-off. To improve NTK and inspired by the Training Speed Estimation (TSE) method, MOTE is designed to model the actual performance of DNNs from macro to micro perspective by draw loss landscape and convergence speed simultaneously. Using two reduction strategies, the MOTE is generated based on a reduced architecture and a reduced dataset. Inspired by evolutionary search, our iterative ranking-based, coarse-to-fine architecture search is highly effective. Experiments on NASBench-201 show MOTE-NAS achieves 94.32% accuracy on CIFAR-10, 72.81% on CIFAR-100, and 46.38% on ImageNet-16-120, outperforming NTK-based NAS approaches. An evaluation-free (EF) version of MOTE-NAS delivers high efficiency in only 5 minutes, delivering a model more accurate than KNAS. Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan |
NeurIPS | 2 |
| 2024 | AIoT-Based Shrimp Larvae Counting System Using Scaled Multilayer Feature Fusion NetworkabstractThe Artificial Intelligence of Things (AIoT) plays a crucial role in shrimp farming by enabling automated and real-time monitoring of shrimp counting, especially larvae. With this counting information, proper feeding control can be maintained by equipping IoT devices to capture real-time data on water quality, temperature, and other environmental factors, ensuring healthy shrimp growth and increasing production. Taking advantage of the power of AIoT, this article proposes a scaled multilayer feature fusion network (SMILES-Net) to allow farmers to remotely and automatically manage the counting of shrimps, effectively reducing the need for manual labor, and enabling swift interventions to sustain shrimp well-being. Since shrimp larvae are extremely small, we frame this counting problem as a density prediction problem, where the sum of the constructed density map is the total number of shrimps predicted. The pooling operation used in convolutional neural networks scales each feature map to 1/4 and causes the rich features of small shrimps to disappear dramatically. To tackle this truncation problem, we propose a novel synthetic fusion module (SFM) and an intrablock fusion module (IFM) to create a smoother scale space, generating better heat maps with fine-grained features for shrimp counting. Furthermore, we introduce a lightweight version of SMILES-Net (LW-SMILES-Net) that enables real-time shrimp counting without compromising accuracy. This method is evaluated on different data sets for shrimp counting and outperforms all State-of-The-Art methods. Overall, integrating SMILES-Net with IoT devices can provide a powerful solution for real-time shrimp larvae counting in shrimp farming, contributing to sustainable increases in shrimp production. The data set is available athttps://github.com/Naughty725/shrimp. Yi-Kuan Hsieh, Jun-Wei Hsieh, Wu-Chih Hu, Yu-Chee Tseng |
IEEE Internet Things J. | 2 |
| 2023 | SARAS-Net: Scale and Relation Aware Siamese Network for Change DetectionabstractChange detection (CD) aims to find the difference between two images at different times and output a change map to represent whether the region has changed or not. To achieve a better result in generating the change map, many State-of-The-Art (SoTA) methods design a deep learning model that has a powerful discriminative ability. However, these methods still get lower performance because they ignore spatial information and scaling changes between objects, giving rise to blurry boundaries. In addition to these, they also neglect the interactive information of two different images. To alleviate these problems, we propose our network, the Scale and Relation-Aware Siamese Network (SARAS-Net) to deal with this issue. In this paper, three modules are proposed that include relation-aware, scale-aware, and cross-transformer to tackle the problem of scene change detection more effectively. To verify our model, we tested three public datasets, including LEVIR-CD, WHU-CD, and DSFIN, and obtained SoTA accuracy. Our code is available at https://github.com/f64051041/SARAS-Net. Chao-Peng Chen, Jun-Wei Hsieh, Ping-Yang Chen, Yi-Kuan Hsieh, Bor-Shiun Wang |
AAAI | 2 |
| 2023 | Fisheye Multiple Object Tracking by Learning Distortions Without DewarpingabstractWe develop a new Multiple Object Tracking (MOT) scheme for fisheye cameras that can directly perform vehicle detection, re-identification, and tracking under fisheye distortions without explicit dewarping. Fisheye cameras provide omnidirectional coverage that is wider than traditional cameras, reducing fewer need of cameras to monitor road intersections. However, the problem of distorted views introduces new challenges for fisheye MOT. In this paper, we propose a Fish-Eye Multiple Object Tracking (FEMOT) approach with two novelties. We develop the Distorted Fisheye Image Augmentation (DFIA) method to improve object detection and re-identification on fisheye cameras, where fisheye model training can be performed on existing datasets of traditional cameras via fisheye data synthesis and augmentation. We also develop the Hybrid Data Association (HDA) method to perform tracking directly on fisheye views, without the need of de-warping. The developed FEMOT framework provides practical design and advancement that enables large-scale use of fisheye cameras in smart city and surveillance applications. Ping-Yang Chen, Jun-Wei Hsieh, Ming-Ching Chang, Munkhjargal Gochoo, Fang-Pang Lin, Yong-Sheng Chen |
ICIP | 2 |
| 2022 | COFENet: Co-Feature Neural Network Model for Fine-Grained Image ClassificationabstractIt is challenging to classify patterns with small inter-class variations but large intra-class variations especially for textured objects with relatively small sizes and blurry boundaries. We propose the Co-Feature Network (COFENet), a novel deep learning network for fine-grained texture-based image classification. State-of-the-art (SoTA) methods on this mostly rely on feature concatenation by merging convolutional features into fully connected layers. Some existing work explored the variation between pair-wise features during learning, they only considered the relations in the feature channels, and did not explore the spatial or structural relations among the image regions where the features are extracted from. We propose to leverage such information among the features and their relative spatial layouts to capture richer pairwise, orientationwise, and distancewise relations among feature channels for end-to-end learning of intra-class and inter-class variations. Bor-Shiun Wang, Jun-Wei Hsieh, Yi-Kuan Hsieh, Ping-Yang Chen |
ICIP | 2 |
| 2022 | SFPN: Synthetic FPN for Object DetectionabstractFPN (Feature Pyramid Network) has become a basic component of most SoTA one stage object detectors. Many previous studies have repeatedly proved that FPN can caputre better multi-scale feature maps to more precisely describe objects if they are with different sizes. However, for most backbones such VGG, ResNet, or DenseNet, the feature maps at each layer are downsized to their quarters due to the pooling operation or convolutions with stride 2. The gap of downscaling-by-2 is large and makes its FPN not fuse the features smoothly. This paper proposes a new SFPN (Synthetic Fusion Pyramid Network) arichtecture which creates various synthetic layers between layers of the original FPN to enhance the accuracy of light-weight CNN backones to extract objects’ visual features more accurately. Finally, experiments prove the SFPN architecture outperforms either the large backbone VGG16, ResNet50 or light-weight backbones such as MobilenetV2 based on AP score. Yu-Ming Zhang, Jun-Wei Hsieh, Chun-Chieh Lee, Kuo-Chin Fan |
ICIP | 2 |
| 2022 | CSL-YOLO: A Cross-Stage Lightweight Object Detector with Low FLOPsabstractThe development of lightweight object detectors is essential due to the limited computation resources. To reduce the computation cost, how to generate features plays a significant role. This paper proposes a new lightweight convolution method Cross-Stage Lightweight Module (CSL-M). It combines the Inverted Residual Block (IRB) and Cross-Stage Partial (CSP) concept. Experiments conducted at CIFAR-10 show that the proposed CSL-Net based on CSL-M performs better with fewer FLOPs than the other lightweight backbones. Finally, we use CSL-Net as the backbone to construct a lightweight detector CSL-YOLO, achieving better detection performance with only 43% FLOPs and 52% parameters than Tiny-YOLOv4. Yu-Ming Zhang, Chun-Chieh Lee, Jun-Wei Hsieh, Kuo-Chin Fan |
ISCAS | 3 |
| 2022 | Multi-fusion feature pyramid for real-time hand detection
Chuan-Wang Chang, Santanu Santra, Jun-Wei Hsieh, Pirdiansyah Hendri, Chi-Fang Lin |
Multim. Tools Appl. | 3 |
| 2022 | Mixed Stage Partial Network and Background Data Augmentation for Surveillance Object DetectionabstractState-of-the-art (SoTA) object detection models and their accuracy have been improved by a large margin via CNNs (Convolutional Neural Networks); however, these models still perform poorly for small road objects. Moreover, the SoTA models are mainly trained on public benchmark datasets such as MS COCO, which include more complicated backgrounds and thus make them robust for object detection. However, for surveillance or road videos, their monotone backgrounds make these SoTA detectors background-over-fitted. In applications such as autonomous driving or traffic flow estimation, the background-over-fitting problem will increase various challenges and lead to accuracy degradation in object detection. One novelty of this paper is to propose an MBA (Mixed Background Augmentation) method to improve detection accuracy without adding new labeling efforts and any pre-training processes. During the inference stage, only one input image is needed for vehicle detection without involving background subtraction. Another novelty of this paper is the design of an efficient MSP (Mixed Stage Partial) network to detect objects more accurately and efficiently from surveillance videos. Extensive experiments on KITTI and UA-DETRAC benchmarks show that the proposed method achieves the SoTA results for highly accurate and efficient vehicle detection. The detection accuracy is improved from 78.53% to 83.59% with 25.7$fps$on the UA-DETRAC data set. The implementation code is available athttps://github.com/pingyang1117/MSPNet. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Multi-Teacher Single-Student Visual Transformer with Multi-Level Attention for Face Spoofing Detection
Yao-Hui Huang, Jun-Wei Hsieh, Ming-Ching Chang, Lipeng Ke, Siwei Lyu, Arpita Samanta Santra |
BMVC | 2 |
| 2021 | Learnable Discrete Wavelet Pooling (LDW-Pooling) for Convolutional Networks
Bor-Shiun Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Lipeng Ke, Siwei Lyu |
BMVC | 2 |
| 2021 | Light-Weight Mixed Stage Partial Network for Surveillance Object Detection with Background Data AugmentationabstractState-of-the-art (SoTA) models have improved object detection accuracy with a large margin via convolutional neural networks, however still with an inferior performance for small objects. Moreover, these models are trained mainly based on the COCO dataset, and its backgrounds are more complicated than road environments, and thus degrade the accuracy of small road object detection. Compared with the COCO dataset, the background of a surveillance video is relatively stable and can be used to enhance the accuracy of road object detection. This paper designs a computationally efficient mixed stage partial (MSP) network to detect road objects. Another novelty of this paper is to propose a mixed background data augmentation method to enhance the detection accuracy without adding new labelling efforts. During inference, only the input image is used to detect road objects without further using any subtraction information. Extensive experiments on KITTI and UA-DETRAC benchmarks show the proposed method achieves the SoTA results for highly-accurate and efficient road object detection. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Yong-Sheng Chen |
ICIP | 2 |
| 2021 | Towards Deep Learning-Based Sarcopenia Screening with Body Joint Composition AnalysisabstractSarcopenia, a newly recognized geriatric syndrome, now prevalent in the rapidly aging region of Asia, is characterized by the age-related decline of skeletal muscle mass plus relatively low muscle strength and/or physical performance. Doctors screen for sarcopenia by observing patients’ habitual gait features without quantification and the performance of gait disturbances differ in various people that are considered to be sarcopenic, which is an important basis along with reduced physical functioning for the diagnosis of sarcopenia. Such a subjective diagnosis has been seen as a problem because diagnostic results may differ among doctors and factors such as fatigue may affect diagnosis. To strengthen and aid the use of these observations, we built a novel automatic deep learning model based on random forest for real-time human body joint detection coupled with a modified Long Short-Term Memory (LSTM) to recognize gait features for further clinical analysis. Aligned with the Asian Working Group for Sarcopenia (AWGS) [1] aims, our goal is to facilitate the implementation of standardized sarcopenia diagnosis in clinical practice by providing an automatic gait analysis system. Our model is recorded from geriatric patients for whole gait understanding. Experimental results demonstrate that our proposed model improves gait recognition performance compared to baseline methods. We believe, the quantitative evaluation provided by our method will assist the clinical diagnosis of sarcopenia and the experimental results on our gait datasets verify the feasibility and effectiveness of the proposed method. Yung-Chih Chen, Jun-Wei Hsieh, Yao-Hong Yang, Chien-Hung Lee, Pei-Yi Yu, Ping-Yang Chen, Arpita Samanta Santa |
ICIP | 2 |
| 2021 | Air-writing recognition using reverse time ordered stroke context
Tsung-Hsien Tsai, Jun-Wei Hsieh, Chuan-Wang Chang, Chin-Rong Lay, Kuo-Chin Fan |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object DetectionabstractThis paper proposes the Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN) for fast and accurate single-shot object detection. Feature Pyramid (FP) is widely used in recent visual detection, however the top-down pathway of FP cannot preserve accurate localization due to pooling shifting. The advantage of FP is weakened as deeper backbones with more layers are used. In addition, it cannot keep up accurate detection of both small and large objects at the same time. To address these issues, we propose a new parallel FP structure with bi-directional (top-down and bottom-up) fusion and associated improvements to retain high-quality features for accurate localization. We provide the following design improvements: (1) A parallel bifusion FP structure with a bottom-up fusion module (BFM) to detect both small and large objects at once with high accuracy. (2) A concatenation and re-organization (CORE) module provides a bottom-up pathway for feature fusion, which leads to the bi-directional fusion FP that can recover lost information from lower-layer feature maps. (3) The CORE feature is further purified to retain richer contextual information. Such CORE purification in both top-down and bottom-up pathways can be finished in only a few iterations. (4) The adding of a residual design to CORE leads to a new Re-CORE module that enables easy training and integration with a wide range of deeper or lighter backbones. The proposed network achieves state-of-the-art performance on the UAVDT17 and MS COCO datasets. Code is available at https://github.com/pingyang1117/PRBNet_PyTorch. Ping-Yang Chen, Ming-Ching Chang, Jun-Wei Hsieh, Yong-Sheng Chen |
IEEE Trans. Image Process. | 3 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2020 | Lownet: Privacy Preserved Ultra-Low Resolution Posture Image ClassificationabstractIndoor posture recognition is vital for monitoring/detecting exercises, activities of daily living, accidental falls, unusual behavior, etc. However, high-resolution image based systems have a high accuracy, they are considered as intrusive and most of the current state-of-the-art image classifiers (VGG, ImageNet, ResNext) are not applicable for ultra-low resolution (<; 32 pixels in extent) image classification due to their downsizing feature extraction architecture. Thus, we propose a shallow LowNet model for classifying privacy preserved 16x16 posture images with its feature preserving architecture, variable ReLU slopes, and a custom loss function. LowNet outperformed, with an Accuracy of 98.94% and F1-score of 79.86%, the existing models (LeNet, ResNet1, ResNet-2) which can run on our Ultra lowresolution Thermal Posture Image (UTPI38) dataset (offered here) with 38 classes (4374 samples) collected from 23 volunteers. More experimental results are discussed on the custom loss, and variable ReLU slopes which gave 8.2% performance increase. Thus, we conclude that LowNet is useful in a multiclass ultra-low-resolution thermal posture image classification task. Munkhjargal Gochoo, Tan-Hsu Tan, Fady Shibata-Alnajjar, Jun-Wei Hsieh, Ping-Yang Chen |
ICIP | 4 |
| 2020 | Deep Real-time Hand Detectoin Using CFPN on Embedded SystemsabstractReal-time HI (Human Interface) systems need accurate and efficient hand detection models to meet the limited resources in budget, dimension, memory, computing, and electric power. In recent years, object detection became a less challenging task with the latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide the desired efficiency and accuracy for HI systems on embedded devices due to their complex time-consuming architecture. In addition, the detection of small hands ( pixels) is still a challenging task for all the above existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) to provide above mentioned performance for small hand detection. The superiority of CFPN is confirmed on a HandFlow dataset with mAP:0.5 of 95.6 and FPS of 33 on Nvidia TX2. The COCO dataset is also used to compare with other state-of-the-art method and shows the highest efficiency and accuracy with the proposed CFPN model. Thus we conclude that the proposed model is useful for real-life small hand detection on embedded devices. Pirdiansyah Hendri, Jun-Wei Hsieh, Ping-Yang Chen, Munkhjargal Gochoo, Yong-Sheng Chen |
ICPR | 2 |
| 2020 | Driver License Field Detection Using Real-Time Deep Networks
Chun-Ming Tsai, Jun-Wei Hsieh, Ming-Ching Chang |
IEA/AIE | 2 |
| 2019 | Real-Time Video-Based Person Re-Identification Surveillance with Light-Weight Deep Convolutional NetworksabstractToday's person re-ID system mostly focuses on accuracy and ignores efficiency. But in most real-world surveillance systems, efficiency is often considered the most important focus of research and development. Therefore, for a person re-ID system, the ability to perform real-time identification is the most important consideration. In this study, we implemented a real-time multiple camera video-based person re-ID system using the NVIDIA Jetson TX2 platform. This system can be used in a field that requires high privacy and immediate monitoring. This system uses YOLOv3-tiny based light-weight strategies and person re-ID technology, thus reducing 46% of computation, cutting down 39.9% of model size, and accelerating 21% of computing speed. The system also effectively upgrades the pedestrian detection accuracy. In addition, the proposed person re-ID example mining and training method improves the model's performance and enhances the robustness of cross-domain data. Our system also supports the pipeline formed by connecting multiple edge computing devices in series. The system can operate at a speed up to 18 fps at 1920×1080 surveillance video stream. The demo of our developed systems can be found at https://sites.google.com/g.ncu.edu.tw/video-based-person-re-id/. Chien-Yao Wang, Ping-Yang Chen, Ming-Chiao Chen, Jun-Wei Hsieh, Hong-Yuan Mark Liao |
AVSS | 4 |
| 2019 | Smaller Object Detection for Real-Time Embedded Traffic Flow Estimation Using Fish-Eye CamerasabstractReal-time embedded traffic flow estimation (RETFE) systems need accurate and efficient vehicle detection models to meet limited resources in budget, dimension, memory, and computing power. In recent years, object detection became a less challenging task with latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide desired performance for RETFE systems due to their complex time-consuming architecture. In addition, small object (<; 30×30 pixels) detection is still a challenging task for existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) that inspired from YOLOv3 to provide above mentioned performance for the smaller object detection. Main contribution is a proposed concatenated block (CB) which has reduced number of convolutional layers and concatenations instead of time-consuming algebraic operations. The superiority of CFPN is confirmed on the COCO and an in-house CarFlow datasets on Nvidia TX2. Thus we conclude that CFPN is useful for real-time embedded smaller object detection task. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2019 | Alignment of Deep Features in 3D Models for Camera Pose Estimation
Jui-Yuan Su, Shyi-Chyi Cheng, Chin-Chun Chang, Jun-Wei Hsieh |
MMM (2) | 4 |
| 2019 | Chronic Kidney Disease Stage Classification Using Renal Artery Doppler-Derived ParametersabstractIn renal medicine, Estimated Glomerular Filtration Rate (eGFR) based method is a standard for the diagnosis of chronic kidney disease. However, this method is invasive, uncomfortable, costly, and could be dangerous because it requires to draw blood from the artery vessels. Researchers have developed several non-invasive Doppler-derived measures based chronic kidney disease (CKD) stage diagnosing or prognosing approaches; however, there is no adequate automatic renal artery Doppler-derived CKD stage classification method in the literature. Thus, we propose a non-invasive, safer, faster, and low cost, SVM-based CKD stage classification method from a sonogram of the renal artery blood flow. The proposed method extracts kurtosis and curvature parameters of the probability distribution that generated from renal artery blood flow waveform. Kurtosis and curvatures are employed to measure the tailedness and curvedness of the probability distribution. We collected a total of 528 sonograms from 110 (49 males) CKD patients during 2010-2013. The experimental results revealed a statistically significant correlation between the parameters and CKD progress stages. Post-voting results revealed the best f1score of 0.956 for Positive (stages 1-5) CKD stages. Munkhjargal Gochoo, Jun-Wei Hsieh, Chien-Hung Lee, Yun-Chih Chen, Yu-Chi Shih |
SMC | 2 |
| 2019 | Novel IoT-Based Privacy-Preserving Yoga Posture Recognition System Using Low-Resolution Infrared Sensors and Deep LearningabstractIn recent years, the number of yoga practitioners has been drastically increased and there are more men and older people practice yoga than ever before. Internet of Things (IoT)-based yoga training system is needed for those who want to practice yoga at home. Some studies have proposed RGB/Kinect camera-based or wearable device-based yoga posture recognition methods with a high accuracy; however, the former has a privacy issue and the latter is impractical in the long-term application. Thus, this paper proposes an IoT-based privacy-preserving yoga posture recognition system employing a deep convolutional neural network (DCNN) and a low-resolution infrared sensor-based wireless sensor network (WSN). The WSN has three nodes (x, y, and z-axes) where each integrates 8 × 8 pixels' thermal sensor module and a Wi-Fi module for connecting the deep learning server. We invited 18 volunteers to perform 26 yoga postures for two sessions each lasted for 20 s. First, recorded sessions are saved as .csv files, then preprocessed and converted to grayscale posture images. Totally, 93200 posture images are employed for the validation of the proposed DCNN models. The tenfold cross-validation results revealed that F1-scores of the models trained with xyz (all 3-axes) and y (only y-axis) posture images were 0.9989 and 0.9854, respectively. An average latency for a single posture image classification on the server was 107 ms. Thus, we conclude that the proposed IoT-based yoga posture recognition system has a great potential in the privacy-preserving yoga training system. Munkhjargal Gochoo, Tan-Hsu Tan, Shih-Chia Huang, Tsedevdorj Batjargal, Jun-Wei Hsieh, Fady Shibata-Alnajjar, Yung-fu Chen |
IEEE Internet Things J. | 5 |
| 2018 | Real-Time Vehicle Re-Identification System Using Symmelets and HOMsabstractA novel vehicle re-identification (VRID) system is proposed to re-identify a vehicle without using features such as license plate, spatial-temporal cues, or 3D information based on only one still image. To detect vehicles from a still image, a symmelet-based approach is derived to determine their ROIs without using any motion feature. A symmelet is a pair of an interest point and its corresponding symmetrical one. This paper modifies the non-symmetrical SURF descriptor into a symmetrical one without adding any time complexity. In order to obtain a set of dense symmelets, a fast interest point extraction method is proposed to detect dense SURF-like points without using a Hessian matrix. After matching with the proposed symmetrical descriptor, the central line of each vehicle can be easily detected from the set of dense symmelets via a projection technique. Then, the desired vehicle ROI can be accurately located along this line. After that, a novel grid-based approach is proposed to re-identify vehicles grid-by-grid by extracting their HOG features for coarse search and refine the final result by using their HOMs (histograms of matching pairs). Without using any GPUs, the VRID system can re-identify the same vehicle very quickly (more than 25 fps) even though a HD-dimensional frame is handled. The accuracy of this VRID system is higher than 94.5% in the FECT dataset and 54.8% in the VeRi-776 dataset . Hung-Chun Chen, Jun-Wei Hsieh, Shiao-Peng Huang |
AVSS | 2 |
| 2018 | Real-Time Vehicle Re-Identification System Using Symmelets and Deep PatchMatch NetsabstractA novel vehicle re-identification (VRID) system is proposed to re-identify a vehicle without using features of license plate, color, and 3D information. To detect vehicles from a still image, a symmelet-based approach is derived to determine their ROIs without using any motion feature. A symmelet is a pair of an interest point and its corresponding symmetrical one. This paper modifies the non-symmetrical SURF descriptor into a symmetrical one without adding any time complexity. In order to obtain a set of dense symmelets, a fast interest point extraction method is proposed to detect dense SURF-like points without using a Hessian matrix. After matching with the proposed symmetrical descriptor and then obtaining a set of dense symmelets, the central line of each vehicle can be easily detected through a projection technique. Then, the desired vehicle ROI can be accurately located along this symmetrical line. After that, a novel patch-based approach is proposed to re-identify vehicles patch-by-patch by extracting their HOG and HOM (histogram of matching pairs) features for coarse search. Then, a novel PatchMatch network is proposed to refine the final result with extreme accuracy. The VRID system can re-identify the same vehicle in real time even though a HD-dimensional frame is handled. The accuracy of this VRID system is higher than 99.3%. Jun-Wei Hsieh, Hung-Chun Chen |
SMC | 1 |
| 2017 | Vehicle Detection in Hsuehshan Tunnel Using Background Subtraction and Deep Belief Network
Bo-Jhen Huang, Jun-Wei Hsieh, Chun-Ming Tsai |
ACIIDS (2) | 2 |
| 2017 | Suspected vehicle detection for driving without license plate using symmelets and edge connectivityabstractThis paper proposes a novel suspected vehicle detection (SVD) system for detecting vehicles moving on roads without a license plate. To detect vehicles from a still image, a symmelet-based approach is derived to determine their ROIs without using any motion feature. A symmelet is a pair of an interest point and its corresponding symmetrical one. This paper modifies the non-symmetrical SURF descriptor into a symmetrical one without adding any time complexity. Then, different symmelets can be very efficiently extracted from road scenes. The set of symmelets can be then used to locate the desired vehicle's ROI using a projection technique. To examine whether a license plate exists within this ROI, an edge connectivity scheme is then proposed to highlight possible character regions for plate detection. This SVD system provides two advantages; there is no need of background subtraction and it is extremely efficient for real-time ITS applications without using any GPU. Jun-Wei Hsieh |
AVSS | 1 |
| 2017 | Air-writing recognition using reverse time ordered stroke contextabstractA novel real-time recognition system is proposed to recognize finger air-writing characters without using any pen-starting-lift information. It presents a novel reverse time ordered stroke context to represent an air-writing trajectory in a backward way so that redundant starting-lift data can be effectively filtered out. Another two challenging problems often happen in the air-writing recognition system, i.e., the multiplicity problem of writing and the confusion problem. The first one means a character is always written differently and the second one means different various characters own similar writing trajectory. To tackle them, a three-layer hierarchical structure to represent an air-writing character with different sampling rates is proposed. All the alphabets (including lowercase, capital, and digital letters) are recognized in this system. Performance evaluation shows that the proposed solution achieves quite higher recognition accuracy (more than 94.7%) even though no starting gesture is required. Tsung-Hsien Tsai, Jun-Wei Hsieh |
ICIP | 2 |
| 2017 | Model-Based 3D Scene Reconstruction Using a Moving RGB-D Camera
Shyi-Chyi Cheng, Jui-Yuan Su, Jing-Min Chen, Jun-Wei Hsieh |
MMM (1) | 4 |
| 2017 | Video action classification using symmelets and deep learningabstractClassification of human actions is very challenging and important in many video-based applications. Two common features, i.e., the hand-crafted and the deep-learned ones are usually adopted for video representation and have been proven to be effective in many famous datasets in the literature. However, the hand-crafted feature lacks the ability to detect the discriminative and semantic features and the deep-learned one fails to outperform previous hand-crafted feature. This paper propose a novel symmelet-based classification approach to improve the accuracy of the state-of-the-art frameworks. "Symmelet" is a symmetrical pair including a SURF point and its corresponding symmetrical point in the same frame. Many symmetrical properties often exist in various video scenes. With symmelets, various redundant (or background) features can be filtered out so that action contents can be more accurately represented. The new approach takes advantages of symmelets, improved dense trajectories (IDT), and trajectory-pooled deep-convolutional descriptor (TDD) to learn useful deep features for represent video contents deeply. Performance evaluation on two challenging datasets, i.e., HMDB51 and UCF101 shows that the proposed solution is superior and can achieve quite higher recognition accuracy than other state-of-art frameworks. Salah Alghyaline, Jun-Wei Hsieh, Chi-Hung Chuang |
SMC | 2 |
| 2016 | Visual location search using symmeletsabstractSURF is a robust and useful feature detector to various vision-based applications but lacks of the ability to detect symmetric objects. This paper proposes a new symmetrical SURF descriptor to detect all possible symmetric pairs via a mirroring transformation. With this symmetrical descriptor, a novel feature named “symmelet” is introduced and used in scene representation and effective mobile visual location search. A symmelet is a symmetrical pair formed by a SURF point and its symmetrical one. Three advantages can be gained from this symmelet-absed representation. Firstly, because the set of symmelets is small, a scene can be represented more compactly and searched more efficiently. Secondly, its symmetrical property can compare image/scene contents more accurately. Thirdly, the geometric structure of a scene can be easily constructed and verified and thus filter out many false matches. Then, given a query image captured by a mobile phone, the descried location can be very efficiently and accurately retrieved even though this phone is with quite limited computational power. Chong-Po Liao, Jun-Wei Hsieh, Hui-Fen Chiang, Yun Tsao |
ICIP | 2 |
| 2016 | Moment-based symmetry detection for scene modeling and recognition using RGB-D imagesabstractIn this paper we present a novel unsupervised feature representation by extracting salient symmetries in RGB-D images using the proposed moment-based symmetric patch detector. A fast indexing structure is also derived to group local symmetric patches into semantically meaningful symmetric parts. Given an RGB-D image, the hash-based symmetric patch indexing speeds up the searches of symmetric patch pairs, which are further grouped into symmetric parts with nearly linear time complexity. In the context of symmetry matching and scene classification, the second part of this work presents a symmetry-based scene modeling, aiming at computing a robust part-based feature set for each image category. To verify the effectiveness of the symmetry detector, based on the pre-learned part-based scene model, a part-based voting scheme is constructed to annotate the scene type of the input RGB-D image. Experimental results show that the proposed approach outperforms the compared methods in terms of detection and recognition accuracy using publicly available datasets. Jui-Yuan Su, Shyi-Chyi Cheng, Jun-Wei Hsieh, Tzu-Hao Hsu |
ICPR | 3 |
| 2016 | Action classification using data mining and Paris of SURF-based trajectoriesabstractA new action classification approach is proposed to improve the accuracy of the state-of-art frameworks from three folds: (1) Association rule mining is used with dense trajectories approach to discover strong relations between different visual words in the video clips, then a new histogram is built for each video clip based on such relations. (2) The second proposed approach is based on SURF descriptor to extract the most similar pairs of dense trajectories' features, and then the most similar trajectories' features are used to describe the video clip. (3) Finally, a symmetrical SURFs approach is used to detect the symmetrical pairs of trajectories in the video; the most symmetrical features in the video clip are extracted and used to describe the video clip. The above three new features are used in addition to the original dense trajectories' features for action classification. The importance of these new features is that many features are not related to the background and can significantly increase the overall recognition accuracy. Salah Alghyaline, Jun-Wei Hsieh, Hui-Fen Chiang, Rui-Yu Lin |
SMC | 2 |
| 2015 | PLSA-based sparse representation for vehicle color classificationabstractThis paper proposes a novel vehicle color classification method which uses the concept of probabilistic latent semantic analysis (pLSA) to overcome the problem of sparse representation in data classification. Sparse representation is widely used and quite successful in many vision-based applications. However, it needs to calculate the sparse reconstruction cost (SRC) of each sample to find the best candidate. Because an optimization process is involved, it is very inefficient. In addition, it uses only the residual and does not consider the arrangement (or distribution) of combination coefficients of visual codes in classification. Thus, it often fails to classify categories if they are similar. In this paper, the pLSA concept is first introduced into the sparse representation to build a new classifier without using the SRC measure. The weakness of the pLSA scheme is the use of EM algorithm for updating the posteriori probability of latent class. Because it is very time-consuming, a novel weighting voting strategy is introduced to improve the pLSA scheme for recognizing objects in real time. The advantages of this classifier are: the accuracy is much higher than the SRC scheme and the efficiency is real-time in data classification. Vehicle color classification is demonstrated in this paper to prove the superiority of the new classifier. Ssu-Ying Wang, Jun-Wei Hsieh, Yilin Yan, Li-Chih Chen, Duan-Yu Chen |
AVSS | 2 |
| 2015 | Real-time vehicle color identification using symmetrical SURFs and chromatic strengthabstractThis paper proposes a new vehicle color classification scheme to identify vehicles with their colors. To detect vehicles from roads, the paper proposes a novel symmetrical descriptor to determine the ROI of each vehicle without using any motion features. This scheme provides two advantages; there is no need of background subtraction and it is extremely efficient for real-time applications. After detection, a novel color-correction technique is proposed to reduce the color changes of vehicles so that vehicles can be more accurately identified. The major challenge in vehicle color identification is there are many shade (or confused) colors among vehicles. This paper proposes a new concept that the vehicles with different chromatic attributes should be separately trained even though they are in the same color category. With this concept, a novel tree-based classifier can be constructed to classify vehicles at different stages according to their chromatic strengths. The separation can significantly improve the accuracy of vehicle color classification even that vehicles are with various shade colors. Li-Chih Chen, Jun-Wei Hsieh, Hui-Fen Chiang, Tsung-Hsien Tsai |
ISCAS | 2 |
| 2015 | Modeling and recognizing action contexts in persons using sparse representation
Hui-Fen Chiang, Jun-Wei Hsieh, Chi-Hung Chuang, Kai-Ting Chuang, Yilin Yan |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Human movement analysis around a view circle using time-order similarity distributions
Chi-Hung Chuang, Jun-Wei Hsieh, Hui-Fen Chiang, Yi-Da Chiou |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Vehicle make and model recognition using sparse representation and symmetrical SURFs
Li-Chih Chen, Jun-Wei Hsieh, Yilin Yan, Duan-Yu Chen |
Pattern Recognit. | 2 |
| 2014 | Vehicle licence plate recognition using super-resolution techniqueabstractDue to the development of economy and technology, the people's demands on cars are growing and so are the problems, such as finding stolen car, banning violation and parking lot management. It will be time-consuming and low-efficiency if we do those jobs only by human because of the limitation of human being's concentration. Therefore, it has been a popular topic to develop intelligent monitoring system under new video technology within a decade. There are still many rooms for future development and application in the car license detection and recognition fields. For finding stolen car, we can integrate car license detection system and road monitoring system to analyze the videos and trace the objects, so we can gain high-efficiency and low-cost results. For parking lot management system, we can have the result of access management and automatic charge to reduce the human resource through car license recognition system. Automated toll of highway can be done through car license recognition system as well. Because the car license recognition system is mostly applied to security monitoring field and business purposes, the demand of the accuracy is quite strict. There are causes which make the inaccuracy of the license recognition, such as the lack of video resolution, too small license due to the distance and the light and shadow. We will discuss the license recognition system and how to use apply the super-resolution to overcome those above problems. Chi-Hung Chuang, Luo-Wei Tsai, Ming-Shan Deng, Jun-Wei Hsieh, Kuo-Chin Fan |
AVSS | 4 |
| 2014 | PLSA-Based Sparse Representation for Object ClassificationabstractThis paper proposes a novel object classification method which uses the concept of probabilistic latent semantic analysis (pLSA) to overcome the problem of sparse representation in data classification. Sparse representation is widely used and quite successful in many vision-based applications. However, it needs to calculate the sparse reconstruction cost (SRC) of each sample to find the best candidate. Because an optimization process is involved, it is very inefficient. In addition, it uses only the residual and does not consider the arrangement (or distribution) of combination coefficients of visual codes in classification. Thus, it often fails to classify categories if they are similar. In this paper, the pLSA concept is first introduced into the sparse representation to build a new classifier without using the SRC measure. The weakness of the pLSA scheme is the use of EM algorithm for updating the posteriori probability of latent class. Because it is very time-consuming, a novel weighting voting strategy is introduced to improve the pLSA scheme for recognizing objects in real time. The advantages of this classifier are: the accuracy is much higher than the SRC scheme and the efficiency is real-time in data classification. Two applications are demonstrated in this paper to prove the superiority of the new classifier, i.e., vehicle make and model recognition, and action analysis. Yilin Yan, Jun-Wei Hsieh, Hui-Fen Chiang, Shyi-Chyi Cheng, Duan-Yu Chen |
ICPR | 2 |
| 2014 | Handheld object detection and its related event analysis using ratio histogram and mixture of HMMs
Jun-Wei Hsieh, Jiun-Cheng Cheng, Li-Chih Chen, Chi-Hung Chuang, Duan-Yu Chen |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Symmetrical SURF and Its Applications to Vehicle Detection and Vehicle Make and Model RecognitionabstractSpeeded-Up Robust Features (SURF) is a robust and useful feature detector for various vision-based applications but it is unable to detect symmetrical objects. This paper proposes a new symmetrical SURF descriptor to enrich the power of SURF to detect all possible symmetrical matching pairs through a mirroring transformation. A vehicle make and model recognition (MMR) application is then adopted to prove the practicability and feasibility of the method. To detect vehicles from the road, the proposed symmetrical descriptor is first applied to determine the region of interest of each vehicle from the road without using any motion features. This scheme provides two advantages: there is no need for background subtraction and it is extremely efficient for real-time applications. Two MMR challenges, namely multiplicity and ambiguity problems, are then addressed. The multiplicity problem stems from one vehicle model often having different model shapes on the road. The ambiguity problem results from vehicles from different companies often sharing similar shapes. To address these two problems, a grid division scheme is proposed to separate a vehicle into several grids; different weak classifiers that are trained on these grids are then integrated to build a strong ensemble classifier. The histogram of gradient and SURF descriptors are adopted to train the weak classifiers through a support vector machine learning algorithm. Because of the rich representation power of the grid-based method and the high accuracy of vehicle detection, the ensemble classifier can accurately recognize each vehicle. Jun-Wei Hsieh, Li-Chih Chen, Duan-Yu Chen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Vehicle make and model recognition using symmetrical SURFabstractSURF (Speeded Up Robust Features) is a robust and useful feature detector for various vision-based applications but lacks the ability to detect symmetrical objects. This paper proposes a new symmetrical SURF descriptor to enrich the power of SURF to detect all possible symmetrical matching pairs through a mirroring transformation. A vehicle make-and-model recognition (MMR) application is then adopted to prove the practicability and feasibility of the method. To detect vehicles from the road, the proposed symmetrical descriptor is first applied to determine the ROI of each vehicle from the road without using any motion features. This scheme provides two advantages; there is no need of background subtraction and it is extremely efficient for real-time applications. Two MMR challenges, i.e., multiplicity and ambiguity problems, are then addressed. The multiplicity problem stems from one vehicle model often having different model shapes on the road. The ambiguity problem results from vehicles from different companies often sharing similar shapes. To address these two problems, a grid division scheme is proposed to separate a vehicle into several grids; different weak classifiers that are trained on these grids are then integrated to build a strong ensemble classifier. Because of the rich representation power of the grid-based method and the high accuracy of vehicle detection, the ensemble classifier can accurately recognize each vehicle. Jun-Wei Hsieh, Li-Chih Chen, Duan-Yu Chen, Shyi-Chyi Cheng |
AVSS | 1 |
| 2012 | Occluded human action analysis using dynamic manifold model
Li-Chih Chen, Jun-Wei Hsieh, Chi-Hung Chuang, Chang-Yu Huang, Duan-Yu Chen |
ICPR | 2 |
| 2012 | Vehicle color classification under different lighting conditions through color correctionabstractThis paper presents a novel color correction technique for classifying vehicles under different lighting conditions using their colors. To reduce the lighting effects, a reference image is first selected for building the mapping function between the current frame and the reference image. With this mapping function, the color distortions between frames can be reduced to minimum. In addition to lighting changes, the effect of sun light will make the vehicle window become white and lead to the errors of vehicle classification. To reduce this effect, a window-removing task is then applied for making vehicle pixels with the same color more concentrated on the foreground region. Then, vehicles can be more accurately classified to their categories even though strong sun light casts on them. To tackle the confusion problem that some vehicle colors are too similar, e.g., “deep-blue” and “deepgreen”, a novel tree-based classifier is then designed for classifying vehicles to more detailed labels. Experimental results have proved that the proposed method is a robust, accurate, and powerful tool for vehicle classification. Jun-Wei Hsieh, Li-Chih Chen, Sin-Yu Chen, Duan-Yu Chen |
ISCAS | 1 |
| 2012 | Template Matching and Monte Carlo Markova Chain for People Counting under Occlusions
Jun-Wei Hsieh, Fu-Jiang Fang, Guo-Jin Lin, Yu-Shi Wang |
MMM | 1 |
| 2010 | Human Smoking Event Detection Using Visual Interaction CluesabstractThis paper presents a novel scheme to automatically and directly detect smoking events in video. In this scheme, a color-based ratio histogram analysis is introduced to extract the visual clues from appearance interactions between lighted cigarette and its human holder. The techniques of color re-projection and Gaussian Mixture Models (GMMs) enable the tasks of cigarette segmentation and tracking over the background pixels. Then, a key problem for event analysis is the non-regular form of smoking events. Thus, we propose a self-determined mechanism to analyze this suspicious event using HHM framework. Due to the uncertainties of cigarette size and color, there is no automatic system which can well analyze human smoking events directly from videos. The proposed scheme is compatible to detect the smoking events of uncertain actions with various cigarette sizes, colors, and shapes, and has capacity to extend visual analysis to human events of similar interaction relationship. Experimental results show the effectiveness and real-time performances of our scheme in smoking event analysis. Pin Wu, Jun-Wei Hsieh, Jiun-Cheng Cheng, Shyi-Chyi Cheng, Shau-Yin Tseng |
ICPR | 2 |
| 2010 | Human behavior recognition from arbitrary viewsabstractThis paper presents a new behavior classification system that can analyze human behaviors from arbitrary views. Technically, if different viewing angle are used for observing a person, his appearances will change significantly. To freely recognize his behaviors, traditional methods tend to adopt 3-D data for behavior analysis. However, its inherent correspondence process will make it inappropriate for real time applications. To tackle this problem, a novel view alignment method is first proposed for mapping each action sequence to a fixed view. To achieve this mapping, two features extracted from spatial and temporal domains are used for representing each action sequence. For the spatial feature, the "centroid context" of each posture is defined and extracted through a triangulation technique. For the temporal feature, the "posture transition probability" is constructed for recording the probabilities of one posture type transferring to another one. After mapping, a novel matrix representation is proposed for describing each action more accurately. After that, the Viterbi algorithm is used for aligning two action sequences and then classifying them to different behavior types. Chi-Hung Chuang, Jun-Wei Hsieh, Yi-Da Chiou, I-Ru Tsay, Ming-Hui Jin |
ISCAS | 2 |
| 2010 | Occluded human body segmentation and its application to behavior analysisabstractThis paper addresses the problem of occluded human segmentation and then uses its results for human behavior recognition. To tackle this ill-posed problem, a novel clustering scheme is proposed for constructing a model space for posture classification. Then, a model-driven approach is proposed for separating an occluded region to individual objects. For reducing the model space, a particle filtering technique is then used for locating possible positions of each occluded object. Then, from the positions, the best model of each occluded object can be then selected using its distance maps. Then, a novel template re-projection technique is proposed for repairing an occluded object to a complete one. Due to occlusions, there will be many posture symbol converting errors in this representation. Instead of using a specific symbol, we code a posture using not only its best matched key posture but also its similarities among other key postures. With the matrix representation, different actions can be more robustly and effectively matched by comparing their KL distance. Jun-Wei Hsieh, Sin-Yu Chen, Chi-Hung Chuang, Miao-Fen Chueh, Shiaw-Shian Yu |
ISCAS | 1 |
| 2010 | Vehicle Orientation Analysis Using Eigen Color, Edge Map, and Normalized Cut ClusteringabstractThis paper proposes a novel approach for estimating vehicles' orientations from still images using "eigen color" and edge map through a clustering framework. To extract the eigen color, a novel color transform model is used for roughly segmenting a vehicle from its background. The model is invariant to various situations like contrast changes, background, and lighting. It does not need to be re-estimated for any new vehicles. In this eigen color space, different vehicle regions can be easily identified. However, since the problem of object segmentation is still ill-posed, only with this model, the shape of a vehicle cannot be well extracted from its background and thus affects the accuracy of orientation estimation. In order to solve this problem, the distributions of vehicle edges and colors are then integrated together to form a powerful but high-dimensional feature space. Since the feature dimension is high, the normalized cut spectral clustering (Ncut) is then used for feature reduction and orientation clustering. The criterion in Ncut tries to minimize the ratio of the total dissimilarity between groups to the total similarity within the groups. Then, the vehicle orientation can be analyzed using the eigenvectors derived from the Ncut result. The proposed framework needs only one still image and is thus very different to traditional methods which need motion features to determine vehicle orientations. Experimental results reveal the superior performances in vehicle orientation analysis. Jui-Chen Wu, Jun-Wei Hsieh, Sin-Yu Chen, Cheng-Min Tu, Yung-Sheng Chen |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2010 | Segmentation of Human Body Parts Using Deformable TriangulationabstractThis paper presents a novel segmentation algorithm to segment a body posture into different body parts using the technique of deformable triangulation. To analyze each posture more accurately, they are segmented into triangular meshes, where a spanning tree can be found from the meshes using a depth-first search scheme. Then, we can decompose the tree into different subsegments, where each subsegment can be considered as a limb. Then, two hybrid methods (i.e., the skeleton-based and model-driven methods) are proposed for segmenting the posture into different body parts according to its occlusion conditions. To analyze occlusion conditions, a novel clustering scheme is proposed to cluster the training samples into a set of key postures. Then, a model space can be used to classify and segment each posture. If the input posture belongs to the nonocclusion category, the skeleton-based method is used to divide it into different body parts that can be refined using a set of Gaussian mixture models (GMMs). For the occlusion case, we propose a model-driven technique to select a good reference model for guiding the process of body part segmentation. However, if two postures' contours are similar, there will be some ambiguity that can lead to failure during the model selection process. Thus, this paper proposes a tree structure that uses a tracking technique so that the best model can be selected not only from the current frame but also from its previous frame. Then, a suitable GMM-based segmentation scheme can be used to finely segment a body posture into the different body parts. The experimental results show that the proposed method for body part segmentation is robust, accurate, and powerful. Jun-Wei Hsieh, Chi-Hung Chuang, Sin-Yu Chen, Chih-Chiang Chen, Kuo-Chin Fan |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2009 | A fast cube-based video shot retrieval using 3D moment-preserving techniqueabstractIn this paper, we describe a novel video shot retrieval where each shot is separated into multiple video cubes. Therefore, every video shot can be represented by a linear combination of video cubes, which are used to calculate the similarity measurements among video shots in terms of video cube similarity. The position of a voxel in a cube is characterized with (x, y, t) 3D coordinates and the spatial-temporal features within video cubes are extracted with a set of analytical formulas derived from the proposed 3D moment-preserving technique. Then, the content of a video cube is approximated by three blocks generated from projecting the cube onto xy, yt and tx planes. Based on the visual patterns of xy, yt, and tx blocks, a fast video shot retrieval scheme is proposed. As compared with other key-frame based representations, the proposed cube-based video retrieval improves the retrieval accuracy without sacrificing the execution speed. Experimental results show the efficiency and effectiveness of the proposed video retrieval. Wei-Kan Huang, Chi-Han Chuang, Shyi-Chyi Cheng, Jun-Wei Hsieh |
ICIP | 4 |
| 2009 | Abnormal Event Analysis Using Patching Matching and Concentric Features
Jun-Wei Hsieh, Sin-Yu Chen, Chao-Hong Chiang |
KES (2) | 1 |
| 2009 | Carried Object Detection Using Ratio Histogram and its Application to Suspicious Event AnalysisabstractThis letter proposes a novel method to detect carried objects from videos and applies it for analysis of suspicious events. First of all, we propose a novel kernel-based tracking method for tracking each foreground object and further obtaining its trajectory. With the trajectory, a novel ratio histogram is then proposed for analyzing the interactions between the carried object and its owner. After color re-projection, different carried objects can be then accurately segmented from the background by taking advantages of Gaussian mixture models. After bag detection, an event analyzer is then designed to analyze various suspicious events from the videos. Even though there is no prior knowledge about the bag (such as shape or color), our proposed method still performs well to detect these suspicious events. As we know, due to the uncertainties of the shape and color of the bag, there is no automatic system that can analyze various suspicious events involving bags (such as robbery) without using any manual effort. However, by taking advantages of our proposed ratio histogram, different carried bags can be well segmented from videos and applied for event analysis. Experimental results have proved that the proposed method is robust, accurate, and powerful in carried object detection and suspicious event analysis. Chi-Hung Chuang, Jun-Wei Hsieh, Luo-Wei Tsai, Sin-Yu Chen, Kuo-Chin Fan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Efficient image matching using concentric sampling features and boosting processabstractThis paper presents a novel template matching method to efficiently match and search image patterns. The method using concentric sampling structures, boosting process, and coarse-to-fine framework, differs from the traditional pattern matching schemes of time-exhausting correlation. The time complexity at searching stage is invariant to the dimension of concerned patterns. The rotation-invariant collection of concentric sub-samples represents as a reliable relaxation process of weak beliefs to efficiently reject the impossible location candidates. The concentric sampling approximation of integral images and the hierarchical scheme enable sifting out the patterns to process with the reduced complexity. Experimental result demonstrates the real-time performance on efficient pattern detection and geometry parameter estimation and the flexibility (on translation-, scaling-, and rotation-variant patterns) for various image analysis applications. Pin Wu, Jun-Wei Hsieh |
ICIP | 2 |
| 2008 | Suspicious object detection using fuzzy-color histogramabstractThis paper proposes a novel method to detect suspicious objects from videos for abnormal event analysis. When considering a robbery event happens, there should be some suspicious object transferring conditions following between the forager and the victim. Since there is no prior knowledge about the object’s property, it is difficult to automatically analyze the conditions without any manual efforts. To tackle this problem, a ratio histogram based on fuzzy c-means algorithm is proposed for finding suspicious objects. Furthermore, we use Gaussian mixture models to model the suspicious object’s visual properties so that it can be accurately segmented from videos. After analyzing its subsequent motion features, different abnormal events like robbery can be effectively detected from videos. Experiment results have proved that the proposed method is robust, accurate, and powerful in abnormal event detection. Chi-Hung Chuang, Jun-Wei Hsieh, Luo-Wei Tsai, Pei-Shiuan Ju, Kuo-Chin Fan |
ISCAS | 2 |
| 2008 | Morphology-based text line extraction
Jui-Chen Wu, Jun-Wei Hsieh, Yung-Sheng Chen |
Mach. Vis. Appl. | 2 |
| 2008 | Boosted string representation and its application to video surveillance
Jun-Wei Hsieh, Yung-Tai Hsu |
Pattern Recognit. | 1 |
| 2008 | Video-Based Human Movement Analysis and Its Application to Surveillance SystemsabstractThis paper presents a novel posture classification system that analyzes human movements directly from video sequences. In the system, each sequence of movements is converted into a posture sequence. To better characterize a posture in a sequence, we triangulate it into triangular meshes, from which we extract two features: the skeleton feature and the centroid context feature. The first feature is used as a coarse representation of the subject, while the second is used to derive a finer description. We adopt a depth-first search (dfs) scheme to extract the skeletal features of a posture from the triangulation result. The proposed skeleton feature extraction scheme is more robust and efficient than conventional silhouette-based approaches. The skeletal features extracted in the first stage are used to extract the centroid context feature, which is a finer representation that can characterize the shape of a whole body or body parts. The two descriptors working together make human movement analysis a very efficient and accurate process because they generate a set of key postures from a movement sequence. The ordered key posture sequence is represented by a symbol string. Matching two arbitrary action sequences then becomes a symbol string matching problem. Our experiment results demonstrate that the proposed method is a robust, accurate, and powerful tool for human movement analysis. Jun-Wei Hsieh, Yung-Tai Hsu, Hong-Yuan Mark Liao, Chih-Chiang Chen |
IEEE Trans. Multim. | 1 |
| 2007 | Road Sign Detection Using Eigen Color
Luo-Wei Tsai, Yun-Jung Tseng, Jun-Wei Hsieh, Kuo-Chin Fan, Jiun-Jie Li |
ACCV (1) | 3 |
| 2007 | Suspicious Object Detection and Robbery Event AnalysisabstractThis paper proposes a novel method to detect suspicious objects from videos for robbery event analysis. First of all, a background subtraction using a minimum filter is used for detecting foreground objects from videos. Then, a novel kernel-based tracking method is proposed for tracking each moving object and obtaining its trajectory. Then, we propose a novel robbery event analysis system to analyze suspicious object transferring conditions between any two persons. Usually, when a robbery event happens, there should some suspicious object transferring conditions happening between the robbery and the victim. Since there is no prior knowledge about the object's property, it is difficult to automatically analyze the conditions without any manual efforts. To tackle this problem, a novel ratio histogram is then proposed for finding suspicious objects and then accurately analyzing their transferring conditions. After color re-projection, we use Gaussian mixture models to model the suspicious object's visual properties so that it can be very accurately segmented from videos. After analyzing its subsequent speed, different robbery events can be then effectively detected from videos. Experiment results have proved that the proposed method is robust, accurate, and powerful in robbery event detection. Chi-Hung Chuang, Jun-Wei Hsieh, Kuo-Chin Fan |
ICCCN | 2 |
| 2007 | Grid-based Template Matching for People CountingabstractThis paper presents a novel template matching method to detect and track pedestrians for people counting in real-time. Firstly, a novel background subtraction method is proposed for extracting all foreground objects from background. Then, a shadow elimination method is used to remove unwanted shadow from the background. In order to identify pedestrians from non-pedestrian objects, this paper proposed a novel grid-based template matching scheme to robustly verify each pedestrian. Usually, a pedestrian will have different appearances at different positions. The grid-based approach can effectively reduce the perspective effects into a minimum since it uses different templates to record the appearance changes at each grid. When more templates are used, the detection process will become more inefficient. To speed up its efficiency, an integral image is used to filter out all impossible candidates in advance. Lastly, a tracking method is applied to tracking the direction of each moving pedestrian so that the real number of passing people per direction can be counted more accurately. Experimental results have proved that the proposed method is robust, accurate, and powerful in people counting. Jun-Wei Hsieh, Cheng-Shuang Peng, Kuo-Chin Fan |
MMSP | 1 |
| 2007 | Vehicle Detection Using Normalized Color and Edge MapabstractThis paper presents a novel vehicle detection approach for detecting vehicles from static images using color and edges. Different from traditional methods, which use motion features to detect vehicles, this method introduces a new color transform model to find important "vehicle color" for quickly locating possible vehicle candidates. Since vehicles have various colors under different weather and lighting conditions, seldom works were proposed for the detection of vehicles using colors. The proposed new color transform model has excellent capabilities to identify vehicle pixels from background, even though the pixels are lighted under varying illuminations. After finding possible vehicle candidates, three important features, including corners, edge maps, and coefficients of wavelet transforms, are used for constructing a cascade multichannel classifier. According to this classifier, an effective scanning can be performed to verify all possible candidates quickly. The scanning process can be quickly achieved because most background pixels are eliminated in advance by the color feature. Experimental results show that the integration of global color features and local edge features is powerful in the detection of vehicles. The average accuracy rate of vehicle detection is 94.9%. Luo-Wei Tsai, Jun-Wei Hsieh, Kuo-Chin Fan |
IEEE Trans. Image Process. | 2 |
| 2006 | Video Object Segmentation Using Kernel-based Models and Spatiotemporal SimilarityabstractThis paper proposes a semantic video object segmentation system which combines spatio-temporal video segmentation and region tracking together to extract important semantic objects from videos. At beginning, the paper uses multiple cues to segment video frames to different regions. The cues include color, edges, motions, and kernel-based models. Since these features are complementary to each other, all desired regions can be well segmented from input frames even though they are captured from a non-stationary camera. Then, according to temporal information of each segmented region, we can construct a region adjacency graph (RAG) which can well record the relative relations between each region. Based on the RAG, we propose a Bayesian classifier which can group regions by properly checking their spatial and temporal similarities such that different regions will be merged and associated together to form a meaningful object. Since a kernel-based analysis is included into the designed classifier, all desired semantic objects can be well extracted even though they are static in videos. Experimental results have proved the superiority of the proposed method in object segmentation. Jun-Wei Hsieh, Jun-Xian Lee |
ICIP | 1 |
| 2006 | Building a Remote Supervisory Control Network System for Smart Home ApplicationsabstractWireless sensor networks are often found in the fields of home security, industrial control and maintenance, medical assistance and traffic monitoring and the appearance of ZigBee/IEEE 802.15.4 indicates a network system which is highly reliable, cost-effective, low power consumption, programmable and fast establishing. Currently, many of the wireless sensor network systems are now using ZigBee to implement the designs. A smart sensor network is the infrastructure of home automation and supervisory control systems, as the proposed functions such as intelligence entrance guards management, home security, environmental monitor and light control can be implemented by a smart network integrating with multimedia access service, image processing, security, and sensor and control technologies. Details of the development process of a smart home network in Taiwan based on ZigBee technology with the combination of smart home appliance communication protocol, SAANet, will be given in this article. Yu-Ping Tsou, Jun-Wei Hsieh, Cheng-Ting Lin |
SMC | 2 |
| 2006 | Motion-based video retrieval by trajectory matchingabstractThis paper proposes a hybrid motion-based video retrieval system to retrieve desired videos from video databases through trajectory matching. The hybrid method includes a sketch-based scheme and a string-based one to analyze and index a trajectory with more syntactic meanings. First of all, this method uses a sampling technique to extract a set of control points from each trajectory as features. Then, the sketch-based method uses a curve fitting technique to interpolate some missed data in this set of control points. Then, the visual distance between any two trajectories can be directly measured by comparing their position data. The visual distance is good in solving the problem of translation-invariant trajectory matching but poor in solving the problem of partial trajectory matching. Therefore, in addition to the visual distance, the hybrid method uses the string-based scheme to compare any two trajectories according to their syntactic meanings. With the help of the syntactic distance, many impossible candidates can be filtered out in advance and thus the accuracy of video retrieval can be much enhanced. In addition, the problem of partial trajectory matching will become easy to be solved. Thus, even though a partial trajectory is queried, all desired video clips still can be very accurately retrieved. Experimental results have proved the superiority of our proposed method. Jun-Wei Hsieh, Shang-Li Yu, Yung-Sheng Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Automatic traffic surveillance system for vehicle tracking and classificationabstractThis paper presents an automatic traffic surveillance system to estimate important traffic parameters from video sequences using only one camera. Different from traditional methods that can classify vehicles to only cars and noncars, the proposed method has a good ability to categorize vehicles into more specific classes by introducing a new "linearity" feature in vehicle representation. In addition, the proposed system can well tackle the problem of vehicle occlusions caused by shadows, which often lead to the failure of further vehicle counting and classification. This problem is solved by a novel line-based shadow algorithm that uses a set of lines to eliminate all unwanted shadows. The used lines are devised from the information of lane-dividing lines. Therefore, an automatic scheme to detect lane-dividing lines is also proposed. The found lane-dividing lines can also provide important information for feature normalization, which can make the vehicle size more invariant, and thus much enhance the accuracy of vehicle classification. Once all features are extracted, an optimal classifier is then designed to robustly categorize vehicles into different classes. When recognizing a vehicle, the designed classifier can collect different evidences from its trajectories and the database to make an optimal decision for vehicle classification. Since more evidences are used, more robustness of classification can be achieved. Experimental results show that the proposed method is more robust, accurate, and powerful than other traditional methods, which utilize only the vehicle size and a single frame for vehicle classification. Jun-Wei Hsieh, Shih-Hao Yu, Yung-Sheng Chen, Wen-Fong Hu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2005 | A Fast Method to Detect and Recognize Scaled and Skewed Road Signs
Yi-Sheng Liu 0001, Der-Jyh Duh, Shu-Yuan Chen, Jun-Wei Hsieh |
ACIVS | 4 |
| 2005 | Vehicle detection using normalized color and edge mapabstractThis paper presents a novel vehicle detection approach for detecting vehicles from static images using color and edges. Different from traditional methods which use motion features to detect vehicles, this method introduces a new color transform model to find important "vehicle color" from images for quickly locating possible vehicle candidates. Since vehicles have different colors under different lighting conditions, there were seldom works proposed for detecting vehicles using colors. This paper proves that the new color transform model has extreme abilities to identify vehicle pixels from backgrounds even though they are lighted under various illumination conditions. Each detected pixel corresponds to a possible vehicle candidate. Then, two important features including edge maps and coefficients of wavelet transform are used for constructing a multi-channel classifier to verify this candidate. According to this classifier, we can perform an effective scan to detect all desired vehicles from static images. Since the color feature is first used to filter out most background pixels, this scan can be extremely quickly achieved. Experimental results show that the integrated scheme is very powerful in detecting vehicles from static images. The average accuracy of vehicle detection is 94.5%. Luo-Wei Tsai, Jun-Wei Hsieh, Kuo-Chin Fan |
ICIP (2) | 2 |
| 2005 | Human Behavior Analysis Using Deformable TriangulationsabstractThis paper presents a new posture classification system to analyze different human behaviors directly from video sequences using the technique of triangulation. For well analyzing each posture in the video sequences, we propose a triangulation-based method to triangulate it to different triangle meshes from which two important posture features are then extracted, i.e., the ones of skeleton and centroid context. The first one is used for a coarse search and the second one is for a finer classification to classify postures in more details. For the first descriptor, we take advantages of a dfs (depth-first search) scheme to extract the skeleton features of a posture from its triangulation result. Then, with the help of skeleton information, we can define a new shape descriptor, i.e., centroid context, to describe a posture up to a semantic level. That is, the centroid context is a finer descriptor to describe a posture not only from its whole shape but also from its body parts. Since the two descriptors are complement to each other, all desired human postures can be compared and classified very accurately. The nice ability of posture classification can help us generate a set of key postures for transferring a behavior sequence to a set of symbols. Then, a novel string matching scheme is proposed to analyze different human behaviors. Experimental results have proved that the proposed method is robust, accurate, and powerful in human behavior analysis Yung-Tai Hsu, Jun-Wei Hsieh, Hai-Feng Kao, Hong-Yuan Mark Liao |
MMSP | 2 |
| 2004 | Novel aircraft type recognition with learning capabilities in satellite images
Jun-Wei Hsieh, Jian-Ming Chen, Chi-Hung Chuang, Kuo-Chin Fan |
ICIP | 1 |
| 2004 | Trajectory-based video retrieval by string matchingabstractThis paper proposes a trajectory-based video retrieval system to retrieve desired videos from a video database through string matching. First, in order to represent each trajectory, a hybrid technique is proposed for representing its semantic meanings and geometrical properties. At this scheme, we use the Bezier basis functions to interpolate some lost control points of each represented trajectory. Then, the distance between any two trajectories can be measured by comparing the positions of sampling points extracted along their Bezier approximations. In addition, this hybrid method uses a novel labeling technique for converting a trajectory into a string. This string representation can give more semantic information in interpreting a trajectory and make important improvements in video classification. More importantly, the problem of partial matching will become easy and can be efficiently solved by a string matching technique. Experimental results have proved the superiority of our proposed method. Jun-Wei Hsieh, Shang-Li Yu, Yung-Sheng Chen |
ICIP | 1 |
| 2004 | Tracking Multiple Moving Objects Using A Level-Set MethodabstractThis paper presents a novel approach to track multiple moving objects using the level-set method. The proposed method can track different objects no matter if they are rigid, nonrigid, merged, split, with shadows, or without shadows. At the first stage, the paper proposes an edge-based camera compensation technique for dealing with the problem of object tracking when the background is not static. Then, after camera compensation, different moving pixels can be easily extracted through a subtraction technique. Thus, a speed function with three ingredients, i.e. pixel motions, object variances and background variances, can be accordingly defined for guiding the process of object boundary detection. According to the defined speed function, different object boundaries can be efficiently detected and tracked by a curve evolution technique, i.e. the level-set-based method. Once desired objects have been extracted, in order to further understand the video content, this paper takes advantage of a relation table to identify and observe different behaviors of tracked objects. However, the above analysis sometimes fails due to the existence of shadows. To avoid this problem, this paper adopts a technique of Gaussian shadow modeling to remove all unwanted shadows. Experimental results show that the proposed method is much more robust and powerful than other traditional methods. Jun-Wei Hsieh, Yung-Sheng Chen, Wen-Fong Hu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | Fast stitching algorithm for moving object detection and mosaic construction
Jun-Wei Hsieh |
Image Vis. Comput. | 1 |
| 2003 | Fast stitching algorithm for moving object detection and mosaic constructionabstractThis paper proposes a novel edge-based stitching method to detect moving objects and construct mosaics from images. The method is a coarse-to-fine scheme which first estimates a good initialization of camera parameters with two complementary methods and then refines the solution through an optimization process. The two complementary methods are the edge alignment and correspondence-based approaches, respectively. Since these two methods are complementary to each other, the desired initial estimate can be obtained more robustly. After that, a Monte-Carlo style method is then proposed for integrating these two methods together. Then, an optimization process is applied to refine the above initial parameters. Since the found initialization is very close to the exact solution and only errors on feature positions are considered for minimization, the optimization process can be very quickly achieved. Experimental results are provided to verify the superiority of the proposed method. Jun-Wei Hsieh |
ICME | 1 |
| 2003 | Shadow elimination for effective moving object detection by Gaussian shadow modeling
Jun-Wei Hsieh, Wen-Fong Hu, Yung-Sheng Chen |
Image Vis. Comput. | 1 |
| 2003 | Spatial template extraction for image retrieval by region matchingabstractThis paper presents a template and its relation extraction and estimation (TREE) algorithm for indexing images from picture libraries with more semantics-sensitive meanings. This algorithm can learn the commonality of visual concepts from multiple images to give a middle-level understanding about image contents. In this approach, each image is represented by a set of templates and their spatial relations as keys to capture the essence of this image. Each template is characterized by a set of dominant regions, which reflect different appearances of an object at different conditions and can be obtained by the template extraction and analysis (TEA) algorithm through region matching. The spatial template relation extraction and measurement (STREAM) algorithm is then proposed for obtaining the spatial relations between these templates. Due to the nature of a template, which can represent object's appearances at different conditions, the proposed approach owns better capabilities and flexibilities to capture image contents than traditional region-based methods. In addition, through maintaining the spatial layout of images, the semantic meanings of the query images can be extracted and lead to significant improvements in the accuracy of image retrieval. Since no time-consuming optimization process is involved, the proposed method learns the visual concepts extremely fast. Experimental results are provided to prove the superiority of the proposed method. Jun-Wei Hsieh, W. Eric L. Grimson |
IEEE Trans. Image Process. | 1 |
| 2002 | Template-based image retrievalabstractThe paper presents a TREE (templates and their relationship extraction and estimation) algorithm for indexing images from picture libraries with more semantics-sensitive meanings. In this approach, each image is represented by a set of templates and their spatial relationships as keys to capture the essence of the image. Each template is characterized by a set of dominant regions, which reflect different appearances of an object at different conditions and can be obtained by the proposed TEA (template extraction and analysis) algorithm through region matching. The STREAM (spatial template relationship extraction and measurement) algorithm is then proposed for obtaining the spatial relations between these extracted templates. Due to the nature of a template, which can represent various appearances of an object at different conditions, the proposed approach can provide better capabilities and flexibilities to capture image contents than other traditional region-based methods. Besides, through maintaining the spatial layout of images, the semantic meanings hidden in the query images can be extracted and lead to significant improvements in the accuracy of image retrieval. Jun-Wei Hsieh, W. Eric L. Grimson |
ICME (1) | 1 |
| 2002 | Multiple-Person Tracking System for Content AnalysisabstractThis paper presents a framework to track multiple persons in real-time. First, a method with real-time and adaptable capability is proposed to extract face-like regions based on skin, motion and silhouette features. Then, an adaptable skin model is used for each detected face to overcome the changes of the observed environment. After that, a two-stage face verification algorithm is proposed to quickly eliminate false faces based on face geometries and the SVM (Support Vector Machine) approach. In order to overcome the effect of lighting changes, during verification, a method of color constancy compensation is proposed. Then, a robust tracking scheme is applied to identify multiple persons based on a face-status table. With the table, the proposed system has powerful capabilities to track different persons at different statuses, which is quite important in face-related applications. Experimental results show that the proposed method is more robust and powerful than other traditional methods, which utilize only color, motion information, and the correlation technique. Jun-Wei Hsieh, Yea-Shuan Huang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2000 | Region-Based Image RetrievalabstractThis paper presents a region-based method which uses multiple regions as the key to retrieve images. The proposed method represents the semantics or concepts embedded in the input images with three ingredients. One is a set of regions with weighted importance; the second is the corresponding feature distributions of the regions and the last is the spatial relationships between these regions. The importance of each segmented region in the input example images can be automatically and efficiently determined through a formulated linear system. In addition, a novel method for matching the spatial relationship between regions is also presented to capture the structural semantics of the content of images. By combining the feature distributions and the spatial relationships of regions with appropriate weights, the experimental results show that the retrieval results are much more accurate than other methods which utilize low-level features, such as color, texture, shape, and so on. Jun-Wei Hsieh, W. Eric L. Grimson, Cheng-Chin Chiang, Yea-Shuan Huang |
ICIP | 1 |
| 1997 | PanoVR SDK - a software development kit for integrating photo-realistic panoramic images and 3-D graphical objects into virtual worldsabstractArticle PanoVR SDK—a software development kit for integrating photo-realistic panoramic images and 3-D graphical objects into virtual worlds Share on Authors: Cheng-Chin Chiang Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C. Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C.View Profile , Alex Huang Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C. Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C.View Profile , Tsing-Shin Wang View Profile , Matthew Huang View Profile , Yunn-Yen Chen Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C. Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C.View Profile , Jun-Wei Hsieh View Profile , Ju-Wei Chen View Profile , Tse Cheng Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C. Advanced Technology Center, Computer and Communication Research Laboratories, Industrial Technology Research Institute, Hsinchu, Taiwan, R.O.C.View Profile Authors Info & Claims VRST '97: Proceedings of the ACM symposium on Virtual reality software and technologySeptember 1997 Pages 147–154https://doi.org/10.1145/261135.261162Online:01 September 1997Publication History 8citation641DownloadsMetricsTotal Citations8Total Downloads641Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Cheng-Chin Chiang, Alex Huang, Tsing-Shin Wang, Matthew Huang, Yunn Yen Chen, Jun-Wei Hsieh, Ju-Wei Chen, Tse Cheng |
VRST | 6 |
| 1997 | Image Registration Using a New Edge-Based Approach
Jun-Wei Hsieh, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Tat Ko, Yi-Ping Hung |
Comput. Vis. Image Underst. | 1 |
| 1997 | New automatic multi-level thresholding technique for segmentation of thermal images
Jung-Shiong Chang, Hong-Yuan Mark Liao, Maw-Kae Hor, Jun-Wei Hsieh, Ming-Yang Chern |
Image Vis. Comput. | 4 |
| 1997 | A new wavelet-based edge detector via constrained optimization
Jun-Wei Hsieh, Ming-Tat Ko, Hong-Yuan Mark Liao, Kuo-Chin Fan |
Image Vis. Comput. | 1 |
| 1996 | A fast algorithm for image registration without predetermining correspondencesabstractA novel approach for efficient image registration is proposed. The proposed method applies wavelet transforms to extract a number of feature points as the basis for registration. From the selected feature points, a subset of possible matching pairs is selected to obtain the desired registration parameters. This subset is chosen by using the orientation difference between two target images as a criterion to eliminate spurious matching pairs. In order to predetermine the orientation difference between two target images, a so-called "angle histogram" is calculated. From the angle histogram, the orientation difference can be decided. Once the orientation difference is obtained, the desired subset can be easily determined. By randomly selecting two matching pairs from this subset, a set of registration parameters can be obtained. By checking how many matching pairs are compatible with the selected parameters, the best estimation can be determined. Compared with conventional algorithms, the proposed scheme is a great improvement in terms of efficiency as well as reliability for the image registration problem. Jun-Wei Hsieh, Hong-Yuan Mark Liao, Kuo-Chin Fan, Ming-Tak Ko |
ICPR | 1 |
| 1995 | Wavelet-Based Shape from Shading
Jun-Wei Hsieh, Hong-Yuan Mark Liao, Ming-Tat Ko, Kuo-Chin Fan |
CVGIP Graph. Model. Image Process. | 1 |
| 1995 | Performance of a Mass-Storage System for Video-on-Demand
Jun-Wei Hsieh, Mengjou Lin, Jonathan C. L. Liu, David Hung-Chang Du, Thomas Ruwart |
J. Parallel Distributed Comput. | 1 |
| 1994 | Wavelet-Based Shape from ShadingabstractThis paper proposes a wavelet-based approach to solving the shape from shading (SFS) problem. The proposed method takes advantage of the nature of wavelet theory, which can be applied to efficiently and accurately represent "things", to develop a faster algorithm for reconstructing better surfaces. In order to improve the robustness of the algorithm, two new constraints are introduced into the objective function to strengthen the relation between an estimated surface and its counterpart in the original image. Thus, solving the SFS problem becomes a constrained optimization process. In the first stage of the process, the set of function variables to be solved is represented by a wavelet format. Due to this format, the set of differential operators of different orders which is involved in the whole process can be approximated with the connection coefficients of Daubechies bases. In each iteration of the optimization process an appropriate step size which will result in maximum decrease of the objective function is determined. After finding correct iterative schemes, the solution of the SFS problem will finally be decided. Compared with conventional algorithms, the proposed scheme makes great improvements on the accuracy as well as the convergence speed of the SFS problem.> Jun-Wei Hsieh, Hong-Yuan Mark Liao, Ming-Tat Ko, Kuo-Chin Fan |
ICIP (2) | 1 |