EDBT 2026 Demo / reviewers in the wild / expert
Chien-Yao Wang
dblp:168/6259
· DBLP profile ↗
28ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-2946-8972ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait ScoringabstractPo-Kai Chen, Bo-Wei Tsai, Shao Kuan Wei, Chien-Yao Wang, Jia-Ching Wang, Yi-Ting Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Po-Kai Chen, Bo-Wei Tsai, Shao-Kuan Wei, Chien-Yao Wang, Jia-Ching Wang, Yi-Ting Huang |
ACL (1) | 4 |
| 2025 | Impact of Glyph Information on Latent Space Diffusion Models for Accurate Handwritten Text GenerationabstractThe generation of high-quality stylized handwritten text images is a challenging task in computer vision and artificial intelligence. While advanced approaches using Latent Diffusion Models (LDMs) for generating stylized handwritten text have shown effectiveness, they often struggle with maintaining the structural integrity of certain characters, resulting in issues such as missing or extraneous strokes. In this work, we propose GlyphLDM, an innovative model that integrates glyph image information into both the diffusion and denoising processes in the latent space, enhancing the structural accuracy of generated text images. In the early training stages, our method demonstrated a significant improvement in the structural accuracy of the generated text images, with the Average Confidence Score increasing by approximately 40% compared to the baseline method. These experimental results indicate that incorporating glyph image information has promising potential to enhance the structural accuracy and overall quality of generated text images. This approach provides an effective solution for generating more accurate and diverse handwritten text images. Ying-Li Lin, Chung-I Huang, Chien-Yao Wang, Jia-Ching Wang |
ICASSP | 4 |
| 2025 | A Key to Effective Multi-task Learning: Separate Query Selection for Task-Synergized Handling and Node UtilizationabstractIn the realm of computer vision, effectively handling multi-tasks simultaneously presents a challenge that necessitates innovative solutions. To better address multiple vision problems, we introduce SeTano, an integrated Graph Neural Network (GNN)-based framework. This framework comprises a Dynamic Edge-Sensing GNN (DES-GNN) backbone, which can dynamically adjust edges to extract more pivotal features, and a downstream design which includes a node reduction and a separate query selection strategy. To validate our approach, we perform multi-task experiments on the ImageNet and MS COCO datasets. The results indicate that the integrated design of SeTano leads to enhanced performance in various vision multi-tasks. Shan-Ya Yang, Chien-Yao Wang, Jia-Ching Wang, Chun-Yi Lee |
ICASSP | 3 |
| 2025 | YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-DictionaryabstractIdentifying and localizing objects within images is a fundamental challenge, and numerous efforts have been made to enhance model accuracy by experimenting with diverse architectures and refining training strategies. Nevertheless, a prevalent limitation in existing models is overemphasizing the current input while ignoring the information from the entire dataset. We introduce an innovative $\textbf{R}etriever-\textbf{D}ictionary$ (RD) module to address this issue. This architecture enables YOLO-based models to efficiently retrieve features from a Dictionary that contains the insight of the dataset, which is built by the knowledge from Visual Models (VM), Large Language Models (LLM), or Visual Language Models (VLM). The flexible RD enables the model to incorporate such explicit knowledge that enhances the ability to benefit multiple tasks, specifically, segmentation, detection, and classification, from pixel to image level. The experiments show that using the RD significantly improves model performance, achieving more than a 3\% increase in mean Average Precision for object detection with less than a 1\% increase in model parameters. Beyond 1-stage object detection models, the RD module improves the effectiveness of 2-stage models and DETR-based architectures, such as Faster R-CNN and Deformable DETR. Code is released at https://github.com/henrytsui000/YOLO. Hao-Tang Tsui, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICLR | 2 |
| 2025 | Generalist YOLO: Towards Real-Time End-to-End Multi-Task Visual Language ModelsabstractGeneralist models, capable of handling multiple modalities and tasks simultaneously, are currently one of the hottest research topics. However, due to interference between different tasks during the training process, existing generalist models require a very large decoder to achieve good results in various tasks, which makes real-time prediction difficult for current generalist models. This paper introduces Generalist YOLO, which takes a significant step towards real-time prediction systems for visual language generalist models. The proposed Generalist YOLO uses a unified encoder to reduce conflicts between different tasks, thereby decreasing the complexity required by the decoder. It also introduces a primary-secondary co-attention mechanism that allows different tasks to learn together more effectively, achieving high efficiency and high accuracy. We propose a semantically consistent asymmetric training strategy, allowing various tasks to benefit from performance improvements brought by the latest research results in various fields. The proposed Generalist YOLO achieves excellent results on various vision and language tasks based on MS COCO. While maintaining high accuracy across all tasks, it is 135 times faster than existing generalist models. The source code is released on GitHub at https://github.com/WongKinYiu/GeneralistYOLO. Hung-Shuo Chang, Chien-Yao Wang, Richard Robert Wang, Gene Chou, Hong-Yuan Mark Liao |
WACV | 2 |
| 2024 | YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao |
ECCV (31) | 1 |
| 2024 | Scene Text Recognition Using Progressive Rectification Network And Spelling Error Correction Language ModelabstractScene text recognition has gained popularity in deep neural network research. Compared to document text recognition, scene text recognition faces challenges such as complex backgrounds, diverse fonts, and blurred characters. Visual and semantic information must be considered in text recognition. While recent research has focused on improving semantic information, most studies have used English text datasets. Directly applying these methods to Chinese text datasets may not be effective. This work proposes a vision model with progressive rectification network and a Chinese scene text recognition method that uses a robust error-correcting language model to correct errors predicted by vision models. Firstly, the proposed progressive rectification network is of more effectiveness on the multi-oriented scene text images compared to present rectification method. On the other hand, the designed language model is used to correct errors predicted by vision models. The language model can handle low-quality images, including blurred, occluded, or nonsensical text. Experiments demonstrate that our method outperforms recent classic and state-of-the-art methods, making it a more powerful and suitable option for Chinese scene text recognition. Ming-Zheng Peng, Phuong-Thi Le, Cheng-Chun Wang, Chien-Yao Wang, Jia-Ching Wang |
ICIP | 5 |
| 2024 | mmAlphabet: Air Writing Alphabet Recognition System Based on mmWave FMCW Radar and Convolutional Neural Network
Chao-Wang Huang, Chien-Yao Wang, Jia-Ching Wang |
ICPR (28) | 2 |
| 2023 | YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object DetectorsabstractReal-time object detection is one of the most important research topics in computer vision. As new approaches regarding architecture optimization and training optimization are continually being developed, we have found two research topics that have spawned when dealing with these latest state-of-the-art methods. To address the topics, we propose a trainable bag-of-freebies oriented solution. We combine the flexible and efficient training tools with the proposed architecture and the compound scaling method. YOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 120 FPS and has the highest accuracy 56.8% AP among all known real-time object detectors with 30 FPS or higher on GPU V100. Source code is released in https://github.com/WongKinYiu/yolov7. Chien-Yao Wang, Aleksei Bochkovskii, Hong-Yuan Mark Liao |
CVPR | 1 |
| 2023 | 3D Face Reconstruction Based on Weakly-Supervised Learning Morphable Face ModelabstractIn this paper, we propose a system for 3D face model reconstruction. Earlier studies on reconstruction methods included the software modeling methods or the instrument scanning modeling methods. But both of the above methods require a lot of development resources and time costs. Therefore, we develop a reconstruction system using a weakly supervised approach combining Convolutional Neural Networks (CNN) and 3D Morphable Face Models (3DMM). Given a sufficient number of 2D face images to train and learn the main features of the face, our system is capable of rapidly constructing 3D face models. The proposed method enhances the efficiency of preprocessing and improves the performance of loss function through image depth feature extraction and regression coefficients. Using two datasets for model evaluation and analysis, this study efficiently reconstructs faces without ground-truth labels. Kai-Wen Liang, Pin-Hsuan Li, Chung-Hsun Lo, Chien-Yao Wang, Yung-Fang Chen, Jia-Ching Wang, Pao-Chi Chang |
ICIP | 4 |
| 2022 | SearchTrack: Multiple Object Tracking with Object-Customized Search and Motion-Aware Features
Zhong-Min Tsai, Yu-Ju Tsai, Chien-Yao Wang, Hong-Yuan Mark Liao, Youn-Long Lin, Yung-Yu Chuang |
BMVC | 3 |
| 2022 | Spectral-Temporal Receptive Field-Based Descriptors and Hierarchical Cascade Deep Belief Network for Guitar Playing Technique ClassificationabstractMusic information retrieval is of great interest in audio signal processing. However, relatively little attention has been paid to the playing techniques of musical instruments. This work proposes an automatic system for classifying guitar playing techniques (GPTs). Automatic classification for GPTs is challenging because some playing techniques differ only slightly from others. This work presents a new framework for GPT classification: it uses a new feature extraction method based on spectral-temporal receptive fields (STRFs) to extract features from guitar sounds. This work applies a supervised deep learning approach to classify GPTs. Specifically, a new deep learning model, called the hierarchical cascade deep belief network (HCDBN), is proposed to perform automatic GPT classification. Several simulations were performed and the datasets of: 1) data on onsets of signals; 2) complete audio signals; and 3) audio signals in a real-world environment are adopted to compare the performance. The proposed system improves upon the F-score by approximately 11.47% in setup 1) and yields an F-score of 96.82% in setup 2). The results in setup 3) demonstrate that the proposed system also works well in a real-world environment. These results show that the proposed system is robust and has very high accuracy in automatic GPT classification. Chien-Yao Wang, Pao-Chi Chang, Jian-Jiun Ding, Tzu-Chiang Tai, Andri Santoso, Yu-Ting Liu, Jia-Ching Wang |
IEEE Trans. Cybern. | 1 |
| 2021 | Scaled-YOLOv4: Scaling Cross Stage Partial NetworkabstractWe show that the YOLOv4 object detection neural network based on the CSP approach, scales both up and down and is applicable to small and large networks while maintaining optimal speed and accuracy. We propose a network scaling approach that modifies not only the depth, width, resolution, but also structure of the network. YOLOv4-large model achieves state-of-the-art results: 55.5% AP (73.4% AP50) for the MS COCO dataset at a speed of ~ 16 FPS on Tesla V100, while with the test time augmentation, YOLOv4-large achieves 56.0% AP (73.3 AP50). To the best of our knowledge, this is currently the highest accuracy on the COCO dataset among any published work. The YOLOv4-tiny model achieves 22.0% AP (42.0% AP50) at a speed of ~443 FPS on RTX 2080Ti, while by using TensorRT, batch size = 4 and FP16-precision the YOLOv4-tiny achieves 1774 FPS. Chien-Yao Wang, Aleksei Bochkovskii, Hong-Yuan Mark Liao |
CVPR | 1 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 5 |
| 2020 | How Incompletely Segmented Information Affects Multi-Object Tracking and Segmentation (MOTS)abstractIn recent years, deep learning has made dramatic advances in computer vision field, especially in improving the performance of object detection as well as instance semantic segmentation. Still, multi-object tracking (MOT) remains a very challenging issue. Even in state-of-the-art deep learning-based object detectors, a preferred paradigm for MOT: tracking-by-detection, can only slightly improve the tracking performance. Pixel-level information is considered more precise and useful for tracking performance improvement than using conventional information, such as foreground or background content in a bounding box. However, the performance of current state-of-the-art models for automatically annotating pixel-level information is still far from the expectation of human beings. Therefore, we shall explore how multi-object tracking and segmentation (MOTS) is affected when the information obtained after applying instance semantic segmentation is incomplete. We propose a mask-guided two-streamed augmentation learning (MGTSAL) algorithm, which can be applied to TrackR-CNN to alleviate significant drop of MOTS performance when encountering incompletely segmented information. We evaluate the proposed approach on MOTS KITTI dataset, and our approach outperforms the baseline model TrackR-CNN in all our experimental settings. The promising experimental results and ablation study validate the effectiveness of the proposed approach. Yu-Sheng Chou, Chien-Yao Wang, Shou-De Lin, Hong-Yuan Mark Liao |
ICIP | 2 |
| 2020 | Sound Events Recognition and Retrieval Using Multi-Convolutional-Channel Sparse Coding Convolutional Neural NetworksabstractThis article proposes two novel deep convolutional neural networks (CNN), which are called the sparse coding convolutional neural network (SC-CNN) and the multi-convolutional-channel SC-CNN (MSC-CNN), to address the sound event recognition and retrieval problem. Unlike the general framework of a CNN, in which the feature learning process is performed hierarchically, the proposed framework models the whole memorization process in the human brain, including encoding, storage, and recollection. In particular, the MSC-CNN is designed to recognize multiple sound events that occur simultaneously. The experimental results indicate that the proposed SC-CNN and MSC-CNN outperforms the state-of-the-art systems in sound event recognition and retrieval. Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Dynamic Gallery for Real-Time Multi-Target Multi-Camera TrackingabstractFor multi-target multi-camera recognition tasks, tracking of objects of interest is one of the essential yet challenging issues due to the fact that the task requires re-identifying identical targets across distinct views. Multi-target multi-camera tracking (MTMCT) applications span a wide range of variety (e.g. crowd behavior analysis, anomaly individual tracking and sport player tracking), so how to make the system perform real-time tracking becomes a crucial research issue. In this paper, we propose an online hierarchical algorithm for extreme clustering based MTMCT framework. The system can automatically create a dynamic gallery with real-time fashion by collecting appearance information of multi-object tracking in single-camera view. We evaluate the effectiveness and efficiency of our framework, and compare the state-of-the-art methods on MOT16 as well as DukeMTMC for single and multiple camera tracking. The high-frame-rate performance and promising tracking results confirm our system can be used in realworld applications. Yu-Sheng Chou, Chien-Yao Wang, Ming-Chiao Chen, Shou-De Lin, Hong-Yuan Mark Liao |
AVSS | 2 |
| 2019 | Real-Time Video-Based Person Re-Identification Surveillance with Light-Weight Deep Convolutional NetworksabstractToday's person re-ID system mostly focuses on accuracy and ignores efficiency. But in most real-world surveillance systems, efficiency is often considered the most important focus of research and development. Therefore, for a person re-ID system, the ability to perform real-time identification is the most important consideration. In this study, we implemented a real-time multiple camera video-based person re-ID system using the NVIDIA Jetson TX2 platform. This system can be used in a field that requires high privacy and immediate monitoring. This system uses YOLOv3-tiny based light-weight strategies and person re-ID technology, thus reducing 46% of computation, cutting down 39.9% of model size, and accelerating 21% of computing speed. The system also effectively upgrades the pedestrian detection accuracy. In addition, the proposed person re-ID example mining and training method improves the model's performance and enhances the robustness of cross-domain data. Our system also supports the pipeline formed by connecting multiple edge computing devices in series. The system can operate at a speed up to 18 fps at 1920×1080 surveillance video stream. The demo of our developed systems can be found at https://sites.google.com/g.ncu.edu.tw/video-based-person-re-id/. Chien-Yao Wang, Ping-Yang Chen, Ming-Chiao Chen, Jun-Wei Hsieh, Hong-Yuan Mark Liao |
AVSS | 1 |
| 2019 | Smaller Object Detection for Real-Time Embedded Traffic Flow Estimation Using Fish-Eye CamerasabstractReal-time embedded traffic flow estimation (RETFE) systems need accurate and efficient vehicle detection models to meet limited resources in budget, dimension, memory, and computing power. In recent years, object detection became a less challenging task with latest deep CNN-based state-of-the-art models, i.e., RCNN, SSD, and YOLO; however, these models cannot provide desired performance for RETFE systems due to their complex time-consuming architecture. In addition, small object (<; 30×30 pixels) detection is still a challenging task for existing methods. Thus, we propose a shallow model named Concatenated Feature Pyramid Network (CFPN) that inspired from YOLOv3 to provide above mentioned performance for the smaller object detection. Main contribution is a proposed concatenated block (CB) which has reduced number of convolutional layers and concatenations instead of time-consuming algebraic operations. The superiority of CFPN is confirmed on the COCO and an in-house CarFlow datasets on Nvidia TX2. Thus we conclude that CFPN is useful for real-time embedded smaller object detection task. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Chien-Yao Wang, Hong-Yuan Mark Liao |
ICIP | 4 |
| 2019 | Object Bounding Transformed Network for End-to-End Semantic SegmentationabstractIn recent years, numerous studies of the use of a Fully Convolutional Network (FCN) for image semantic segmentation have been published. This work introduces an end-to-end Object Bounding Transformed Network (OBTNet) which combines the advantages of the Object Boundary Guided (OBG) and Doman Transform (DT). OBG is an object boundary based approach that increases the integrity of object shape. Based on OBG, we propose an Object Boundary Network (OBN) as the object region and object boundary generator. In addition, our system achieves object region preserving and object boundary preserving by employing DT. The proposed system uses the pretrained multi-scale ResNet101 as the base network and uses multi-scale atrous convolution to preserve the dimensions of the feature map, increasing the accuracy of semantic segmentation. Experiments show that our system yielded a mean IOU of 77.74% and outperformed the baseline model on the VOC2012 test set. Kuan-Chung Wang, Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang |
ICIP | 2 |
| 2018 | Sound Event Recognition Using Auditory-Receptive-Field Binary Pattern and Hierarchical-Diving Deep Belief NetworkabstractAutomatic sound event recognition (SER) has recently attracted renewed interest. Although practical SER system has many useful applications in everyday life, SER is challenging owing to the variations among sounds and noises in the real-world environment. This paper presents a novel feature extraction and classification method to solve the problem of SER. An audio-visual descriptor, called the auditory-receptive-field binary pattern, is designed based on the spectrogram image feature, the cepstral features, and the human auditory receptive field model. The extracted features are then fed into a classifier to perform event classification. The proposed classifier, called the hierarchical-diving deep belief network, is a deep neural network system that hierarchically learns the discriminative characteristics from physical feature representation to the abstract concept. The performance of our proposed system was verified using several experiments under various conditions. Using the RWCP dataset, the proposed system achieved a recognition rate of 99.27% for real-world sound data in 105 categories. Under noisy conditions, the developed system is very robust, with which it achieved 95.06% recognition rate with 0 dB signal-to-noise ratio. Using the TUT sound event dataset, the proposed system achieves error rates of 0.81 and 0.73 in sound event detection in home and residential area scenes. The experimental results reveal that the proposed system outperformed the other systems in this field. Chien-Yao Wang, Jia-Ching Wang, Andri Santoso, Chin-Chin Chiang, Chung-Hsien Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Fully complex deep neural network for phase-incorporating monaural source separationabstractDeep neural network (DNN) have become a popular means of separating a target source from a mixed signal. Most of DNN-based methods modify only the magnitude spectrum of the mixture. The phase spectrum is left unchanged, which is inherent in the short-time Fourier transform (STFT) coefficients of the input signal. However, recent studies have revealed that incorporating phase information can improve the quality of separated sources. To estimate simultaneously the magnitude and the phase of STFT coefficients, this work paper developed a fully complex-valued deep neural network (FCDNN) that learns the nonlinear mapping from complex-valued STFT coefficients of a mixture to sources. In addition, to reinforce the sparsity of the estimated spectra, a sparse penalty term is incorporated into the objective function of the FCDNN. Finally, the proposed method is applied to singing source separation. Experimental results indicate that the proposed method outperforms the state-of-the-art DNN-based methods. Yuan-Shan Lee, Chien-Yao Wang, Shu-Fan Wang, Jia-Ching Wang, Chung-Hsien Wu 0001 |
ICASSP | 2 |
| 2017 | Dynamic tracking attention model for action recognitionabstractThis paper proposes a dynamic tracking attention model (DTAM), which mainly comprises a motion attention mechanism, a convolutional neural network (CNN) and long short-term memory (LSTM), to recognize human action in a video sequence. In the motion attention mechanism, the local dynamic tracking is used to track moving objects in feature domain and global dynamic tracking corrects the motion in the spectral domain. The CNN is utilized to perform feature extraction, while the LSTM is applied to handle sequential information about actions that is extracted from videos. It effectively fetches information between consecutive frames in a video sequence and has an even higher recognition rate than does the CNN-LSTM. Combining the DTAM with the visual attention model, the proposed algorithm has a recognition rate that is 3.6% and 4.5% higher than that of the CNN-LSTMs with and without the visual attention model, respectively. Chien-Yao Wang, Chin-Chin Chiang, Jian-Jiun Ding, Jia-Ching Wang |
ICASSP | 1 |
| 2017 | Hierarchical joint-guided networks for semantic image segmentationabstractSemantic image segmentation is now an exciting area of research owing to its various useful applications in daily life. This paper introduces a hierarchical joint-guided network (HJGN) which is mainly composed of proposed hierarchical joint learning convolutional networks (HJLCNs) and proposed joint-guided and making networks (JGMNs). HJLCNs exhibit high robustness in the segmentation of unseen objects that are not contained in training categories. JGMNs are very effective in filling holes and preventing incorrect segmentation predictions. The proposed HJGNs outperform the state-of-the-art methods on the PASCAL VOC 2012 testing set, reaching a mean IU of 80.4%. Chien-Yao Wang, Jyun-Hong Li, Seksan Mathulaprangsan, Chin-Chin Chiang, Jia-Ching Wang |
ICASSP | 1 |
| 2017 | Recognition and retrieval of sound events using sparse coding convolutional neural networkabstractThis paper proposes a novel deep convolutional neural network (CNN), called sparse coding convolutional neural network (SC-CNN), to address the problem of sound event recognition and retrieval task. Unlike the general framework of a CNN, in which feature learning process is performed hierarchically, the proposed framework models the whole memorizing procedures in the human brain, including encoding, storage, and recollection. Sound data from the RWCP sound scene dataset with added noise from NOISEX-92 noise dataset are used to compare the performance of the proposed system with the state-of-the-art baselines. The experimental results indicated that the proposed SC-CNN outperformed the state-of-the-art systems in sound event recognition and retrieval. In the sound event recognition task, the proposed system achieved an accuracy of 94.6%, 100% and 100% under 0db, 10db and clean noise conditions, respectively. In the retrieval task, the proposed system improves the mAP rate of the general CNN by approximately 6%. Chien-Yao Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001, Jia-Ching Wang |
ICME | 1 |
| 2017 | Spectral-temporal receptive fields and MFCC balanced feature extraction for robust speaker recognition
Jia-Ching Wang, Chien-Yao Wang, Yu-Hao Chin, Yu-Ting Liu, En-Ting Chen, Pao-Chi Chang |
Multim. Tools Appl. | 2 |
| 2016 | Locality-preserving K-SVD Based Joint Dictionary and Classifier Learning for Object RecognitionabstractThis paper concerns the development of locality-preserving methods for object recognition. The major purpose is consideration of both descriptor-level locality and image-level locality throughout the recognition process. Two dual-layer locality-preserving methods are developed, in which locality-constrained linear coding (LLC) is used to represent an image. In the learning phase, the discriminative locality-preserving K-SVD (DLP-KSVD) in which the label information is incorporated into the locality-preserving term is proposed. In addition to using class labels to learn a linear classifier, the label-consistent LP-KSVD (LCLP-KSVD) is proposed to enhance the discriminability of the learned dictionary. In LCLP-KSVD, the objective function includes a label-consistent term that penalizes sparse codes from different classes. For testing, additional information about the locality of query samples is obtained by treating the locality-preserving matrix as a feature. The recognition results that were obtained in experiments with the Caltech101 database indicate that the proposed method outperforms existing sparse coding based approaches. Yuan-Shan Lee, Chien-Yao Wang, Seksan Mathulaprangsan, Jia Hao Zhao, Jia-Ching Wang |
ACM Multimedia | 2 |
| 2015 | Kernel Sparse Representation Classifier with Center Enhanced SPM for Vehicle ClassificationabstractIn this paper, we proposes a visual-based vehicle classification system, in which it involves visual feature representation and classification step. In the feature representation step, we present a center enhanced spatial pyramid matching (CE-SPM) to extract the feature from images. In this work, we defined additional region in the center of each images to calculate the histograms of visual words and then pool them together with some weights to construct the feature representation vector of an image. In the classification step, kernel sparse representation classifier is used to address the problem of visual-based vehicle classification. The kernel function maps the features from original space into higher space dimension. The modified active-set algorithm for l1 non-negative least square problem is adopted to solve the optimization problem. The experimental results show the improvement of proposed method over the original SPM. The proposed method can achieve the performance of 93.7% using particular vehicle image dataset. Andri Santoso, Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang |
COMPSAC | 2 |