Chuan-Kai Yang

dblp:29/1787 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-4782-3621ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1Theory of computation · 1
YearPublicationVenuePosition
2026 A real-time bird sound recognition app via deep learning techniques
abstract
This study presents a real-time, on-device bird sound recognition system developed using deep transfer learning and optimized for mobile deployment. A curated Xeno-canto corpus, an open-access repository of wildlife sound recordings contributed by citizen scientists worldwide, comprising 610 Taiwanese bird species was used to evaluate six deep learning architectures: Residual Network-18 (ResNet-18), Yet Another Mobile Network (YAMNet), Visual Geometry Group-like Network for Audio Classification (VGGish), Convolutional Neural Network–Long Short-Term Memory (CNN-LSTM), Attention-based Convolutional Neural Network (Attention-CNN), and a Deep Neural Network (DNN) baseline. All models were trained using class weighting, batch normalization, a dropout rate of 0.2, and targeted data augmentation, including pitch shifting (±2 semitones), time stretching (0.8–1.2 $$\times $$ ), and time shifting ( $$\le $$ 16,000 samples). Among these, ResNet-18 achieved the best balance between accuracy and computational efficiency, with an overall accuracy of 0.955, macro-precision of 0.95, macro-recall of 0.94, and macro-F1 of 0.945 across all 610 classes. The model performs inference in 25.9 milliseconds with only 3.03 megabytes of memory (approximately 795,000 parameters), outperforming heavier architectures such as VGGish (0.8975 accuracy, 42.2 milliseconds, 587 megabytes) while remaining competitive with compact alternatives like YAMNet (0.935 accuracy, 27.0 milliseconds, 10.19 megabytes). Furthermore, Gradient-weighted Class Activation Mapping (Grad-CAM) visualizations confirm that predictions are driven by species-specific temporal–spectral patterns rather than background noise. Converting the optimized model to TensorFlow Lite enables fully offline inference on Android devices, eliminating cloud latency and ensuring user privacy. Overall, this lightweight, high-accuracy framework offers a scalable and practical solution for real-time biodiversity monitoring and conservation research.
Hailemariam Abebe Endalamaw, Chuan-Kai Yang, Cheng-Hung Hsu
Multim. Tools Appl.2
2026 Two-stage fine-tuning of HuBERT for multi-label bird species recognition in overlapping acoustic environments
abstract
Automated recognition of bird species from audio is critical for biodiversity monitoring, yet it remains difficult in practice because field recordings often contain multiple birds vocalizing at the same time, strong environmental noise, and limited labeled data. Most existing systems either assume single-species recordings, require clean inputs, or depend on manually engineered preprocessing, such as source separation. This work introduces a novel two-stage fine-tuning framework that adapts a large self-supervised speech model (HuBERT) to the highly non-speech, polyphonic, multi-label setting of wild bird soundscapes. The proposed approach departs from conventional direct fine-tuning by using a two-stage curriculum. In Stage 1, HuBERT is fine-tuned on clean single-species recordings to learn discriminative, species-specific acoustic representations without interference. In Stage 2, the model is then transferred and further fine-tuned on synthetically constructed overlapping vocalizations, enabling it to generalize to real noisy soundscapes where multiple species co-occur. This two-stage adaptation strategy bridges the acoustic gap between human speech pretraining and avian bioacoustics, and allows robust multi-label prediction directly on overlapping audio without requiring explicit source separation. Extensive experiments on ten bird species show that the proposed two-stage HuBERT achieves an F1-score of 0.94 on overlapping recordings, outperforming (i) HuBERT variants trained only on clean or overlapping audio, and (ii) state-of-the-art CNN, RNN, graph-based, and transformer baselines reported in prior studies. These results demonstrate that two-stage self-supervised adaptation is an effective and scalable direction for real-time, multi-species bird monitoring in complex natural environments.
Hailemariam Abebe Endalamaw, Chuan-Kai Yang
Nat. Comput.2
2025 Visualizing NBA information via storylines
Chuan-Kai Yang, Chiun-How Kao
Comput. Graph.2
2025 Water-sensitive paper detection and spray analysis
abstract
Global population has grown from 1 billion in the 19th century to 7.9 billion today, and is projected to reach 9.2 billion by 2050. To meet the growing demand for food, we need to increase crop yields. However, it is a challenging task to keep crop yields high all the time. In the process of food production, crops are particularly vulnerable to pests and diseases. Spraying pesticides is a common and effective way to control pests and diseases in modern agriculture. However, excessive use of pesticides can damage the ecological environment, and pesticide residues can also harm human health. Therefore, it is essential to be able to evaluate spraying effects conveniently and accurately. This paper adopts water-sensitive paper (WSP) as a spraying evaluation tool. Previous studies have also moved the task of evaluating spraying results to mobile devices. Although they can analyze the results in real time, there are many limitations, such as the need for the paper to be photographed vertically and the image to be clear and free of noise. Therefore, this paper proposes a paper image processing flow that can break through the limitations of previous studies.per uses a Faster R-CNN model to identify the paper in the image, uses Grab Cut to extract the paper image, then calibrates the paper to a standard size, removes shadows, and uses the Otsu threshold method and K-means clustering method to segment the spray droplet area. The bounding box MIoU of paper identification is 0.9766; the MIoU of the WSP is 0.9824. In the subsequent experiments, the spraying coverage percentage evaluation results of three different types of paper images, namely regular paper (Regular WSP), scanned paper (Scan WSP), and outdoor paper (Outdoor WSP), are presented. The results of droplet segmentation are also compared with those of the DropLeaf water-sensitive paper analysis App proposed in previous studies. In the study of water-sensitive paper, it was found that there is currently no source of water-sensitive paper database. Due to the lack of a suitable database, this paper provides a method for synthesizing paper images. By inputting a scanned paper (Scan WSP) as the input image, a synthesized paper image can be generated, which will benefit researchers who want to conduct different experiments on water-sensitive paper in the future.
Chia-Lin Wu, Chuan-Kai Yang, Ji-Yang Lin
Multim. Tools Appl.2
2025 A WebGL-based 3D furniture modeling system via light-field descriptor and interactive force-directed visualization
abstract
In recent years, the demand for designing and modeling tools for non-professional users has been rising. Among the available tools and methods, customizing an item of furniture according to personal preferences has also become increasingly popular. However, for non-professional modeling users, it is still not easy to design a customized 3D model from scratch. Therefore, this paper proposes an online 3D furniture modeling system to assist users to quickly and conveniently generate an ideal 3D furniture model. This system makes use of existing furniture models. We first extract the features of a 3D model via the LightField Descriptor and perform similarity calculation. We then use the features of furniture data to form a Force-Directed Layout to show the similarity among models. Our system allows users to find and select rough furniture models quickly. And through a simple interactive modeling visualization interface, users can change the components of furniture models. In addition, our system provides suggestions for similar components, as well as component adjustment tools, so that users can easily replace and edit components. As a result, an ideal furniture model can be produced. A user study has been conducted to show the effectiveness of our proposed system.
San-Chi Yeh, Chuan-Kai Yang, Ling-Chi Cheng
Multim. Tools Appl.2
2023 Video visualization via face and speaker clustering
Dehvari Mojiborrahman, Chuan-Kai Yang
Multim. Tools Appl.2
2022 Privacy protection and beautification of cornea images
Chia-Lin Wu, Chuan-Kai Yang
Multim. Tools Appl.2
2022 Realistic video generation for american sign language
Meng-Chen Xu, Chuan-Kai Yang
Multim. Tools Appl.2
2021 An object recognition system based on convolutional neural networks and angular resolutions
Achmad Lukman, Chuan-Kai Yang
Multim. Tools Appl.2
2020 A Multi-Person Selfie System via Augmented Reality
abstract
Abstract The limited length of a selfie stick always poses the problem of distortion in a selfie, in spite of the prevalence of selfie stick in recent years. We propose a technique, based on modifying existing augmented reality technology, to support the selfie of multiple persons, through properly aligning different photographing processes. It can be shown that our technique helps avoiding the common distortion drawback of using a selfie stick, and facilitates the composition process of a group photo. It can also be used to create some special effects, including creating an illusion of having multiple appearances of a person.
Chuan-Kai Yang
Comput. Graph. Forum2
2019 Automatic generation of video navigation from Google Street View data with car detection and inpainting
Yuan-Bang Cheng, Chuan-Kai Yang, Guan-Chung Chang, Teng-Wen Chang
Multim. Tools Appl.2
2017 Using data visualization technique to detect sensitive information re-identification problem of real open dataset
Chiun-How Kao, Chih-Hung Hsieh, Yu-Feng Chu, Yu-Ting Kuang, Chuan-Kai Yang
J. Syst. Archit.5
2016 Integrated model/scene construction through context-based search, data-driven suggestion and component replacement
Po-An Chen, Chuan-Kai Yang
Multim. Tools Appl.2
2016 Encryption domain content-based image retrieval and convolution through a block-based transformation algorithm
Jia-Kai Chou, Chuan-Kai Yang, Hsing-Ching Chang
Multim. Tools Appl.2
2016 Automatic hair extraction from 2D images
Chuan-Kai Yang, Chia-Ning Kuo
Multim. Tools Appl.1
2016 Obfuscated volume rendering
Jia-Kai Chou, Chuan-Kai Yang
Vis. Comput.2
2015 Video Object Retrieval by Trajectory and Appearance
abstract
The prevalence of video recording capability, either on surveillance systems or mobile devices, has contributed to the popularity of video data. As a result, video management has become relatively more important than before and, in particular, video retrieval has been one of the main issues in this regard. Traditional video retrieval systems take texts as the inputs to look for similar information from the title, annotation or embedded textual data of a video, in a way that is very similar to the keyword search adopted by a common search engine. However, the lack of visual information specification during a search often makes the result rather inaccurate or even useless. For this reason, video retrieval systems using images or videos as the inputs have also been proposed; nevertheless, the associated ambiguity and complexity have made the implementation of such systems relatively difficult and, therefore, those systems are not as successful as desired. To address this, in this paper, we propose to perform a video retrieval of a desired object through the inputs of its trajectory and/or appearance, together with the help of a 3-D graphical user interface for more intuitive interactions, so that more satisfactory results can be achieved. We firmly believe that such a framework could serve as the foundation for behavior analysis used in many surveillance systems.
Yuan-Hao Lai, Chuan-Kai Yang
IEEE Trans. Circuits Syst. Video Technol.2
2013 A Study on Enhancing Timeline-Like Visualization with Verbal Text
abstract
There has been a long-standing question of whether sound or audio annotations may assist in data visualization tasks. This paper presents our study focusing on enhancing timeline/storyline like visualizations with audio annotations. Timeline visualizations are widely used to illustrate interactions among different entities over time. This type of visualizations facilitate reviewing important activities and their associations. For a long timeline, however, it is difficult for the viewer to remember all the essential information found in the visualization process. Since hearing is another primary human sense for perceiving information, we conjecture that adding audio annotations to selected sections of a timeline visualization can help improve the viewer's recall of important events. We have designed a user study based on augmenting storyline visualizations with verbal text, and tested subjects with three different settings: visual cues only, verbal text only, and both. While our test results do not give a strong indication of the clear advantage of adding verbal text, the lessons learned in our study suggest directions for further study and will help us and others design audio-augmented visualization systems.
Jia-Kai Chou, Isaac Liao, Kwan-Liu Ma, Chuan-Kai Yang
CW4
2013 Simulation of face/hairstyle swapping in photographs with skin texture synthesis
Jia-Kai Chou, Chuan-Kai Yang
Multim. Tools Appl.2
2012 Face-off: automatic alteration of facial features
Jia-Kai Chou, Chuan-Kai Yang, Sing-Dong Gong
Multim. Tools Appl.2
2012 Fast architecture prototyping through 3D collage
Chuan-Kai Yang, Ching-Yang Tsai
Multim. Tools Appl.1
2011 PaperVis: Literature Review Made Easy
abstract
Abstract Reviewing literatures for a certain research field is always important for academics. One could use Google‐like information seeking tools, but oftentimes he/she would end up obtaining too many possibly related papers, as well as the papers in the associated citation network. During such a process, a user may easily get lost after following a few links for searching or cross‐referencing. It is also difficult for the user to identify relevant/important papers from the resulting huge collection of papers. Our work, called PaperVis, endeavors to provide a user‐friendly interface to help users quickly grasp the intrinsic complex citation‐reference structures among a specific group of papers. We modify the existing Radial Space Filling (RSF) and Bullseye View techniques to arrange involved papers as a node‐link graph that better depicts the relationships among them while saving the screen space at the same time. PaperVis applies visual cues to present node attributes and their transitions among interactions, and it categorizes papers into semantically meaningful hierarchies to facilitate ensuing literature exploration. We conduct experiments on the InfoVis 2004 Contest Dataset to demonstrate the effectiveness of PaperVis.
Jia-Kai Chou, Chuan-Kai Yang
Comput. Graph. Forum2
2010 Digital product transaction mechanism for electronic auction environment
abstract
The rapid development in electronic commerce and information technology drives the traditional physical product trading evolved to digital product trading. With the effect of the multi-agents system in the Internet environment and the promotions of Government, digital product industry grows fast. The authors proposed a digital product transaction mechanism for electronic auction in the multi-agents system environment. The research introduced a convenient platform to protect the privacies of both buyers and sellers, and track digital product further in an electronic auction environment. In addition, by using simple cryptography techniques supplemented with encryption, the authors ensure the security of information transactions, thereby providing a mechanism of safe and fair digital product electronic auction.
Chih-Ta Yen, Tzong-Chen Wu, Ming-Huang Guo, Chuan-Kai Yang, Han-Chieh Chao
IET Inf. Secur.4
2009 Synthesizing solid particle textures via a visual hull algorithm
Hsing-Ching Chang, Chuan-Kai Yang, Jia-Wei Chiou, Shih-Hsien Liu
Comput. Graph.2
2008 An Optimal Algorithm for the Minimum Disc Cover Problem
Min-Te Sun, Chih-Wei Yi, Chuan-Kai Yang, Ten-Hwang Lai
Algorithmica3
2008 An interactive facial expression generation system
Chuan-Kai Yang, Wei-Ting Chiang
Multim. Tools Appl.1
2008 Realization of Seurat's pointillism via non-photorealistic rendering
Chuan-Kai Yang, Hui-Lin Yang
Vis. Comput.1
2006 Integration of volume decompression and out-of-core iso-surface extraction from irregular volume data
Chuan-Kai Yang, Tzi-cker Chiueh
Vis. Comput.1
2005 A Simple and Novel Seed-Set Finding Approach for Iso-Surface Extraction
abstract
Iso-surface extraction is one of the most important approaches for volume rendering, and iso-contouring is one of the most effective methods for iso-surface extraction. Unlike most other methods having their search domain to be the whole dataset, iso-contouring does its search only on a relatively small subset of the original data-set. This subset, called a seed-set, has the property that every iso-surface must intersect with it, and it could be built at the preprocessing time. When an iso-value is given at the run time, iso-contouring algorithm starts from the intersected cells in the seed-set, and gradually propagates to form the whole iso-surface. As smaller seed-sets offer less cell searching time, most existing iso-contouring algorithms concentrates on how to identify an optimal seed-set. In this paper, we propose a new and linear-time approach for seed-set construction. This presented algorithm could reduce the size of the generated seed-sets by up to one or two orders of magnitude, compared with other previously proposed fast (linear time) algorithms.
Chiang-Han Hung, Chuan-Kai Yang
EuroVis2
2000 On-the-Fly rendering of losslessly compressed irregular volume data
abstract
Very large irregular-grid data sets are represented as tetrahedral meshes and may incur significant disk I/O access overhead in the rendering process. An effective way to alleviate the disk I/O overhead associated with rendering a large tetrahedral mesh is to reduce the I/O bandwidth requirement through compression. Existing tetrahedral mesh compression algorithms focus only on compression efficiency and cannot be readily integrated into the mesh rendering process, and thus demand that a compressed tetrahedral mesh be decompressed before it can be rendered into a 2D image. This paper presents an integrated tetrahedral mesh compression and rendering algorithm called Gatun, which allows compressed tetrahedral meshes to be rendered incrementally as they are being decompressed, thus leading to an efficient irregular grid rendering pipeline. Both compression and rendering algorithms in Gatun exploit the same local connectivity information among adjacent tetrahedra, and thus can be tightly integrated into a unified implementation framework. Our tetrahedral compression algorithm is specifically designed to facilitate the integration with an irregular grid renderer without any compromise in compression efficiency. A unique performance advantage of Gatun is its ability to reduce the runtime memory footprint requirement by releasing memory allocated to tetrahedra as early as possible.
Chuan-Kai Yang, Tulika Mitra, Tzi-cker Chiueh
IEEE Visualization1
2000 Zodiac: A history-based interactive video authoring system
Tzi-cker Chiueh, Tulika Mitra, Anindya Neogi, Chuan-Kai Yang
Multim. Syst.4
1998 Zodiac: A History-Based Interactive Video Authoring System
abstract
streams with other media types such as audio is an es-Fasy-io-use audio[video authoting took play a cmcial Ta2ein moving multimedia pTogTamsfrom TeseaTchcu-Tiosify to main-stTeam applications.This papeT desckbes the design and implementation of an inteTactitre video authoTing system called Zodiac, which employs an innovative edit histoTy abstraction to suppoTt Sevemt unique editing jeafuTes notfoun~in existing commercial and Te-seaTch video editing systems.Zo disc pTOVideSUSeTSa concepti~ally clean and semantically poweTfil bran&ing history model of edit operations to oTganize the authoTingpTocess, and to navigate among the design aL teTnatives.In addition, by analyzing the edit histo~, Zodiac is able to Teiiably detect a composed stTeam5 shot and scene boznda~es, ~hi& faci~itate intemctive ~~ideobTowsiny.Zodiac ako featuTes a video object annotation capability that allows meTs to associate annotations to moving objects in a video sequence.The annotations themselves could be text, image, audio, OT ~'ideo.Zodiac is built on top of hf3fFS, a file system specifica~~ydesigned foT inteTac~~ve mu~~imedia de- velaprnent environments, and implements an internal bufieT manage~that suppoTk transparent lossless com pT~ssion/decompression.Shot/scene detection, video object annotation, and bu~eT management all exploit the edit historr information foT peTfomance optimization.A complete digitd video authoring environment must support two fundamental functions: capturing and generation of raw video cfips, and temporal arrangement of video segments with special inter-segment transition eflects.b addition, the abiity to synchronize video Permissionto make digilzl or hard copies of all or part of thrs \trork for personzl or cl=
Tzi-cker Chiueh, Tulika Mitra, Anindya Neogi, Chuan-Kai Yang
ACM Multimedia4
1997 Integrated volume compression and visualization
abstract
Volumetric data sets require enormous storage capacity even at moderate resolution levels. The excessive storage demands not only stress the capacity of the underlying storage and communications systems, but also seriously limit the speed of volume rendering due to data movement and manipulation. A novel volumetric data visualization scheme is proposed and implemented in this work that renders 2D images directly from compressed 3D data sets. The novelty of this algorithm is that rendering is performed on the compressed representation of the volumetric data without pre-decompression. As a result, the overheads associated with both data movement and rendering processing are significantly reduced. The proposed algorithm generalizes previously proposed whole-volume frequency-domain rendering schemes by first dividing the 3D data set into subcubes, transforming each subcube to a frequency-domain representation, and applying the Fourier projection theorem to produce the projected 2D images according to given viewing angles. Compared to the whole-volume approach, the subcube-based scheme not only achieves higher compression efficiency by exploiting local coherency, but also improves the quality of resultant rendering images because it approximates the occlusion effect on a subcube by subcube basis.
Tzi-cker Chiueh, Chuan-Kai Yang, Taosong He, Hanspeter Pfister, Arie E. Kaufman
IEEE Visualization2