EDBT 2026 Demo / reviewers in the wild / expert
Chuan-Kai Yang
dblp:29/1787
· DBLP profile ↗
33ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-4782-3621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A real-time bird sound recognition app via deep learning techniquesabstractThis study presents a real-time, on-device bird sound recognition system developed using deep transfer learning and optimized for mobile deployment. A curated Xeno-canto corpus, an open-access repository of wildlife sound recordings contributed by citizen scientists worldwide, comprising 610 Taiwanese bird species was used to evaluate six deep learning architectures: Residual Network-18 (ResNet-18), Yet Another Mobile Network (YAMNet), Visual Geometry Group-like Network for Audio Classification (VGGish), Convolutional Neural Network–Long Short-Term Memory (CNN-LSTM), Attention-based Convolutional Neural Network (Attention-CNN), and a Deep Neural Network (DNN) baseline. All models were trained using class weighting, batch normalization, a dropout rate of 0.2, and targeted data augmentation, including pitch shifting (±2 semitones), time stretching (0.8–1.2 $$\times $$ ), and time shifting ( $$\le $$ 16,000 samples). Among these, ResNet-18 achieved the best balance between accuracy and computational efficiency, with an overall accuracy of 0.955, macro-precision of 0.95, macro-recall of 0.94, and macro-F1 of 0.945 across all 610 classes. The model performs inference in 25.9 milliseconds with only 3.03 megabytes of memory (approximately 795,000 parameters), outperforming heavier architectures such as VGGish (0.8975 accuracy, 42.2 milliseconds, 587 megabytes) while remaining competitive with compact alternatives like YAMNet (0.935 accuracy, 27.0 milliseconds, 10.19 megabytes). Furthermore, Gradient-weighted Class Activation Mapping (Grad-CAM) visualizations confirm that predictions are driven by species-specific temporal–spectral patterns rather than background noise. Converting the optimized model to TensorFlow Lite enables fully offline inference on Android devices, eliminating cloud latency and ensuring user privacy. Overall, this lightweight, high-accuracy framework offers a scalable and practical solution for real-time biodiversity monitoring and conservation research. Hailemariam Abebe Endalamaw, Chuan-Kai Yang, Cheng-Hung Hsu |
Multim. Tools Appl. | 2 |
| 2026 | Two-stage fine-tuning of HuBERT for multi-label bird species recognition in overlapping acoustic environmentsabstractAutomated recognition of bird species from audio is critical for biodiversity monitoring, yet it remains difficult in practice because field recordings often contain multiple birds vocalizing at the same time, strong environmental noise, and limited labeled data. Most existing systems either assume single-species recordings, require clean inputs, or depend on manually engineered preprocessing, such as source separation. This work introduces a novel two-stage fine-tuning framework that adapts a large self-supervised speech model (HuBERT) to the highly non-speech, polyphonic, multi-label setting of wild bird soundscapes. The proposed approach departs from conventional direct fine-tuning by using a two-stage curriculum. In Stage 1, HuBERT is fine-tuned on clean single-species recordings to learn discriminative, species-specific acoustic representations without interference. In Stage 2, the model is then transferred and further fine-tuned on synthetically constructed overlapping vocalizations, enabling it to generalize to real noisy soundscapes where multiple species co-occur. This two-stage adaptation strategy bridges the acoustic gap between human speech pretraining and avian bioacoustics, and allows robust multi-label prediction directly on overlapping audio without requiring explicit source separation. Extensive experiments on ten bird species show that the proposed two-stage HuBERT achieves an F1-score of 0.94 on overlapping recordings, outperforming (i) HuBERT variants trained only on clean or overlapping audio, and (ii) state-of-the-art CNN, RNN, graph-based, and transformer baselines reported in prior studies. These results demonstrate that two-stage self-supervised adaptation is an effective and scalable direction for real-time, multi-species bird monitoring in complex natural environments. Hailemariam Abebe Endalamaw, Chuan-Kai Yang |
Nat. Comput. | 2 |
| 2025 | Visualizing NBA information via storylines
Chuan-Kai Yang, Chiun-How Kao |
Comput. Graph. | 2 |
| 2025 | Water-sensitive paper detection and spray analysisabstractGlobal population has grown from 1 billion in the 19th century to 7.9 billion today, and is projected to reach 9.2 billion by 2050. To meet the growing demand for food, we need to increase crop yields. However, it is a challenging task to keep crop yields high all the time. In the process of food production, crops are particularly vulnerable to pests and diseases. Spraying pesticides is a common and effective way to control pests and diseases in modern agriculture. However, excessive use of pesticides can damage the ecological environment, and pesticide residues can also harm human health. Therefore, it is essential to be able to evaluate spraying effects conveniently and accurately. This paper adopts water-sensitive paper (WSP) as a spraying evaluation tool. Previous studies have also moved the task of evaluating spraying results to mobile devices. Although they can analyze the results in real time, there are many limitations, such as the need for the paper to be photographed vertically and the image to be clear and free of noise. Therefore, this paper proposes a paper image processing flow that can break through the limitations of previous studies.per uses a Faster R-CNN model to identify the paper in the image, uses Grab Cut to extract the paper image, then calibrates the paper to a standard size, removes shadows, and uses the Otsu threshold method and K-means clustering method to segment the spray droplet area. The bounding box MIoU of paper identification is 0.9766; the MIoU of the WSP is 0.9824. In the subsequent experiments, the spraying coverage percentage evaluation results of three different types of paper images, namely regular paper (Regular WSP), scanned paper (Scan WSP), and outdoor paper (Outdoor WSP), are presented. The results of droplet segmentation are also compared with those of the DropLeaf water-sensitive paper analysis App proposed in previous studies. In the study of water-sensitive paper, it was found that there is currently no source of water-sensitive paper database. Due to the lack of a suitable database, this paper provides a method for synthesizing paper images. By inputting a scanned paper (Scan WSP) as the input image, a synthesized paper image can be generated, which will benefit researchers who want to conduct different experiments on water-sensitive paper in the future. Chia-Lin Wu, Chuan-Kai Yang, Ji-Yang Lin |
Multim. Tools Appl. | 2 |
| 2025 | A WebGL-based 3D furniture modeling system via light-field descriptor and interactive force-directed visualizationabstractIn recent years, the demand for designing and modeling tools for non-professional users has been rising. Among the available tools and methods, customizing an item of furniture according to personal preferences has also become increasingly popular. However, for non-professional modeling users, it is still not easy to design a customized 3D model from scratch. Therefore, this paper proposes an online 3D furniture modeling system to assist users to quickly and conveniently generate an ideal 3D furniture model. This system makes use of existing furniture models. We first extract the features of a 3D model via the LightField Descriptor and perform similarity calculation. We then use the features of furniture data to form a Force-Directed Layout to show the similarity among models. Our system allows users to find and select rough furniture models quickly. And through a simple interactive modeling visualization interface, users can change the components of furniture models. In addition, our system provides suggestions for similar components, as well as component adjustment tools, so that users can easily replace and edit components. As a result, an ideal furniture model can be produced. A user study has been conducted to show the effectiveness of our proposed system. San-Chi Yeh, Chuan-Kai Yang, Ling-Chi Cheng |
Multim. Tools Appl. | 2 |
| 2023 | Video visualization via face and speaker clustering
Dehvari Mojiborrahman, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2022 | Privacy protection and beautification of cornea images
Chia-Lin Wu, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2022 | Realistic video generation for american sign language
Meng-Chen Xu, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2021 | An object recognition system based on convolutional neural networks and angular resolutions
Achmad Lukman, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2020 | A Multi-Person Selfie System via Augmented RealityabstractAbstract The limited length of a selfie stick always poses the problem of distortion in a selfie, in spite of the prevalence of selfie stick in recent years. We propose a technique, based on modifying existing augmented reality technology, to support the selfie of multiple persons, through properly aligning different photographing processes. It can be shown that our technique helps avoiding the common distortion drawback of using a selfie stick, and facilitates the composition process of a group photo. It can also be used to create some special effects, including creating an illusion of having multiple appearances of a person. Chuan-Kai Yang |
Comput. Graph. Forum | 2 |
| 2019 | Automatic generation of video navigation from Google Street View data with car detection and inpainting
Yuan-Bang Cheng, Chuan-Kai Yang, Guan-Chung Chang, Teng-Wen Chang |
Multim. Tools Appl. | 2 |
| 2017 | Using data visualization technique to detect sensitive information re-identification problem of real open dataset
Chiun-How Kao, Chih-Hung Hsieh, Yu-Feng Chu, Yu-Ting Kuang, Chuan-Kai Yang |
J. Syst. Archit. | 5 |
| 2016 | Integrated model/scene construction through context-based search, data-driven suggestion and component replacement
Po-An Chen, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2016 | Encryption domain content-based image retrieval and convolution through a block-based transformation algorithm
Jia-Kai Chou, Chuan-Kai Yang, Hsing-Ching Chang |
Multim. Tools Appl. | 2 |
| 2016 | Automatic hair extraction from 2D images
Chuan-Kai Yang, Chia-Ning Kuo |
Multim. Tools Appl. | 1 |
| 2016 | Obfuscated volume rendering
Jia-Kai Chou, Chuan-Kai Yang |
Vis. Comput. | 2 |
| 2015 | Video Object Retrieval by Trajectory and AppearanceabstractThe prevalence of video recording capability, either on surveillance systems or mobile devices, has contributed to the popularity of video data. As a result, video management has become relatively more important than before and, in particular, video retrieval has been one of the main issues in this regard. Traditional video retrieval systems take texts as the inputs to look for similar information from the title, annotation or embedded textual data of a video, in a way that is very similar to the keyword search adopted by a common search engine. However, the lack of visual information specification during a search often makes the result rather inaccurate or even useless. For this reason, video retrieval systems using images or videos as the inputs have also been proposed; nevertheless, the associated ambiguity and complexity have made the implementation of such systems relatively difficult and, therefore, those systems are not as successful as desired. To address this, in this paper, we propose to perform a video retrieval of a desired object through the inputs of its trajectory and/or appearance, together with the help of a 3-D graphical user interface for more intuitive interactions, so that more satisfactory results can be achieved. We firmly believe that such a framework could serve as the foundation for behavior analysis used in many surveillance systems. Yuan-Hao Lai, Chuan-Kai Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | A Study on Enhancing Timeline-Like Visualization with Verbal TextabstractThere has been a long-standing question of whether sound or audio annotations may assist in data visualization tasks. This paper presents our study focusing on enhancing timeline/storyline like visualizations with audio annotations. Timeline visualizations are widely used to illustrate interactions among different entities over time. This type of visualizations facilitate reviewing important activities and their associations. For a long timeline, however, it is difficult for the viewer to remember all the essential information found in the visualization process. Since hearing is another primary human sense for perceiving information, we conjecture that adding audio annotations to selected sections of a timeline visualization can help improve the viewer's recall of important events. We have designed a user study based on augmenting storyline visualizations with verbal text, and tested subjects with three different settings: visual cues only, verbal text only, and both. While our test results do not give a strong indication of the clear advantage of adding verbal text, the lessons learned in our study suggest directions for further study and will help us and others design audio-augmented visualization systems. Jia-Kai Chou, Isaac Liao, Kwan-Liu Ma, Chuan-Kai Yang |
CW | 4 |
| 2013 | Simulation of face/hairstyle swapping in photographs with skin texture synthesis
Jia-Kai Chou, Chuan-Kai Yang |
Multim. Tools Appl. | 2 |
| 2012 | Face-off: automatic alteration of facial features
Jia-Kai Chou, Chuan-Kai Yang, Sing-Dong Gong |
Multim. Tools Appl. | 2 |
| 2012 | Fast architecture prototyping through 3D collage
Chuan-Kai Yang, Ching-Yang Tsai |
Multim. Tools Appl. | 1 |
| 2011 | PaperVis: Literature Review Made EasyabstractAbstract Reviewing literatures for a certain research field is always important for academics. One could use Google‐like information seeking tools, but oftentimes he/she would end up obtaining too many possibly related papers, as well as the papers in the associated citation network. During such a process, a user may easily get lost after following a few links for searching or cross‐referencing. It is also difficult for the user to identify relevant/important papers from the resulting huge collection of papers. Our work, called PaperVis, endeavors to provide a user‐friendly interface to help users quickly grasp the intrinsic complex citation‐reference structures among a specific group of papers. We modify the existing Radial Space Filling (RSF) and Bullseye View techniques to arrange involved papers as a node‐link graph that better depicts the relationships among them while saving the screen space at the same time. PaperVis applies visual cues to present node attributes and their transitions among interactions, and it categorizes papers into semantically meaningful hierarchies to facilitate ensuing literature exploration. We conduct experiments on the InfoVis 2004 Contest Dataset to demonstrate the effectiveness of PaperVis. Jia-Kai Chou, Chuan-Kai Yang |
Comput. Graph. Forum | 2 |
| 2010 | Digital product transaction mechanism for electronic auction environmentabstractThe rapid development in electronic commerce and information technology drives the traditional physical product trading evolved to digital product trading. With the effect of the multi-agents system in the Internet environment and the promotions of Government, digital product industry grows fast. The authors proposed a digital product transaction mechanism for electronic auction in the multi-agents system environment. The research introduced a convenient platform to protect the privacies of both buyers and sellers, and track digital product further in an electronic auction environment. In addition, by using simple cryptography techniques supplemented with encryption, the authors ensure the security of information transactions, thereby providing a mechanism of safe and fair digital product electronic auction. Chih-Ta Yen, Tzong-Chen Wu, Ming-Huang Guo, Chuan-Kai Yang, Han-Chieh Chao |
IET Inf. Secur. | 4 |
| 2009 | Synthesizing solid particle textures via a visual hull algorithm
Hsing-Ching Chang, Chuan-Kai Yang, Jia-Wei Chiou, Shih-Hsien Liu |
Comput. Graph. | 2 |
| 2008 | An Optimal Algorithm for the Minimum Disc Cover Problem
Min-Te Sun, Chih-Wei Yi, Chuan-Kai Yang, Ten-Hwang Lai |
Algorithmica | 3 |
| 2008 | An interactive facial expression generation system
Chuan-Kai Yang, Wei-Ting Chiang |
Multim. Tools Appl. | 1 |
| 2008 | Realization of Seurat's pointillism via non-photorealistic rendering
Chuan-Kai Yang, Hui-Lin Yang |
Vis. Comput. | 1 |
| 2006 | Integration of volume decompression and out-of-core iso-surface extraction from irregular volume data
Chuan-Kai Yang, Tzi-cker Chiueh |
Vis. Comput. | 1 |
| 2005 | A Simple and Novel Seed-Set Finding Approach for Iso-Surface ExtractionabstractIso-surface extraction is one of the most important approaches for volume rendering, and iso-contouring is one of the most effective methods for iso-surface extraction. Unlike most other methods having their search domain to be the whole dataset, iso-contouring does its search only on a relatively small subset of the original data-set. This subset, called a seed-set, has the property that every iso-surface must intersect with it, and it could be built at the preprocessing time. When an iso-value is given at the run time, iso-contouring algorithm starts from the intersected cells in the seed-set, and gradually propagates to form the whole iso-surface. As smaller seed-sets offer less cell searching time, most existing iso-contouring algorithms concentrates on how to identify an optimal seed-set. In this paper, we propose a new and linear-time approach for seed-set construction. This presented algorithm could reduce the size of the generated seed-sets by up to one or two orders of magnitude, compared with other previously proposed fast (linear time) algorithms. Chiang-Han Hung, Chuan-Kai Yang |
EuroVis | 2 |
| 2000 | On-the-Fly rendering of losslessly compressed irregular volume dataabstractVery large irregular-grid data sets are represented as tetrahedral meshes and may incur significant disk I/O access overhead in the rendering process. An effective way to alleviate the disk I/O overhead associated with rendering a large tetrahedral mesh is to reduce the I/O bandwidth requirement through compression. Existing tetrahedral mesh compression algorithms focus only on compression efficiency and cannot be readily integrated into the mesh rendering process, and thus demand that a compressed tetrahedral mesh be decompressed before it can be rendered into a 2D image. This paper presents an integrated tetrahedral mesh compression and rendering algorithm called Gatun, which allows compressed tetrahedral meshes to be rendered incrementally as they are being decompressed, thus leading to an efficient irregular grid rendering pipeline. Both compression and rendering algorithms in Gatun exploit the same local connectivity information among adjacent tetrahedra, and thus can be tightly integrated into a unified implementation framework. Our tetrahedral compression algorithm is specifically designed to facilitate the integration with an irregular grid renderer without any compromise in compression efficiency. A unique performance advantage of Gatun is its ability to reduce the runtime memory footprint requirement by releasing memory allocated to tetrahedra as early as possible. Chuan-Kai Yang, Tulika Mitra, Tzi-cker Chiueh |
IEEE Visualization | 1 |
| 2000 | Zodiac: A history-based interactive video authoring system
Tzi-cker Chiueh, Tulika Mitra, Anindya Neogi, Chuan-Kai Yang |
Multim. Syst. | 4 |
| 1998 | Zodiac: A History-Based Interactive Video Authoring Systemabstractstreams with other media types such as audio is an es-Fasy-io-use audio[video authoting took play a cmcial Ta2ein moving multimedia pTogTamsfrom TeseaTchcu-Tiosify to main-stTeam applications.This papeT desckbes the design and implementation of an inteTactitre video authoTing system called Zodiac, which employs an innovative edit histoTy abstraction to suppoTt Sevemt unique editing jeafuTes notfoun~in existing commercial and Te-seaTch video editing systems.Zo disc pTOVideSUSeTSa concepti~ally clean and semantically poweTfil bran&ing history model of edit operations to oTganize the authoTingpTocess, and to navigate among the design aL teTnatives.In addition, by analyzing the edit histo~, Zodiac is able to Teiiably detect a composed stTeam5 shot and scene boznda~es, ~hi& faci~itate intemctive ~~ideobTowsiny.Zodiac ako featuTes a video object annotation capability that allows meTs to associate annotations to moving objects in a video sequence.The annotations themselves could be text, image, audio, OT ~'ideo.Zodiac is built on top of hf3fFS, a file system specifica~~ydesigned foT inteTac~~ve mu~~imedia de- velaprnent environments, and implements an internal bufieT manage~that suppoTk transparent lossless com pT~ssion/decompression.Shot/scene detection, video object annotation, and bu~eT management all exploit the edit historr information foT peTfomance optimization.A complete digitd video authoring environment must support two fundamental functions: capturing and generation of raw video cfips, and temporal arrangement of video segments with special inter-segment transition eflects.b addition, the abiity to synchronize video Permissionto make digilzl or hard copies of all or part of thrs \trork for personzl or cl= Tzi-cker Chiueh, Tulika Mitra, Anindya Neogi, Chuan-Kai Yang |
ACM Multimedia | 4 |
| 1997 | Integrated volume compression and visualizationabstractVolumetric data sets require enormous storage capacity even at moderate resolution levels. The excessive storage demands not only stress the capacity of the underlying storage and communications systems, but also seriously limit the speed of volume rendering due to data movement and manipulation. A novel volumetric data visualization scheme is proposed and implemented in this work that renders 2D images directly from compressed 3D data sets. The novelty of this algorithm is that rendering is performed on the compressed representation of the volumetric data without pre-decompression. As a result, the overheads associated with both data movement and rendering processing are significantly reduced. The proposed algorithm generalizes previously proposed whole-volume frequency-domain rendering schemes by first dividing the 3D data set into subcubes, transforming each subcube to a frequency-domain representation, and applying the Fourier projection theorem to produce the projected 2D images according to given viewing angles. Compared to the whole-volume approach, the subcube-based scheme not only achieves higher compression efficiency by exploiting local coherency, but also improves the quality of resultant rendering images because it approximates the occlusion effect on a subcube by subcube basis. Tzi-cker Chiueh, Chuan-Kai Yang, Taosong He, Hanspeter Pfister, Arie E. Kaufman |
IEEE Visualization | 2 |