Chee Siang Leow

dblp:232/2856 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0009-0008-1382-8962ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reference-Free Handwritten Japanese Character Generation via CLIP-Conditioned Diffusion Models
Koki Fujita, Hideaki Yajima, Chee Siang Leow, Hiromitsu Nishizaki
ICDAR (3)3
2025 Design and Training of a Sound Classification Model on a Resource-Constrained Edge Device for Fruit Theft Prevention
abstract
In recent years, fruit theft has become a serious social problem in Japan, where conventional surveillance measures are limited due to the unique challenges in agricultural areas, such as limited power supply and large areas. To address this, previous studies have developed battery-powered, microphone-based theft detection systems that operate on microcontrollers; however, achieving high classification accuracy under severe resource constraints remains a challenge. This paper presents a design and training methodology for a compact, high-accuracy sound classification model for fruit theft prevention. Our approach combines Depthwise Separable Convolution (DSC) for efficient feature extraction with Ensemble Distillation (EnD), which transfers knowledge from two teacher models with distinct feature extraction approaches. Experiments on a three-class task (environmental sounds, speech, and footsteps) demonstrated that the proposed model achieved an F1-score of 88.61%, an improvement of 12.00 percentage points over the baseline. The inference memory usage was limited to 49 kB, and including the entire system for processes other than deep learning inference, the total memory usage remained within 160 kB, well below the 256 kB RAM capacity of the target hardware. Moreover, the model achieved an inference time of approximately 0.4519 seconds per 1-second audio segment on a Seeeduino XIAO nRF52840 micro-controller. These results confirm the practical feasibility of the proposed system as a robust AI-based surveillance solution under strict resource constraints for real-world orchard applications.
Haruki Endo, Hideaki Yajima, Chee Siang Leow, Tsutomu Tanzawa, Koji Makino, Kazuyoshi Ishida, Hiromitsu Nishizaki
IECON3
2025 L2H-NeRF: low- to high-frequency-guided NeRF for 3D reconstruction with a few input scenes
Taoqi Bao, Jiangnan Ye 0002, Zhankong Bao, Chee Siang Leow, Haoji Hu, Issei Fujishiro
Vis. Comput.4
2025 Non-invasive estimation of Shine Muscat grape color and sensory evaluation from standard camera images
abstract
Abstract This study proposes a non-invasive method to estimate both color and sensory attributes of Shine Muscat grapes from standard camera images. First, we focus on color estimation by integrating a Vision Transformer (ViT) feature extractor with interquartile range (IQR)-based outlier removal. Experimental results show that our approach achieves 97.2% accuracy, significantly outperforming Convolutional Neural Network (CNN) models. This improvement underscores the importance of capturing global contextual information to differentiate subtle color variations in grape ripeness. Second, we address human sensory evaluation by collecting questionnaire responses on 13 attributes (e.g., “Sweetness,” “Overall taste rating”), each rated on a five-point scale. Because these ratings tend to cluster around midrange values (labels “2,” “3,” and “4”), we initially limit the dataset to the extreme labels “1” (“lowest grade”) and “5” (“highest grade”) for binary classification. Three attributes—“Overall color,” “Sweetness,” and “Overall taste rating”—exhibit relatively high classification accuracies of 79.9%, 75.1%, and 75.7%, respectively. By contrast, the other 10 attributes reach only 50%–66%, suggesting that subjective variations and limited visual cues pose significant challenges. Overall, the proposed approach demonstrates the feasibility of an image-based system that integrates color estimation and sensory evaluation to support more objective, data-driven harvest timing decisions for Shine Muscat grapes.
Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Xiaoyang Mao, Wan-Young Chung, Hiromitsu Nishizaki
Vis. Comput.2
2024 High Quality Color Estimation of Shine Muscat Grape Using Vision Transformer
abstract
Currently, skilled farmers judge the ripeness of the Shine Muscat grape variety by looking at the color on the surface of the grapes. However, the color of Shine Muscat grapes does not change much as they grow, and there are individual differences in the way the color is perceived. Furthermore, the same color can look very different depending on the exposure to sunlight and shadows. Therefore, there is a need for a system that can quantitatively determine the color of Shine Muscat grapes to pass on the harvesting techniques of experienced farmers to amateurs and inexperienced farmers. This research aims to improve the accuracy of the color estimation of Shine Muscat grapes using deep learning. We propose a method to estimate the color of individual grapes using a color estimation model with a self-attention mechanism, from which the color of the whole bunch is estimated. A Vision Transformer model with a self-attention mechanism was found to improve the color estimation accuracy to $96.9 \%$. Furthermore, by eliminating outliers using the interquartile range, a color estimation accuracy of $97.2 \%$ could be achieved, demonstrating the effectiveness of the new color estimation model.
Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki
CW2
2024 Development of a Fruit Theft Reporting System Using a Compact Microcontroller with Deep Learning Based on Suspicious Sounds
abstract
In Japanese fruit-growing regions, fruit theft is a significant issue. Traditional anti-theft methods are fraught with various shortcomings and are often ineffective. To address this, we have developed a new approach for preventing fruit theft: a device that detects suspicious sounds by integrating a sound sensor with a compact, low-power microcontroller. This paper presents a suspicious sound detection system designed for a fruit theft alert device. It details the development of a deep learning model for detecting suspicious sounds, covering aspects from data collection to model training and evaluation. Despite operating on a microcontroller with limited memory and computational power, the system under development has achieved a classification accuracy of up to 78.5% in F1-score for footsteps, speech, and other environmental sounds.
Chee Siang Leow, Tsutomu Tanzawa, Tze Yaw Bong, Koji Makino, Kazuyoshi Ishida, Hiromitsu Nishizaki
IECON1
2022 Handwritten Character Generation using Y-Autoencoder for Character Recognition Model Training
abstract
It is well-known that the deep learning-based optical character recognition (OCR) system needs a large amount of data to train a high-performance character recognizer. However, it is costly to collect a large amount of realistic handwritten characters. This paper introduces a Y-Autoencoder (Y-AE)-based handwritten character generator to generate multiple Japanese Hiragana characters with a single image to increase the amount of data for training a handwritten character recognizer. The adaptive instance normalization (AdaIN) layer allows the generator to be trained and generate handwritten character images without paired-character image labels. The experiment shows that the Y-AE could generate Japanese character images then used to train the handwritten character recognizer, producing an F1-score improved from 0.8664 to 0.9281. We further analyzed the usefulness of the Y-AE-based generator with shape images, out-of-character (OOC) images, which have different character images styles in model training. The result showed that the generator could generate a handwritten image with a similar style to that of the input character.
Tomoki Kitagawa, Chee Siang Leow, Hiromitsu Nishizaki
LREC2
2022 Appropriate grape color estimation based on metric learning for judging harvest timing
abstract
Abstract The color of a bunch of grapes is a very important factor when determining the appropriate time for harvesting. However, judging whether the color of the bunch is appropriate for harvesting requires experience and the result can vary by individuals. In this paper, we describe a system to support grape harvesting based on color estimation using deep learning. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of pest grains, and color estimation are required, for which deep learning-based approaches are adopted. In this study, YOLOv5, an object detection model that considers both accuracy and processing speed, is adopted for bunch detection and grain detection. For the detection of diseased grains, an autoencoder-based anomaly detection model is also employed. Since color is strongly affected by brightness, a color estimation model that is less affected by this factor is required. Accordingly, we propose multitask learning that uses metric learning. The color estimation model in this study is based on AlexNet. Metric learning was applied to train this model. Brightness is an important factor affecting the perception of color. In a practical experiment using actual grapes, we empirically selected the best three image channels from RGB and CIELAB (L*a*b*) color spaces and we found that the color estimation accuracy of the proposed multi-task model, the combination with “L” channel from L*a*b color space and “GB” from RGB color space for the grape image (represented as “LGB” color space), was 72.1%, compared to 21.1% for the model which used the normal RGB image. In addition, it was found that the proposed system was able to determine the suitability of grapes for harvesting with an accuracy of 81.6%, demonstrating the effectiveness of the proposed system.
Tatsuyoshi Amemiya, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki
Vis. Comput.2
2021 Development of a Support System for Judging the Appropriate Timing for Grape Harvesting
abstract
The color of grape bunches is a significant factor when harvesting grapes at the appropriate timing. Judging the suitable color for shipment requires experience and varies from one person to another. We herein describe a support system for grape harvesting based on color estimation. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of diseased grains, and color estimation should be performed. Models based on deep learning are employed for this series of processes. Since color is strongly affected by sunlight, we propose a multitask model that considers sunlight exposure to achieve a robust color estimation model that exhibits decreased sensitivity to sunlight. Our results show that the color estimation accuracy of the model is 76% when sunlight exposure is not considered and 81% when sunlight exposure is considered. In addition, we performed a practical field test of the developed harvest support system in an actual grape field. The results show that our support system can determine the appropriateness of grape harvest with an accuracy of 90%, demonstrating the effectiveness of the system.
Tatsuyoshi Amemiya, Kodai Akiyama, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki
CW3
2021 End-to-End Inflorescence Measurement for Supporting Table Grape Trimming with Augmented Reality
abstract
Inflorescence trimming is a crucial process to produce high-quality table grapes. It can eliminate nutrient competition in a bunch and makes it less vulnerable to disease development. After trimming, the remaining part of the inflorescence should have a target length decided by the grape variety. This is challenging for novice farmers because of the time constraint. The farmer needs to finish trimming the inflorescence before the berries develop. This paper proposes a novel end-to-end inflorescence length measurement method for supporting a trimming process with augmented reality technology. The proposed technique makes use of the state-of-the-art deep neural network model for detecting the inflorescence area, as well as the scissors from the images captured with a camera installed on an optical see-through head-mounted display. A new algorithm is designed to estimate the length of the remaining inflorescence with the screw of the scissors loop as the calibrator. The estimated length is then visualized on the head-mounted display to support the farmer in performing the trimming correctly and efficiently. The experiment, conducted with real inflorescence trimming tasks, shows that the mean absolute error of the length estimation is only 0.19 cm, which is small enough for use in real applications.
Prawit Buayai, Kabin Yok-In, Daisuke Inoue 0004, Chee Siang Leow, Hiromitsu Nishizaki, Koji Makino, Xiaoyang Mao
CW4
2021 Language and Speaker-Independent Feature Transformation for End-to-End Multilingual Speech Recognition
Tomoaki Hayakawa, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki
Interspeech2
2021 Voice Activity Detection for Live Speech of Baseball Game Based on Tandem Connection with Speech/Noise Separation Model
Yuto Nonaka, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki
Interspeech2