Yi-Zeng Hsieh

dblp:95/2392 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Parallel Concatenated Feature Pyramid Network for Dehazing a Single Image on Smartphone Images
abstract
ABSTRACT Smartphones capturing images in outdoor environments are often affected by adverse weather conditions, resulting in low‐quality images. This paper introduces the Parallel Concatenated Feature Pyramid Network (C‐FPN) to address the challenge of dehazing single smartphone images. Dehazing a single image on smartphones is considered an ill‐posed problem. While the Feature Pyramid Network (FPN) is widely used in computer vision tasks, its feature extraction is limited by the max‐pooling operator. Furthermore, it cannot retain the hazy feature and restore the image at the same time, which fails to preserve critical hazy image features. Additionally, most existing methods struggle to balance preserving haze‐relevant information with effective image restoration. To address these limitations, this study proposes a novel parallel concatenated FP architecture that estimates atmosphere light and calculates transmission information on smartphones. The key contributions of this paper include (1) designing a parallel concatenated FP architecture capable of retrieving hazy features across various environments in deeper layers, (2) incorporating a concatenation structure to retain hazy information, enabling depth estimation and the generation of a transmission map, (3) using the transmission map as an input for a convolutional neural network with a dehazing loss function to calculate atmosphere light under different environments, and (4) implementing a skipping connection in the C‐FPN to retain essential features, facilitating an end‐to‐end learning structure. The proposed method demonstrates superior performance on the SOTS, NH‐HAZE 2, and synthetic hazy image indoor datasets. The PSNR/SSIM achieve 26.58/0.948, 26.28/0.966 and 17.15/0.761, respectively. In addition to dehazing, the method achieves excellent object detection performance.
Yu-Shiuan Tsai, Yi-Zeng Hsieh, Kai-en Lin, Pin-hsiang Wang
IET Image Process.2
2025 Integrating cycleGAN and BERT for Chinese text style transfer
Chien-Hsing Chou, Cheng-Hou Chou, Yi-Zeng Hsieh, Tzu-Shien Yang
Multim. Tools Appl.3
2025 Integrating self-organizing feature map with graph convolutional network for enhanced superpixel segmentation and feature extraction in non-Euclidean data structure
Yi-Zeng Hsieh, Chia-Hsuan Wu
Multim. Tools Appl.1
2025 Strumming in the Metaverse: A Deep-Learning-Enabled Virtual Air Guitar System in VR With Enhanced Chord Recognition and Simulated Pedal Effects
abstract
Virtual reality (VR) is increasingly capable and inexpensive, and VR devices have become indispensable in many domains, such as gaming, videoconferencing, education, and healthcare. VR has also been applied to music performance and learning. Virtual instruments, such as virtual pianos and drums, free users from the need to own physical forms of these (often bulky and expensive) instruments. VR devices enable users to enjoy music anytime and anywhere without constraints. Virtual concerts, including spatial audio simulations and reconstructions of historical performances, are becoming increasingly common. Previous studies have primarily examined virtual guitars in non-VR environments and air guitar chord recognition. However, systematic research on virtual air guitar systems in VR remains scarce. Virtual guitar games that are available on the market cannot recognize hand gestures accurately and thus cannot accurately identify the strumming patterns and chords played by the player. To overcome this problem, we propose a VR-based virtual air guitar system that can recognize 30 chords and various strumming techniques through deep learning and visual feedback. Employing a black-box approach, we combine WaveNet and FiLM to simulate electric guitar pedal effects with a knob difference loss mechanism, which simulates the turning of knobs on a guitar effects pedal, for enhanced accuracy.
Yi-Zeng Hsieh, Ji-Jie Lin, Mu-Chun Su, Wei-Jen Lin
IEEE Trans. Multim.1
2024 Blurred Facial Recognition Based on AdaFace
abstract
Many campuses in Taiwan feature open access, complicating the control and identification of individuals entering. This issue, along with emerging campus safety concerns, underscores the critical need for efficient security measures. The low resolution of surveillance cameras and the inefficiency of manual identification methods exacerbate this challenge. This study suggests the use of artificial intelligence, specifically the Adaface model with its unique Marginal-based loss functions, to enhance the identification process even with low-quality, blurred images. By automating the classification and processing of facial data–distinguishing between registered and unregistered faces–the system can perform real-time comparisons and logging. Our experiments with Adaface have demonstrated improved multi-face recognition capabilities and the ability to track unrecognized faces over time. Nonetheless, variations in scene context can still affect feature visibility and model stability.
Quan-Bin Zhang, Po-Yang Chi, Shang-Ze Lin, Chia-Xsuan Wu, Yi-Zeng Hsieh
COMPSAC5
2024 The development of assisted- visually impaired people robot in the indoor environment based on deep learning
Yi-Zeng Hsieh, Xiang-Long Ku, Shih-Syun Lin
Multim. Tools Appl.1
2023 Object Detection via Fisheye Camera
abstract
During the competition, several factors that could decrease the effectiveness of the training result was quickly identified, such as the lack of distortion of provided training data, the high similarity between multiple images, and the extreme imbalance of quantity between classes. Due to the short duration of the competition, we proposed six simple-to-conduct yet proven very effective methods: 1) data filtering by skipping neighboring images, 2) data filtering to reduce high quantity classes, 3) use other datasets to replenish total quantity and abundancy, 4) develop an algorithm to distort a plain image accurately, 5) Pretrain on selected MS-COCO. Ablation studies were conducted to find out the effect of each part of the data. And found that using a selected set of MS-COCO to pre-train could increase its effect. Our work is straightforward in terms of methods used, but the result of our evaluation shows that it is reasonably solid.
Yi-Zeng Hsieh, Hau-Ching Chen, Yi-Hung Yeh
MMAsia1
2022 Aerial face recognition and absolute distance estimation using drone and deep learning
Ying-Hung Pu, Po-Sheng Chiu, Yu-Shiuan Tsai, Meng-Tsung Liu, Yi-Zeng Hsieh, Shih-Syun Lin
J. Supercomput.5
2020 Development of a wearable guide device based on convolutional neural network for blind or visually impaired persons
Yi-Zeng Hsieh, Shih-Syun Lin, Fu-Xiong Xu
Multim. Tools Appl.1
2017 Music emotion recognition using PSO-based fuzzy hyper-rectangular composite neural networks
abstract
This study proposed a novel system for recognising emotional content in music, and the proposed system is based on particle swarm optimisation (PSO)‐based fuzzy hyper‐rectangular composite neural networks (PFHRCNNs), which integrates three computational intelligence tools, i.e. hyper‐rectangular composite neural networks (HRCNNs), fuzzy systems, and PSO. PFHRCNN is flexible to the complex data due to the fuzzy membership estimation, and an optimisation of the parameters is provided by PSO. First, raw features are extracted from each music clips. After feature extraction, a HRCNN is separately constructed for each class. Each trained HRCNN will result in a set of crisp rules. A problem associated with these generated crisp rules is that some of them are ineffective; therefore, a crisp rule is transformed into a fuzzy rule incorporated with a confidence factor. Next, PSO is adopted to simultaneously trim the rules, search a set of good confidence factors, and fine‐tune the locations of the selected hyper‐rectangles to increase their effectiveness. Finally, a PFHRCNN consisted of a set of fuzzy rules can be generated to recognise the emotion state of music. The experimental result shows that the proposed system has a good performance.
Yu-Hao Chin, Yi-Zeng Hsieh, Mu-Chun Su, Shu-Fang Lee, Miao-Wen Chen, Jia-Ching Wang
IET Signal Process.2
2016 The Effectiveness of e-Portfolios for Enhancing College Freshmen's Reflection and Aesthetic Literacy
abstract
This paper aims to explore the theories and interrelationships of the effectiveness of creating e-portfolios for enhancing the reflection and aesthetic literacy of first-year college students. First, the paper re-examined the reflection scale and the aesthetic literacy scale with a questionnaire survey of first-year university students in the central of Taiwan. Moreover, this paper justified the relationships between the reflection and aesthetic literacy scales and their dimensions. Furthermore, the research results show that the reflection is highly correlation to the depth of the digital files and aesthetic literacy is highly correlation to the visual presentation of the digital files. Both variables can suitably reflect the student’s characteristics in the e-portfolio., which provides a foundation for future research. Finally, the research attempted to encourage the school educational incorporating the use of e-portfolios in the 12-year compulsory education curricula.
Shih-yun Lu, Wei-Her Hsieh, Chih-Cheng Lo, Yi-Zeng Hsieh, Tsai-Cheng Chang
CSEDU (1)4
2016 A Q-learning-based swarm optimization algorithm for economic dispatch problem
Yi-Zeng Hsieh, Mu-Chun Su
Neural Comput. Appl.1
2014 To design an interactive learning system for child by integrating blocks with Kinect
abstract
In this study, an interactive block-building system named as e-Block system is developed for children to learn the concepts of geometric structures and space. First, the system displays a picture (e.g., car or house) of the target object intended for the child to assemble. The child then follows the instructions provided by the system and uses various blocks to build the object. After the child has completed the task, the system employs a pattern recognition algorithm to automatically compare the assembled object with the picture and determine whether the shape is identical. The experimental results show that the proposed system achieves high accuracy rate, and children in testing are enjoy this system and have more motivation to play with building blocks.
Ke-Wei Chen, Feng-Chih Hsu, Yi-Zeng Hsieh, Chien-Hsing Chou
EDUCON3
2014 A PSO-based rule extractor for medical diagnosis
Yi-Zeng Hsieh, Mu-Chun Su, Pa-Chun Wang
J. Biomed. Informatics1
2011 A SOMO-based approach to the operating room scheduling problem
Mu-Chun Su, Shih-Chang Lai, Pa-Chun Wang, Yi-Zeng Hsieh, Shih-Chieh Lin
Expert Syst. Appl.4
2010 A Robot-based Learning Companion for Storytelling
Yi-Zeng Hsieh, Mu-Chun Su, Sherry Y. Chen, Gow-Dong Chen, Shih-Chieh Lin
ICCE1