Yen-Lin Chen

dblp:15/4143 · DBLP profile ↗
← Back
28ranked-venue papers
13as first author
6since 2021 · last 2026
0000-0001-7717-9393ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2026 MonoTDF: Temporal deep feature learning for generalizable monocular 3D object detection
abstract
Monocular 3D object detection has gained significant attention due to its cost-effectiveness and practicality in real-world applications. However, existing monocular methods often struggle with depth estimation and spatial consistency, limiting their accuracy in complex environments. In this work, we introduce a Temporal Deep Feature Learning framework, which enhances monocular 3D object detection by integrating temporal features across sequential frames. Our approach leverages a novel deep feature auxiliary module based on convolutional recurrent structures, effectively capturing spatiotemporal information to improve depth perception and detection robustness. The proposed module is model-agnostic and can be seamlessly integrated into various existing monocular detection frameworks. Extensive experiments across multiple state-of-the-art monocular 3D object detection models demonstrate consistent performance improvements, particularly in detecting small or partially occluded objects. Our results highlight the effectiveness and generalizability of the proposed approach, making it a promising solution for real-world autonomous perception systems. The source code of this work is at: https://github.com/Shuray36/MonoTDF-Temporal-Deep-Feature-Learning-for-Generalizable-Monocular-3D-Object-Detection .
Xiu-Zhi Chen, Yi-Kai Chiu, Chih-Sheng Huang, Yen-Lin Chen
Pattern Recognit.4
2025 Self-Attention Enhanced Deep Learning Models for Immune Cell Deconvolution from Bulk RNA-Seq
abstract
Accurate immune cell composition profiling is crucial for understanding immunological dynamics and disease mechanisms. Bulk RNA sequencing (bulk RNA-seq) is widely employed due to its cost-effectiveness and scalability; however, it lacks the resolution to identify cell-specific gene expression. To address this limitation, we propose a self-attention enhanced deep learning model designed for precise immune cell deconvolution from bulk RNA-seq data. We systematically annotated immune cell types from four single-cell RNA-seq (scRNA-seq) peripheral blood mononuclear cell (PBMC) datasets and validated these annotations against established automated identification tools (SingleR, Seurat, scPred, ScType). Leveraging these annotations, we generated realistic pseudo-bulk RNA-seq training samples using Dirichlet-distribution-based composition sampling, significantly enhancing the model’s performance, particularly for rare cell populations. Comparative evaluations demonstrated that our self-attention enhanced deep learning model consistently outperformed existing approaches, including CIBERSORTx and Scaden, achieving lower prediction errors and higher correlations on benchmark PBMC datasets. Integrating multi-head self-attention allowed the model to dynamically capture intricate dependencies among gene expression features, substantially improving deconvolution accuracy for specific cell subsets. While demonstrating robust performance on PBMC datasets, we acknowledge that broader validation is essential due to potential limitations in generalizability across different tissue types and conditions. Our study highlights the potential of self-attention mechanisms and realistic training data generation strategies to enhance computational deconvolution techniques, providing valuable tools for clinical diagnostics and translational immunology research.
Chia-Ru Chung, Yen-Lin Chen, Justin Bo-Kai Hsu, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng
CIBCB2
2025 MPPQNet: A Moment-Preserving Product Quantization Neural Network for Progressive 3D Point Cloud Transmission
Shyi-Chyi Cheng, Yen-Lin Chen, Shih-Yu Li
MMM (3)2
2025 Improving Model Flexibility of Electrical Characteristic Prediction of a-IGZO-TFTs
abstract
This paper presents a lightweight and flexible artificial neural network (ANN) model for predicting the transfer characteristics of amorphous Indium Gallium Zinc Oxide (a-IGZO) thin-film transistors (TFTs). Traditional Technology Computer-Aided Design (TCAD) simulations, while accurate, are computationally intensive and inflexible to rapid design variations. Prior efforts using variational autoencoders (VAEs) showed promise but were limited by rigid input formats that required retraining for any changes in curve dimension or voltage range. To overcome this, we propose an ANN architecture that treats gate voltage as a dynamic input, enabling continuous and accurate predictions across a wide voltage spectrum without the need for retraining. Experimental evaluation through 5-fold cross-validation confirms the ANN’s competitive performance, achieving an average ℝ2score of 0.9898, outperforming VAE models in flexibility and robustness. This approach offers a fast, accurate, and hardware-efficient alternative for a-IGZO TFT modeling and optimization, facilitating the monolithic 3D (M3D) integration and the next generation of display.
Khean Thye Bea, Jo-An Liao, Yen-Ting Chen, Hsin-Hui Hu, Yen-Lin Chen, Wai-Khuen Cheng, Kun-Ming Chen
SMC5
2025 An Auto-Labeling tool for Occupancy Grid and BEV in Autonomous Driving Dataset
abstract
Accurate 3D semantic labeling is critical for scene understanding and navigation in autonomous driving systems. However, manual annotation of 3D point clouds is labor-intensive and difficult to scale across diverse environments. To address this challenge, we propose a modular-based automated labeling framework that leverages 2D semantic cues to facilitate 3D scene understanding. The system integrates state-of-the-art models—Grounding DINO, SAM2, YOLOv7, and Generative AI (GAI) —to perform high-quality 2D semantic segmentation, which is then processing into semantic LiDAR point clouds. Afterward, the framework employs multi-frame fusion and static-dynamic scene separation to construct dense 3D semantic occupancy grids. Additionally, GAI serves as a prompt-refinement module, improving segmentation accuracy at object boundaries and in distant regions. Experiments on real-world street-view datasets demonstrate that AutoLabel-2D significantly enhances labeling efficiency and segmentation completeness, offering strong generalization across scenes. This framework provides a scalable and effective solution for high-definition mapping and semantic perception in autonomous driving applications.
Khean Thye Bea, Yu Chen Yang, Shih Chi Tseng, Chieh-Sheng Huang, Yen-Lin Chen
SMC5
2023 Prediction of Electrical Characteristics of a-IGZO TFT Based on Transfer Learning-Based Variational Autoencoder
abstract
In this study, we proposed a transfer-learning based variational autoencoder model for predicting the electrical characteristics in the parameter tuning process of a-IGZO TFT structure design. The result achieve a high R2 score of 0.9704 with a low-computing-power hardware-friendly method that reduced time consumption significantly compared to prior approaches. The findings have practical implications for mitigating the time-consuming nature of TCAD simulations, and the method can expand to various types of input data while ensuring high performance and generalization. We demonstrated significant improvement in generalization and accuracy through a k-fold validation.
Khean Thye Bea, Shih-Shin Hu, Wei-Hsuan Lin, Da Zheng Lin, Hsin-Hui Hu, Xiu-Zhi Chen, Ting-Ru Lin, Yen-Lin Chen, Kun-Ming Chen, Wai-Khuen Cheng
SMC8
2018 Applying Machine Learning Concept to Provide Adaptable Digital Tour Guide System
Kai-Yi Chin, Ko-Fong Lee, Ya-Chuan Kao, Yen-Lin Chen
ICCE4
2016 Edge Snapping-Based Depth Enhancement for Dynamic Occlusion Handling in Augmented Reality
abstract
Dynamic occlusion handling is critical for correct depth perception in Augmented Reality (AR) applications. Consequently it is a key component to ensure realistic and immersive AR experiences. Existing solutions to tackle this challenge typically suffer from various limitations, e.g. assumption of a static scene or high computational complexity. In this work, we propose an algorithm for depth map enhancement for dynamic occlusion handling in AR applications. The key of our algorithm is an edge snapping approach, formulated as discrete optimization, that improves the consistency of object boundaries between RGB and depth data. The optimization problem is solved efficiently via dynamic programming and our system runs in near real-time on the tablet platform. Experimental evaluations demonstrate that our approach largely improves the raw sensor data and is particularly suitable compared to several related approaches in terms of both speed and quality. Furthermore, we demonstrate visually pleasing dynamic occlusion effects for multiple AR use cases based on our edge snapping results.
Yen-Lin Chen, Mao Ye 0005, Liu Ren 0001
ISMAR2
2016 Energy-efficient video decoding schemes for embedded handheld devices
Yen-Lin Chen, Ming-Feng Chang, Wen-Yew Liang
Multim. Tools Appl.1
2014 Real-time eye detection and event identification for human-computer interactive control for driver assistance
abstract
Eye movements can provide important information for human-computer interactive applications. Due to the progress of computer technology, the detecting accuracy and speed of pattern recognition are promoted. Additionally, according to research advances in embedded systems, the applications of digital cameras, such as internet cameras, smart phones and smart TV, and smart cars, are widely used in nowadays. Therefore, we propose a set of real-time human-eye detection and tracking systems with human-computer interaction applications. This technique can obtain eye movements and can be adopted as interactive control commands on driver assistance systems. The proposed system is implemented on an OMAP4430 for embedded applications, and experimental results show that the proposed architecture is capable of effective and real-time eye position detection and event identification for human-computer interactive applications on driver assistance systems.
Yen-Lin Chen, Chao-Wei Yu, Chuan-Yen Chiang, Chin-Hsuan Liu, Wei-Chen Sun, Hsin-Han Chiang, Tsu-Tian Lee
SMC1
2014 Data broadcasting for dependent information using multiple channels in wireless broadcast environments
Ta-Chih Su, Jenq-Haur Wang, Yen-Lin Chen
J. Parallel Distributed Comput.4
2013 Accurate and Robust 3D Facial Capture Using a Single RGBD Camera
abstract
This paper presents an automatic and robust approach that accurately captures high-quality 3D facial performances using a single RGBD camera. The key of our approach is to combine the power of automatic facial feature detection and image-based 3D nonrigid registration techniques for 3D facial reconstruction. In particular, we develop a robust and accurate image-based nonrigid registration algorithm that incrementally deforms a 3D template mesh model to best match observed depth image data and important facial features detected from single RGBD images. The whole process is fully automatic and robust because it is based on single frame facial registration framework. The system is flexible because it does not require any strong 3D facial priors such as blend shape models. We demonstrate the power of our approach by capturing a wide range of 3D facial expressions using a single RGBD camera and achieve state-of-the-art accuracy by comparing against alternative methods.
Yen-Lin Chen, Hsiang-Tao Wu, Fuhao Shi, Xin Tong 0001, Jinxiang Chai
ICCV1
2012 Development of hand-cleaning service-oriented autonomous navigation robot
abstract
This paper proposes the development of an autonomous navigation robot with hand-cleaning service in indoor environments. To navigate in unknown environments and provide service, the robot is with several intelligent behaviors including wall-following, obstacle avoidance, autonomous navigation, and human detection. A laser-sensor-based approach is used in the wall-following and obstacle avoidance behavior controllers. A preliminary map-matching algorithm is applied in the localization strategy of autonomous navigation in which the robot can acquire the current location and then move toward to the target position. In this study a hand-cleaning mechanism is embedded into the robot and the service will activate while a human is recognized within the designated range. The overall robotic system is carried out using a two-wheeled driving mobile robot with LabVIEW as an integration tool. The experimental results demonstrate the practicable application of the proposed approach.
Chia-Long Chan, Hsin-Han Chiang, Yen-Lin Chen, Geng-Yen Chen, Tsu-Tian Lee
SMC3
2012 A knowledge-based system for extracting text-lines from mixed and overlapping text/graphics compound document images
Yen-Lin Chen, Zeng-Wei Hong, Cheng-Hung Chuang
Expert Syst. Appl.1
2011 Developing Ubiquitous Multi-touch Sensing and Displaying Systems with Vision-Based Finger Detection and Event Identification Techniques
abstract
This study presents efficient vision-based finger detection, tracking, and event identification techniques, as well as a low-cost hardware framework for multi-touch sensing and display applications. A fast bright-blob segmentation process based on automatic multilevel histogram thresholding is performed to extract pixels of touch blobs from the captured image sequences obtained from the scattered infrared lights by the video camera. Given the touch blobs extracted from each of the captured frames, a blob tracking and event recognition process is then conducted to analyze the spatial and temporal information of these touch blobs from consecutive frames and determine the possible touch events issued by users. This process also refines the detection results and corrects for errors and occlusions caused by noise and errors during the blob extraction processes. Our proposed blob tracking and touch event recognition process includes two phases. First, the phase of blob tracking associates the motion correspondence of blobs in succeeding frames by analyzing their spatial and temporal features. Then the phase of touch event recognition process can identify meaningful touch events activated by users from the motion information of touch blobs. Experimental results demonstrate that the proposed vision-based finger detection, tracking, and event identification system is feasible and effective for multi-touch sensing applications in various operational environments and conditions.
Yen-Lin Chen, Chuan-Yen Chiang, Wen-Yew Liang, Tung-Ju Hsieh, Da-Cheng Lee, Shyan-Ming Yuan, Yang-Lang Chang
HPCC1
2011 Cyclic twill-woven objects
Ergun Akleman, Jianer Chen, Yen-Lin Chen, Qing Xing, Jonathan L. Gross
Comput. Graph.3
2011 A fuzzy hierarchy integral analytic expert decision process in evaluating foreign investment entry mode selection for Taiwanese bio-tech firms
Hsu-Hua Lee, Tsau-Tang Yang, Chie-Bein Chen, Yen-Lin Chen
Expert Syst. Appl.4
2010 A knowledge-based approach for textual information extraction from mixed text/graphics complex document images
abstract
A new knowledge-based technique for extracting and identifying text-lines from various real-life mixed text/graphics complex document images is presented in this paper. The proposed technique first decompose the document image into distinct object planes to separate homogeneous objects including textual regions of interest, non-text objects such as graphics and pictures, and background textures. Then a knowledge-based text extraction and identification method is performed on the resultant planes to obtain text-lines with different characteristics in each plane. This proposed system can offer high flexibility and expandability by just updating new rules for coping with more various types of real-life and future complex document images. From the experimental and comparative results, the proposed knowledge-based technique demonstrates its effectiveness and advantages on extracting text-lines with various illuminations, sizes, and font styles from various types of mixed text/graphics complex document images.
Yen-Lin Chen
SMC1
2010 Embedded on-road nighttime vehicle detection and tracking system for driver assistance
abstract
This study presents an effective method for detecting vehicles in front of the camera-assisted car during nighttime driving and implements it on an embedded system. The proposed method detects vehicles based on detecting and locating vehicle headlights and taillights using techniques of image segmentation and pattern analysis. Firstly, to effectively extract bright objects of interest, a segmentation process based on automatic multilevel thresholding applied on the grabbed road-scene images. Then the extracted bright objects are processed by to identify and tracking the vehicles by locating and analyzing the spatial and temporal features of vehicle light patterns and to estimate their distances to the camera-assisted car. Finally, we also implement the above vision-based techniques on a real-time system mounted in the host car. The proposed vision-based techniques are integrated and implemented on an ARM-Linux embedded platform, as well as the peripheral devices, including image grabbing devices, voice reporting module, and other in-vehicle control devices, will be also integrated to accomplish an in-vehicle embedded vision-based nighttime driver assistance system
Yen-Lin Chen, Chuan-Yen Chiang
SMC1
2009 3D Reconstruction of Human Motion and Skeleton from Uncalibrated Monocular Video
Yen-Lin Chen, Jinxiang Chai
ACCV (1)1
2009 Flexible registration of human motion data with parameterized motion models
abstract
This paper presents an efficient model-based approach for automatic human motion registration, which builds temporal correspondences between structurally similar but distinctive motion examples. The key idea of the model-based registration process is to construct a parameterized motion model from a set of preregistered motion examples. With such a model, we can register an input motion with the parameterized motion model by continuously deforming the model to best match the input motion. We formulate the registration process in a gradient-based nonlinear optimization framework by minimizing an objective function that measures differences between the input motion and deforming motion. We also develop a multi-resolution optimization process to efficiently estimate the model parameters as well as the temporal correspondences between the input motion and deforming motion. We demonstrate the performance of our approach by testing the algorithm on difficult motion sequences and comparing with alternative approaches.
Yen-Lin Chen, Jianyuan Min, Jinxiang Chai
SI3D1
2009 Real-time Vision-based Multiple Vehicle Detection and Tracking for Nighttime Traffic Surveillance
abstract
This study presents an effective system for detecting and tracking moving vehicles in nighttime traffic scene for traffic surveillance. The proposed method identifies vehicles based on detecting and locating vehicle headlights and taillights by using the techniques of image segmentation and pattern analysis. First, to effectively extract bright objects of interest, a fast bright-object segmentation process based on automatic multilevel histogram thresholding is applied on the nighttime road-scene images. This automatic multilevel thresholding approach can provide robustness and adaptability for the detection system to be operated well under various illumination conditions at night. The extracted bright objects are processed by a spatial clustering and tracking procedure by locating and analyzing the spatial and temporal features of vehicle light patterns, and then identifying and classifying the moving cars and motorbikes in the traffic scenes. Experimental results demonstrate that the proposed approach is feasible and effective for vehicle detection and identification in various nighttime environments for traffic surveillance.
Yen-Lin Chen, Bing-Fei Wu, Chung-Jui Fan
SMC1
2009 A collaborative desktop tagging system for group knowledge management based on concept space
Wen-Tai Hsieh, Jay Stu, Yen-Lin Chen, Seng-cho Timothy Chou
Expert Syst. Appl.3
2009 A multi-plane approach for text segmentation of complex document images
Yen-Lin Chen, Bing-Fei Wu
Pattern Recognit.1
2009 Interactive generation of human animation with deformable motion models
abstract
This article presents a new motion model deformable motion models for human motion modeling and synthesis. Our key idea is to apply statistical analysis techniques to a set of precaptured human motion data and construct a low-dimensional deformable motion model of the form x = M (α, γ), where the deformable parameters α and γ control the motion's geometric and timing variations, respectively. To generate a desired animation, we continuously adjust the deformable parameters' values to match various forms of user-specified constraints. Mathematically, we formulate the constraint-based motion synthesis problem in a Maximum A Posteriori (MAP) framework by estimating the most likely deformable parameters from the user's input. We demonstrate the power and flexibility of our approach by exploring two interactive and easy-to-use interfaces for human motion generation: direct manipulation interfaces and sketching interfaces.
Jianyuan Min, Yen-Lin Chen, Jinxiang Chai
ACM Trans. Graph.2
2008 Vision-based nighttime vehicle detection and range estimation for driver assistance
abstract
This paper presents a real-time vision system for assisting driver during nighttime driving. The proposed system provides the following features: 1) effectively detection and tracking of oncoming and preceding vehicles based on image segmentation and pattern analysis techniques. 2) Robust and adaptive vehicle detection under various illuminated conditions at nighttime urban environments benefited by a novel automatic object segmentation scheme. 3) Providing beneficial information for assisting the driver to perceive surrounding traffic conditions outside the car during nighttime driving. 4) Providing a versatile control strategy for in-vehicle facilities of the autonomous vehicles. 5) Offering real-time traffic event-driven video surveillance machinery for recording evidences of possible traffic accidents. Experimental results demonstrate the feasibility and effectiveness of the proposed system on nighttime driver assistance issues.
Yen-Lin Chen, Chuan-Tsai Lin, Chung-Jui Fan, Chih-Ming Hsieh, Bing-Fei Wu
SMC1
2006 Text Extraction from Complex Document Images Using the Multi-plane Segmentation Technique
abstract
This study presents a new method for extracting characters from various real-life complex document images. The proposed method applies a multi-plane segmentation technique to separate homogeneous objects including text blocks, non-text graphical objects, and background textures into individual object planes. It consists of two stages-automatic localized multilevel thresholding, and multi-plane region matching and assembling. Then a text extraction process can be performed on the resultant planes to detect and extract characters with different characteristics in the respective planes. The proposed method processes document images regionally and adaptively according to their respective local features. This allows preservation of detailed characteristics from extracted characters, especially small characters with thin strokes, as well as gradational illuminations of characters. This also permits background objects with uneven, gradational, and sharp variations in contrast, illumination, and texture to be handled easily and well. Experimental results on real-life complex document images demonstrate that the proposed method is effective in extracting characters with various illuminations, sizes, and font styles from various types of complex document images.
Yen-Lin Chen, Bing-Fei Wu
SMC1
2005 Multi-layer segmentation of complex document images
abstract
Text is commonly printed on a complex background. Segmenting text is an important part in document analysis. In the past some methods have been shown for the segmentation of texts with images. However, previous studies have not sufficiently addressed complex compound documents. This investigation presents an algorithm for the segmentation of text in various document images. The proposed segmentation algorithm applies a new multilayer segmentation method to separate the text from various compound document images, independent from the text and background overlapping or not. This method solves various problems associated with the complexity of background images. Experimental results obtained using various document images scanned from book covers, advertisements, brochures and magazines, reveal that the proposed algorithm can successfully segment Chinese and English text strings from various backgrounds, regardless of whether the texts are over a simple, slowly varying or rapidly varying background texture.
Bing-Fei Wu, Yen-Lin Chen, Chung-Cheng Chiu
Int. J. Pattern Recognit. Artif. Intell.2