VLDB 2026 Research / reviewers in the wild / expert
Yung-Yao Chen
dblp:39/10697
· DBLP profile ↗
22ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-6852-8862ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D Multi-Modal Object Detection Based on Cross-Attention Feature FusionabstractIn Advanced Driver Assistance Systems (ADAS), environmental perception and object detection are crucial for ensuring safe autonomous driving. Single-modality systems often struggle under adverse weather conditions, underscoring the need for multi-modal approaches. Current fusion methods typically rely on simplistic concatenation of multi-modal features, which neglects semantic alignment and does not fully exploit inter-modal correlations. This paper proposes a crossattention feature fusion specifically designed to enhance the global correlation between camera and radar features. By dynamically adjusting feature weights through cross-attention, our approach significantly improves feature integration. Furthermore, we propose a depth-weighted voting fusion strategy to select the most accurate sensor depth, thereby enhancing decision-making stability. Experimental results on the nuScenes dataset show substantial improvements, with mean Average Precision (mAP) of 0.399 and mean Average Translation Error (mATE) of 0.602, highlighting the effectiveness of our approach in enhancing the robustness and accuracy of multi-modal fusion. Sin-Ye Jhong, Min-Hsuan Ho, Si-Yu Lu, Yung-Yao Chen |
ICRA | 4 |
| 2025 | Hierarchical Spatiotemporal Fusion for Event-Visible Object DetectionabstractTraditional visible light cameras are prone to performance degradation under varying weather and lighting conditions. To address this challenge, we introduce an eventbased camera and propose a novel hierarchical spatiotemporal fusion approach for event-visible object detection. Our method enhances detection performance by integrating data from both event-based and visible light cameras. We have designed three key modules: The Gated Event Accumulation Representation module (GEAR), the Temporal Feature Selection module (TFS), and the Adaptive Fusion module (AF). GEAR and TFS enhance temporal feature fusion at both image and feature levels, while AF effectively integrates multi-modal features with low computational complexity. Our approach has been trained and validated on the publicly available DSEC-Detection dataset, achieving mAP50 and mAP50-95 scores of 67.2% and 45.6%, respectively, demonstrating superior detection performance and validating the effectiveness of the proposed method. Sin-Ye Jhong, Hsin-Chun Lin, Tzu-Chi Liu, Kai-Lung Hua, Yung-Yao Chen |
ICRA | 5 |
| 2025 | Deterministic Optimization-Based Path Planning Techniques for Obstacle Avoidance in Human-Robot Collaborative ScenariosabstractIn the context of Smart Manufacturing, robots are increasingly designed to operate alongside humans, with collaborative robots playing a central role. Ensuring safety in such Human-Robot Collaboration (HRC) scenarios require advanced path planning systems capable of detecting and responding to obstacles, including humans, in real time, thereby enabling safe and efficient cooperation. This paper presents a comparative study of optimization-based path planning techniques applied in collaborative robotics for obstacle avoidance. A structured comparison is conducted among different deterministic optimization strategies, such as Newton’s Method, Conjugate Gradient, and Gradient Descent, each employing various methods for obstacle pose identification, estimation, and danger factor modeling. Through an extensive review of these methods and their outcomes, the study evaluates their performance in terms of path safety, computational efficiency, and adaptability to dynamic environments. The analysis highlights the strengths and limitations of each optimization-based model and provides guidance for selecting suitable path planning approaches for different robotic applications. Brijesh Patel 0002, Yung-Chieh Chang, Po Ting Lin, Chao-Lung Yang, Yung-Yao Chen, Kai-Lung Hua, Meng-Kun Liu |
SMC | 5 |
| 2025 | Radiance Field-Based Pose Estimation via Decoupled Optimization Under Challenging Initial Conditions
Si-Yu Lu, Yung-Yao Chen, Yi-Tong Wu, Hsin-Chun Lin, Sin-Ye Jhong, Wen-Huang Cheng |
WACV | 2 |
| 2025 | Deep integration of conditional gan, attention mechanism, and image clustering for automated color separation and correction in textile screen printingabstractManual color separation and image processing are needed for pretreatment in traditional textile screen printing but the shortage of engineers has created a significant bottleneck in improving printing quality. Therefore, this research aimed to introduce an innovative method by proposing the squeeze-excitation attention and Pix2PixHD (SEAPix) incorporating KMeans clustering algorithm, forming KMeans+SEAPix model. The proposed SEAPix integrates a squeeze and excitation attention (SEA) mechanism into the conditional generative adversarial network (cGAN), Pix2PixHD, architecture. KMeans algorithm facilitates the separation of similar colors into distinct branches. Furthermore, SEAPix method improves the accuracy of high-resolution image generation by focusing on important regions and reducing irrelevant noise. To verify the sequence of color separation and image processing, this research developed two different frameworks, namely the Image Generation-Separation Framework (PF1) and the Image Generation-Separation Framework (PF2). PF1 starts with image generation, followed by Kmeans color separation, while PF2 adopted a different method by using color separation first, followed by generating images with cGAN. The experimental results of this research showed that PF2 was superior to PF1. Specifically, PF2 had better structural similarity index measure (SSIM), peak signal-to-noise ratio (PSNR), intersection over union (IoU), and pixel accuracy increased by 15.5643%, 0.5249%, 11.3757%, and 1.6508%, respectively. PF2 also has lower learned perceptual image patch similarity (LPIPS) and mean squared error (MSE) decreased by 24.3518% and 2.5963% compared with PF1. In conclusion, the proposed KMeans+SEAPix could revolutionize traditional textile screen printing methods by providing automated color separation and image correction. Chao-Lung Yang, Yulius Harjoseputro, Chi-Hao Chien, Yung-Yao Chen |
Appl. Intell. | 4 |
| 2025 | Hybrid CNN-ViT architecture to exploit spatio-temporal feature for fire recognition trained through transfer learning
Hong-Cyuan Wang, Yung-Yao Chen, Kai-Lung Hua |
Multim. Tools Appl. | 3 |
| 2025 | A hybrid approach of simultaneous segmentation and classification for medical image analysis
Chao-Lung Yang, Yulius Harjoseputro, Yung-Yao Chen |
Multim. Tools Appl. | 3 |
| 2023 | MCGAN: mask controlled generative adversarial network for image retargeting
Jilyan Bianca Dy, John Jethro Virtusio, Daniel Stanley Tan, Yong-Xiang Lin, Joel P. Ilao, Yung-Yao Chen, Kai-Lung Hua |
Neural Comput. Appl. | 6 |
| 2023 | Controllable Model Compression for Roadside Camera Depth EstimationabstractIn the Cooperative Intelligent Transportation System (C-ITS) paradigm, vehicles could communicate with roadside units to augment their traffic knowledge. Smart roadside units could provide second-order information (e.g., vehicle count) from raw first-order data (e.g., visual feed, point clouds), and this “smart” feature is usually provided using deep neural network models. However, implementing these useful models implies a cost for computational complexity that could hinder the future deployment of smart roadside units needed for sustainability in transportation systems. In this paper, we propose to use model compression on deep image processing models to promote its feasibility for usage in smart sensors. We formulated a controllable convolutional model compression (CCMC) algorithm that can perform filter-wise evolutionary pruning on image processing networks, along with a predefined compression ratio. CCMC is applicable for image processing networks, which have multiple possible traffic data sources (e.g., road camera surveillance). Furthermore, CCMC has a definable target compression ratio that is useful for controlling the trade-off between resource consumption and output performance. We tested our proposed method on depth estimation, which is useful for scene understanding and mapping the locations of objects in the 3D space. Our experiments show that the pruned model has minimal performance discrepancy from the original one, supporting the sustainability features needed for intelligent transportation systems. Jose Jaena Mari Ople, Shang-Fu Chen, Yung-Yao Chen, Kai-Lung Hua, Mohammad Hijji, Po Yang 0001, Khan Muhammad 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Code generation from a graphical user interface via attention-based encoder-decoder model
Wen-Yin Chen, Pavol Podstreleny, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
Multim. Syst. | 4 |
| 2022 | VDNet: video deinterlacing network based on coarse adaptive module and deformable recurrent residual network
Yin-Chen Yeh, Jilyan Bianca Dy, Tai-Ming Huang, Yung-Yao Chen, Kai-Lung Hua |
Neural Comput. Appl. | 4 |
| 2021 | Explainable AI: A Multispectral Palm-Vein Identification System with New Augmentation FeaturesabstractRecently, as one of the most promising biometric traits, the vein has attracted the attention of both academia and industry because of its living body identification and the convenience of the acquisition process. State-of-the-art techniques can provide relatively good performance, yet they are limited to specific light sources. Besides, it still has poor adaptability to multispectral images. Despite the great success achieved by convolutional neural networks (CNNs) in various image understanding tasks, they often require large training samples and high computation that are infeasible for palm-vein identification. To address this limitation, this work proposes a palm-vein identification system based on lightweight CNN and adaptive multi-spectral method with explainable AI. The principal component analysis on symmetric discrete wavelet transform (SMDWT-PCA) technique for vein images augmentation method is adopted to solve the problem of insufficient data and multispectral adaptability. The depth separable convolution (DSC) has been applied to reduce the number of model parameters in this work. To ensure that the experimental result demonstrates accurately and robustly, a multispectral palm image of the public dataset (CASIA) is also used to assess the performance of the proposed method. As result, the palm-vein identification system can provide superior performance to that of the former related approaches for different spectrums. Yung-Yao Chen, Sin-Ye Jhong, Chih-Hsien Hsia, Kai-Lung Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Sketch-guided Deep Portrait GenerationabstractGenerating a realistic human class image from a sketch is a unique and challenging problem considering that the human body has a complex structure that must be preserved. Additionally, input sketches often lack important details that are crucial in the generation process, hence making the problem more complicated. In this article, we present an effective method for synthesizing realistic images from human sketches. Our framework incorporates human poses corresponding to locations of key semantic components (e.g., arm, eyes, nose), seeing that its a strong prior for generating human class images. Our sketch-image synthesis framework consists of three stages: semantic keypoint extraction, coarse image generation, and image refinement. First, we extract the semantic keypoints using Part Affinity Fields (PAFs) and a convolutional autoencoder. Then, we integrate the sketch with semantic keypoints to generate a coarse image of a human. Finally, in the image refinement stage, the coarse image is enhanced by a Generative Adversarial Network (GAN) that adopts an architecture carefully designed to avoid checkerboard artifacts and to generate photo-realistic results. We evaluate our method on 6,300 sketch-image pairs and show that our proposed method generates realistic images and compares favorably against state-of-the-art image synthesis methods. Trang-Thi Ho, John Jethro Virtusio, Yung-Yao Chen, Chih-Ming Hsu, Kai-Lung Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Spatially-Aware Domain Adaptation for Semantic Segmentation of Urban ScenesabstractIt is very expensive and time consuming to collect a large enough dataset with pixel-level annotations to train a semantic segmentation model. Synthetic datasets are common alternatives for training segmentation models, however models trained on synthetic data do not necessarily perform well on real world images due to the domain shift problem. Domain adaptation techniques address this problem by leveraging on adversarial training to align features. Prior works have mostly performed global feature alignment. They do not consider the positions of objects. However, objects in urban scenes are highly correlated with their spatial locations. For example, the sky will always appear on top while cars will usually appear in the middle of the image. Based on this insight, we propose a spatial-aware discriminator that accounts for the spatial prior on the objects in order to improve the feature alignment. We demonstrate in our experiments that our model outperforms several state-of-the-art baselines in terms of mean intersection over union (mIoU). Yong-Xiang Lin, Daniel Stanley Tan, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
ICIP | 4 |
| 2019 | Cloud image watermarking: high quality data hiding and blind decoding scheme based on block truncation coding
Yung-Yao Chen, Kuan-Yu Chi |
Multim. Syst. | 1 |
| 2018 | Pedestrian Detection from Lidar Data via Cooperative Deep and Hand-Crafted FeaturesabstractAutopilot systems need to be able to detect pedestrians with high precision and recall regardless of whether it is during the day or night. This means that we cannot rely on normal cameras to sense the surroundings due to its sensitivity to lighting conditions. An alternative for images is to use light detection and ranging sensors (LiDAR) that produces three-dimensional point clouds where each point represents the distance to an object. However, most pedestrian detection systems are designed for image inputs and not on distance point clouds. In this paper, we propose a method for detecting pedestrians using only the three-dimensional point clouds generated by the LiDAR. Our approach first projects the three-dimensional point cloud into a two-dimensional plane. We then extract both hand-crafted features and learned features from a convolutional neural network in order to train a support vector machine (SVM) to detect pedestrians. Our proposed method achieved significant improvements in terms of F1-measurement over prior state-of-the-art methods. Tzu-Chieh Lin, Daniel Stanley Tan, Hsueh-Ling Tang, Shih-Che Chien, Feng-Chia Chang, Yung-Yao Chen, Wen-Huang Cheng, Kai-Lung Hua |
ICIP | 6 |
| 2018 | Vehicle Detection in Thermal Images Using Deep Neural NetworkabstractIn today's world, it becomes critical for a self-driving car to detect the vehicles irrespective of it being a day or night. We propose a real-time vehicle detection using a sequence of night-time thermal images. Moreover, the thermal images have the capability of retaining even the minuscule vehicle details in a dim environment. For an efficient vehicle detection, the thermal image dataset collected during the dusk and night is used for training purposes. Subsequently, the contrast enhancement and sharpening of these images are performed using the Thermal Feature Enhancement (TFE). Then the concatenated images are supplied as the input to allow the model to learn more effectively. Besides, we also propose an improved convolution network model entitled as the Thermal Image Only Looked Once (TOLO) model for vehicle detection. Additionally, we propose a method called as Low Probability Candidate Filter (LPCF) to compensate the probability of not-easy-to-detect vehicles. Our proposed method produces better results for the F1-measure in comparison with existing methods. Chin-Wei Chang, Kathiravan Srinivasan, Yung-Yao Chen, Wen-Huang Cheng, Kai-Lung Hua |
VCIP | 3 |
| 2018 | High-quality blind watermarking in halftones using random toggle approach
Yung-Yao Chen, Wei-Sheng Chen |
Multim. Tools Appl. | 1 |
| 2017 | Multi-cue pedestrian detection from 3D point cloud dataabstractPedestrian detection is one of the key technologies of driver assistance system. In order to prevent potential collisions, pedestrians should be always accurately identified whether during the day or at night. Since the visual images of the night are not clear, this paper proposes a method for recognizing pedestrians by using a high-definition LIDAR without visual images. In order to handle the long-distance sparse point problem, a novel solution is introduced to improve the performance. The proposed method maps the three-dimensional point cloud to the two-dimensional plane by a distance-aware expansion approach and the corresponding 2D contour and its associated 2D features are then extracted. Based on both 2D and 3D cues, the proposed method obtains significant performance boosts over state-of-the-art approaches by 13% in terms of F1-measure. Hsueh-Ling Tang, Shih-Che Chien, Wen-Huang Cheng, Yung-Yao Chen, Kai-Lung Hua |
ICME | 4 |
| 2016 | The Lattice-Based Screen Set: A Square N-Color All-Orders Moiré-Free Screen SetabstractPeriodic clustered-dot screens are widely used for electrophotographic printers due to their print stability. However, moiré is a ubiquitous problem that arises in color printing due to the beating together of the clustered-dot, periodic halftone patterns that are used to represent different colorants. The traditional solution in the graphic arts and printing industry is to rotate identical square screens to angles that are maximally separated from each other. However, the effectiveness of this approach is limited when printing with more than four colorants, i.e., N -color printing, where N > 4 . Moreover, accurately achieving the angles that have maximum angular separation requires a very high-resolution plate writer, as is used in commercial offset printing. Commercially available high-end digital printers cannot achieve this resolution. In this paper, we propose a systematic way to design color screen sets for periodic, clustered-dot screens that offer more explicit control of the moiré properties of the resulting screens when used in color printing. We develop a principled approach for the moiré-free screen design that is called lattice-based screen design. The basic concept behind our approach is the creation of the screen set on a 2D lattice in the frequency domain, and then picking each fundamental frequency vector of the individual colorant planes in the created spectral lattice according to the desired properties. The lattice-based screen design offers more flexibility in designing N -color screen sets with different halftone geometries, and all of them are guaranteed to be all-orders moiré-free. We demonstrate the efficacy of our proposed method by introducing several new screen designs, and a comparison with published screen designs. Yung-Yao Chen, Tamar Kashti, Mani Fischer, Doron Shaked, Robert Ulichney, Jan P. Allebach |
IEEE Trans. Image Process. | 1 |
| 2011 | Design of color screen sets for robustness to color plane misregistrationabstractPeriodic clustered-dot screens are widely used for electrophotographic printers due to their homogeneous halftone texture and their robustness to dot gain. However, when applied to color printing, there are two important phenomena that limit the quality of printed color halftones generated using a screening technology: (1) moire´ due to the superposition halftone patterns corresponding to different periodicity matrices, and (2) appearance changes due to misregistration between different colorant planes. This paper focuses on analyzing the registration sensitivity of periodic, clustered-dot screens. To quantitatively measure the effect of registration errors, we introduce two new functions: (1) cost, and (2) risk of registration errors. We propose the notion of “visual equivalence”, and derive three propositions under which visual equivalence can be achieved, even when registration errors occur. Yung-Yao Chen, Mani Fischer, Omri Shacham, Carl Staelin, Jan P. Allebach |
ICIP | 2 |
| 2006 | A Vision-Based Parking Lot Management SystemabstractGoals of parking lot management system include counting the number of parked vehicles, monitoring the changes of the parked vehicles over the time, and identifying the stalls available. To decrease the cost of the production, an integrated vision-based system is a good choice. In this paper, we propose a vision-based parking management system to manage an outdoor parking lot by four cameras set up at loft of buildings around it, sending information, including real-time display, to database of ITS center via internet. This system enables drivers to find parking spaces available or monitoring the parking lot where they parked their cars easily by wireless communication device. To increase accuracy, in the beginning, color manage is done to all input images, maintaining color consistency. Then, an adaptive parking lot background model is generated. The adequate color of each parking space is found out using statistical method in color image sequences captured by a camera, and foreground is extracted based on color information. The result will be further modified by shadow detection based on luminance analysis. Vision-based parking management system can manage large area by just several cameras. Adjusting position of the camera can easily make this system suitable for most cases. Besides, this system is endurable and is easy-installed because of its simple equipment. Sheng-Fuu Lin, Yung-Yao Chen, Sung-Chieh Liu |
SMC | 2 |