EDBT 2026 Demo / reviewers in the wild / expert
Che-Tsung Lin
dblp:119/3959
· DBLP profile ↗
21ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-5843-7294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Texture: Advanced Facial Privacy Protection via Hierarchical Diffusion Autoencoder
Ting-Yi Lu, Che-Tsung Lin, Christopher Zach, Shang-Hong Lai |
ICPR (7) | 2 |
| 2025 | Text in the dark: Extremely low-light text image enhancement
Che-Tsung Lin, Chun Chet Ng, Zhi Qin Tan, Wan Jun Nah, Xinyu Wang 0010, Jie-Long Kew, Po-Hao Hsu, Shang-Hong Lai, Chee Seng Chan, Christopher Zach |
Signal Process. Image Commun. | 1 |
| 2024 | When IC meets text: Towards a rich annotated integrated circuit text dataset
Chun Chet Ng, Che-Tsung Lin, Zhi Qin Tan, Xinyu Wang 0010, Jie-Long Kew, Chee Seng Chan, Christopher Zach |
Pattern Recognit. | 2 |
| 2023 | Rethinking Long-Tailed Visual Recognition with Dynamic Probability Smoothing and Frequency Weighted FocusingabstractDeep learning models trained on long-tailed (LT) datasets often exhibit bias towards head classes with high frequency. This paper highlights the limitations of existing solutions that combine class- and instance-level re-weighting loss in a naive manner. Specifically, we demonstrate that such solutions result in overfitting the training set, significantly impacting the rare classes. To address this issue, we propose a novel loss function that dynamically reduces the influence of outliers and assigns class-dependent focusing parameters. We also introduce a new long-tailed dataset, ICText-LT, featuring various image qualities and greater realism than artificially sampled datasets. Our method has proven effective, outperforming existing methods through superior quantitative results on CIFAR-LT, Tiny ImageNet-LT, and our new ICText-LT datasets. The source code and new dataset are available at https://github.com/nwjun/FFDS-Loss. Wan Jun Nah, Chun Chet Ng, Che-Tsung Lin, Yeong Khang Lee, Jie-Long Kew, Zhi Qin Tan, Chee Seng Chan, Christopher Zach, Shang-Hong Lai |
ICIP | 3 |
| 2023 | Decentralized Training of 3D Lane Detection with Automatic Labeling Using HD MapsabstractTo have competent 3D lane detection for real-world driving, a massive amount of data from all over the world is needed, but data collection and manual annotation are costly and time-consuming. The diversity of data collected by developmental cars might still be limited compared to the data collected by a large fleet of customer cars. Federated learning enables training models on edge without transferring data out of devices. However, training supervised learning tasks at the edge is directly tied to having access to high-quality labels, which is limited at the edge.In this paper, we propose a fully automatic method to generate 3D lane labels at the edge using a pre-recorded HD map to enable the federated training of the 3D lane detection model. As a reference, a semi-automatic method is applied for creating a 3D-lane dataset used as ground truth. Our experimental results show that the model can achieve comparable performance when training on the same dataset in both a centralized and a decentralized manner. And the models trained on semi-automatic labeled datasets slightly outperform those trained on fully-automatically labeled datasets. This study shows that a well-performing 3D lane detection model can be trained in a supervised and fully decentralized manner, and most importantly, data privacy at the edge is guaranteed. Yadong Mao, Zhuqi Xiao, Che-Tsung Lin, Pedro Porto Buarque de Gusmão, Nicholas D. Lane, Christopher Zach, Mina Alibeigi |
VTC2023-Spring | 3 |
| 2023 | Cycle-object consistency for image-to-image domain adaptation
Che-Tsung Lin, Jie-Long Kew, Chee Seng Chan, Shang-Hong Lai, Christopher Zach |
Pattern Recognit. | 1 |
| 2022 | AdaSTE: An Adaptive Straight-Through Estimator to Train Binary Neural NetworksabstractWe propose a new algorithm for training deep neural networks (DNNs) with binary weights. In particular, we first cast the problem of training binary neural networks (BiNNs) as a bilevel optimization instance and subsequently construct flexible relaxations of this bilevel program. The resulting training method shares its algorithmic simplicity with several existing approaches to train BiNNs, in particular with the straight-through gradient estimator successfully employed in BinaryConnect and subsequent methods. Infact, our proposed method can be interpreted as an adaptive variant of the original straight-through estimator that conditionally (but not always) acts like a linear mapping in the backward pass of error propagation. Experimental results demonstrate that our new algorithm offers favorable performance compared to existing approaches.11This work was partially supported by theWallenberg AI, Autonomous Systems and Software Program (WASP) funded by the Knut and Alice Wallenberg Foundation. Huu Le, Rasmus Kjær Høier, Che-Tsung Lin, Christopher Zach |
CVPR | 3 |
| 2022 | CyEDA: Cycle-Object Edge Consistency Domain AdaptationabstractA difficulty of global-level translation is to preserve instance-level details in an image. Although some instance level translation methods can retain the details, most of them require either pre-trained object detection/segmentation network or annotation labels. In this work, we propose a novel method namely CyEDA to perform global level domain adaptation that can preserve image contents without any pre-trained networks integration or annotation labels. Specifically, we introduce blending masks and cycle-object edge consistency loss which exploit the preservation of image objects. We show that our approach can outperform other SOTAs in terms of image quality and FID score in both BDD100K and GTA datasets. The code and pre-trained models are publicly available at https://github.com/bjc1999/CyEDA. Jing Chong Beh, Kam Woh Ng, Jie-Long Kew, Che-Tsung Lin, Chee Seng Chan, Shang-Hong Lai, Christopher Zach |
ICIP | 4 |
| 2022 | Extremely Low-Light Image Enhancement with Scene Text RestorationabstractDeep learning-based methods have made impressive progress in enhancing extremely low-light images - the image quality of the reconstructed images has generally improved. However, we found out that most of these methods could not sufficiently recover the image details, for instance, the texts in the scene. In this paper, a novel image enhancement framework is proposed to precisely restore the scene texts, as well as the overall quality of the image simultaneously under extremely low-light conditions. Mainly, we employed a self-regularised attention map, an edge map, and a novel text detection loss. In addition, leveraging the synthetic low-light images is beneficial for image enhancement on the genuine ones in terms of text detection. The quantitative and qualitative experimental results have shown that the proposed model outperforms state-of-the-art methods in image restoration, text detection, and text spotting on See In the Dark and ICDAR15 datasets. Po-Hao Hsu, Che-Tsung Lin, Chun Chet Ng, Jie-Long Kew, Mei Yih Tan, Shang-Hong Lai, Chee Seng Chan, Christopher Zach |
ICPR | 2 |
| 2022 | Energy-based Models for Deep Probabilistic RegressionabstractIt is desirable that a deep neural network trained on a regression task not only achieves high prediction accuracy, but its prediction posteriors are also well-calibrated, especially in safety-critical settings. Recently, energy-based models specifically to enrich regression posteriors have been proposed and achieve state-of-art results in object detection tasks. However, applying these models at prediction time is not straightforward as the resulting inference methods require to minimize an underlying energy function. Furthermore, these methods empirically do not provide accurate prediction uncertainties. Inspired by recent joint energy-based models for classification, in this work, we propose to utilize a joint energy model for regression tasks and describe architectural differences needed in this setting. Within this framework, we apply our methods to three computer vision regression tasks. We demonstrate that joint energy-based models for deep probabilistic regression improve the calibration property, do not require expensive inference, and yield competitive accuracy in terms of the mean absolute error (MAE). Xixi Liu 0001, Che-Tsung Lin, Christopher Zach |
ICPR | 2 |
| 2022 | Effortless Training of Joint Energy-Based Models with Sliced Score MatchingabstractStandard discriminative classifiers can be upgraded to joint energy-based models (JEMs) by combining the classification loss with a log-evidence loss. Hence, such models intrinsically allow detection of out-of-distribution (OOD) samples, and empirically also provide better calibrated posteriors, i.e. prediction uncertainties. However, the training procedure suggested for JEMs (using stochastic gradient Langevin dynamics—or SGLD— to maximize the evidence) is reported to be brittle. In this work we propose to utilize score matching—in particular sliced score matching—to obtain a stable training method for JEMs. We observe empirically that the combination of score matching with the standard classification loss leads to improved OOD detection and better calibrated classifiers for otherwise identical DNN architectures. Additionally, we also analyze the impact of replacing the regular soft-max layer for classification with a gated soft-max one in order to improve the intrinsic transformation invariance and generalization ability.1 Xixi Liu 0001, D. Staudt, Che-Tsung Lin, Christopher Zach |
ICPR | 3 |
| 2021 | GAN-Based Day-to-Night Image Style Transfer for Nighttime Vehicle DetectionabstractData augmentation plays a crucial role in training a CNN-based detector. Most previous approaches were based on using a combination of general image-processing operations and could only produce limited plausible image variations. Recently, GAN (Generative Adversarial Network) based methods have shown compelling visual results. However, they are prone to fail at preserving image-objects and maintaining translation consistency when faced with large and complex domain shifts, such as day-to-night. In this paper, we propose AugGAN, a GAN-based data augmenter which could transform on-road driving images to a desired domain while image-objects would be well-preserved. The contribution of this work is three-fold: (1) we design a structure-aware unpaired image-to-image translation network which learns the latent data transformation across different domains while artifacts in the transformed images are greatly reduced; (2) we quantitatively prove that the domain adaptation capability of a vehicle detector is not limited by its training data; (3) our object-preserving network provides significant performance gain in the difficult day-to-night case in terms of vehicle detection. AugGAN could generate more visually plausible images compared to competing methods on different on-road image translation tasks across domains. In addition, we quantitatively evaluate different methods by training Faster R-CNN and YOLO with datasets generated from the transformed results and demonstrate significant improvement on the object detection accuracies by using the proposed AugGAN model. Che-Tsung Lin, Sheng-Wei Huang, Yen-Yi Wu, Shang-Hong Lai |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Multimodal Structure-Consistent Image-to-Image TranslationabstractUnpaired image-to-image translation is proven quite effective in boosting a CNN-based object detector for a different domain by means of data augmentation that can well preserve the image-objects in the translated images. Recently, multimodal GAN (Generative Adversarial Network) models have been proposed and were expected to further boost the detector accuracy by generating a diverse collection of images in the target domain, given only a single/labelled image in the source domain. However, images generated by multimodal GANs would achieve even worse detection accuracy than the ones by a unimodal GAN with better object preservation. In this work, we introduce cycle-structure consistency for generating diverse and structure-preserved translated images across complex domains, such as between day and night, for object detector training. Qualitative results show that our model, Multimodal AugGAN, can generate diverse and realistic images for the target domain. For quantitative comparisons, we evaluate other competing methods and ours by using the generated images to train YOLO, Faster R-CNN and FCN models and prove that our model achieves significant improvement and outperforms other methods on the detection accuracies and the FCN scores. Also, we demonstrate that our model could provide more diverse object appearances in the target domain through comparison on the perceptual distance metric. Che-Tsung Lin, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong Lai |
AAAI | 1 |
| 2019 | Recognizing Chinese Texts with 3D Convolutional Neural NetworkabstractIn this paper, we propose a deep learning system to localize and recognize Chinese texts in scenes with signage and road marks through 3D convolutional neural network. The proposed system adopts YOLO for detecting target location and exploits 3D convolutional neural network for recognizing the contents. The proposed design outperforms the existing designs based on LSTM and achieves real-time processing performance, which is feasible to be implemented on embedded platforms. The proposed system reaches over 90% accuracy in recognizing Chinese texts on bird's-eye viewing road marks in a self-driving vehicle equipped with a fisheye camera. In addition, this system can achieve 20 fps execution speed with NVIDIA DIGITS DevBox with 1080Ti GPU, which is fast enough for autonomous driving applications. Kuan-Chou Chen, Guan-Ting Lin, Che-Tsung Lin, Jiun-In Guo |
ICIP | 3 |
| 2019 | Cross Domain Adaptation for on-Road Object Detection Using Multimodal Structure-Consistent Image-to-Image TranslationabstractImage-to-image translation is potential to boost the detection accuracy of a CNN-based object detector in a different domain. Despite recent GAN (Generative Adversarial Network) based methods have shown compelling visual results, they are prone to fail at preserving image-objects and maintaining structure consistency when faced with large and complex domain shifts such as day-to-night, which reduces their practicality on tasks such as generating large-scale training data for different domains. In this work, we introduce image-translation-structure and cycle-structure consistency for generating diverse and structure-preserved translated images across complex domains, such as between day and night, for object detector training. Given only a single/labelled image at daytime, our model could generate a diverse collection of images at nighttime with different ambient light levels and rear lamp conditions (on/off) but with the same vehicle type, color and locations. Qualitative results show that our model can generate diverse and realistic images in the target domain data. For quantitative comparisons, we evaluate other competing methods and ours by using the generated images to train the Faster R-CNN and YOLO detectors and prove that our model achieves significant improvement and outperforms other methods on detection accuracy. Che-Tsung Lin |
ICIP | 1 |
| 2018 | AugGAN: Cross Domain Adaptation with GAN-Based Data Augmentation
Sheng-Wei Huang, Che-Tsung Lin, Shu-Ping Chen, Yen-Yi Wu, Po-Hao Hsu, Shang-Hong Lai |
ECCV (9) | 2 |
| 2016 | Robust and Efficient Tracking with Large Lens Distortion for Vehicular Technology ApplicationsabstractAdvances in video technology have enabled its wide adoption in the auto industry. Today, many vehicles are equipped with backup, front-looking, and side-looking cameras that allow the driver to easily monitor traffic around the vehicle for enhancing safety. One difficulty with performing automated image analysis using a vehicle's onboard video has to do with the significant lens distortion of these sensors to cover a large field of view around the vehicle. This paper reports our research on proposing a tracking scheme that improves the accuracy and denseness of object tracking in the presence of large lens distortion. The contribution of our research is 4-fold: (1) We evaluated a large collection of state-of-the-art trackers to understand their deficiency when applied to videos with large lens distortion, (2) we showed how to derive useful evaluation metrics from public-domain, real-world driving videos that do not come with ground-truth information on pixel tracking, (3) we identified many enhancement techniques that can potentially help improve the poor performance of current trackers on videos of large lens distortion, and (4) we performed a systematic study to validate the efficacy of these enhancement techniques and proposed a new tracker design that achieved substantial improvement over the state-of-the- art, in terms of both accuracy and density, based on a rigorous precision vs. recall analysis. Che-Tsung Lin, Long-Tai Chen, Pai-Wei Cheng, Yuan-Fang Wang |
VTC Fall | 1 |
| 2015 | Evaluation, Design and Application of Object Tracking Technologies for Vehicular Technology ApplicationsabstractAdvances in video technology have enabled its wide adoption in the auto industry. Today, many vehicles are equipped with backup, front-looking, and side-looking cameras that allow the driver to easily monitor the traffic around the vehicle for enhanced safety. This paper reports our research on evaluating many existing object tracking techniques, and proposes a new tracker design and its application for 3D environmental mapping in vehicular technology applications. The contribution of our research is 4-fold: (1) We evaluate a large collection of state-of-the-art trackers using multiple criteria relevant to vehicular technology applications, (2) we show how to derive useful evaluation metrics from public-domain, real-world driving videos that do not come with ground-truth information on pixel tracking, (3) we propose a new tracker that is geared specifically for vehicular technology application and show that it achieves tracking accuracy that outperforms SIFT and is on-par with state-of-the-art optical-flow tracking algorithm, which has the best accuracy in our evaluation. Furthermore, we show that our tracker is 600 times more efficient than optical flow and 7 times more efficient than SIFT, and (4) we validated our new tracker design for 3D environmental map building application and showed that the new tracker can obtain comparable results as SIFT but with a significant saving in runtime. Che-Tsung Lin, Long-Tai Chen, Yuan-Fang Wang |
VTC Fall | 1 |
| 2014 | Enhancing Vehicular Safety in Adverse Weather Using Computer Vision AnalysisabstractThe goal of the project is to design intelligent and robust image-processing and augmented-reality algorithms for driver assistance and enhanced vehicular safety. In particular, the focuses were two-fold: (1) realizing the abilities to identify and localize in a vehicle''s on-board video the sweeping windshield wipers during raining days and (2) designing and implementing an in-painting technique to remove the image of the windshield wipers and replace it with the corresponding pixels (not blocked by the wipers) from an adjacent video frame. Che-Tsung Lin, Long-Tai Chen, Yuan-Fang Wang |
VTC Fall | 1 |
| 2013 | Front Vehicle Blind Spot Translucentization Based on Augmented RealityabstractThis paper proposes a new vehicle blind spot elimination system which utilizes the on-board videos captured from other vehicles and the host vehicle. Such information is exchanged by WAVE/DSRC devices to achieve collaborative safety. The preceding vehicle which fully or partially blocks the field of view of the host vehicle could be translucentized in the video captured by the host vehicle and the driving environment of the front vehicle could be then visually checked by the host driver. Che-Tsung Lin, Long-Tai Chen, Yuan-Fang Wang |
VTC Fall | 1 |
| 2012 | Self-learning-based rain streak removal for image/videoabstractRain removal from an image/video is a challenging problem and has been recently investigated extensively. In our previous work, we have proposed the first single-image-based rain streak removal framework via properly formulating it as an image decomposition problem based on morphological component analysis (MCA) solved by performing dictionary learning and sparse coding. However, in this previous work, the dictionary learning process cannot be fully automatic, where the two dictionaries used for rain removal were selected heuristically or by human intervention. In this paper, we extend our previous work to propose an automatic self-learning-based rain streak removal framework for single image. We propose to automatically self-learn the two dictionaries used for rain removal without additional information or any assumption. We then extend our single-image-based method to video-based rain removal in a static scene by exploiting the temporal information of successive frames and reusing the dictionaries learned by the former frame(s) in a video while maintaining the temporal consistency of the video. As a result, the rain component can be successfully removed from the image/video while preserving most original details. Experimental results demonstrate the efficacy of the proposed algorithm. Li-Wei Kang, Chia-Wen Lin, Che-Tsung Lin |
ISCAS | 3 |