Li-Yun Wang

dblp:57/9696 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0003-4288-2569ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Night-to-Day: Unpaired Image-to-Image Translation for Nighttime Pedestrian Detection
abstract
In this paper, we show that exploiting Generative Adversarial Networks (GANs) to transform nighttime images into daytime representation increases the robustness of pedestrian detection in low-light conditions. Our work aims at first learning the image translation to transfer the style from daytime images to nighttime images with unpaired GAN training. Second, we use our end-to-end trained GAN model to translate night images as a pre-processing step before feeding them into an object detector that is pre-trained on daytime images only. To demonstrate the effectiveness of our translation approach, we conducted experiments on two real-world pedestrian datasets using both one-stage and two-stage object detectors. Our results outperform the baseline in all experiments and show highly competitive detection performance compared with other GAN-based approaches while holding the most lightweight architecture. We believe that our approach is an effective pre-processing first step that helps in bridging the performance gap between day and night at no expense of re-training object detector networks with more night images.
Afnan Althoupety, Li-Yun Wang, Wu-chi Feng, Banafsheh Rekabdar
ECAI2
2023 Illuminating the Bias in Pedestrian Detection
abstract
Despite major advancements in state-of-the-art object detectors and low-light image enhancements, nighttime pedestrian detection remains a challenge. One commonly used solution to remedy this domain shift problem is re-training object detectors with extra nighttime scenes which is not only computationally expensive but also not generalizable. In this paper, we explore a new solution and aim to systematically analyze and understand the efficiency of a lightening algorithm on pedestrian detection in low-light conditions. In our analysis, we first explore the effect of normalizing image luminance based on the ground truth bounding boxes to adaptively adjust global image luminance and evaluate its effects on detection performance. Second, unlike general low-light image enhancements that rely on global or local image statistics, we design a pedestrian-luminance-aware lightening algorithm to automatically correct nighttime images luminance so that pedestrians can be more robustly detected. Through extensive experiments, our algorithm not only achieves competitive detection results compared to the baseline on two real-world nighttime datasets but also elevates the confidence score of detected pedestrians.
Afnan Althoupety, Li-Yun Wang, Wu-chi Feng, Banafsheh Rekabdar
ISM2
2023 Video as Text: A New Paradigm for Flexible Video Analysis
abstract
The ability to capture and distribute digital videos has been available for many years. Users can easily capture high-quality video streams with mobile devices and distribute them to end users through varying platforms. This paper presents the design and implementation of a new multimedia framework called Video as Text (vText), which analyzes and manipulates video data as trivially as we handle text data in most Unix and Linux systems. In most Unix systems, it is easy to accomplish highly complex textual analysis and processing by combining relatively simple programs (e.g., grep, awk, sed, and cut) through Unix pipes. The vText paradigm seeks to mimic such programs. We demonstrate the design and implementation of vText linking video codecs with computer vision and image-processing algorithms. The experimental results indicate that the combination of simple programs provides high flexibility to users but does not incur high overhead when processing the video data.
Li-Yun Wang, Wu-chi Feng
ISM1
2022 Class Specialized Knowledge Distillation
Li-Yun Wang, Anthony Rhodes, Wu-chi Feng
ACCV (2)1
2021 Adversarial Perturbation Suppression using Adaptive Gaussian Smoothing and Color Reduction
abstract
This paper presents a novel context-aware image denoising algorithm that combines an adaptive image smoothing technique and color reduction techniques to remove perturbation from adversarial images. Adaptive image smoothing is achieved using auto-threshold canny edge detection to produce an accurate edge map used to produce a blurred image that preserves more edge features. The proposed algorithm then uses color reduction techniques to reconstruct the image using only a few representative colors. Through this technique, the algorithm can reduce the effects of adversarial perturbations on images. We also discuss experimental data on classification accuracy. Our results showed that the proposed approach reduces adversarial perturbation in adversarial attacks and increases the robustness of the deep convolutional neural network models.
Li-Yun Wang
ISM1
2021 CCAP: Cooperative Context Aware Pruning for Neural Network Model Compression
abstract
In this paper, we propose a new cross-domain model compression technique to yield a compact target model. We utilize a Cooperative Context-Aware Pruning (CCAP) module to produce sparse attention maps. They are then used to transmit the source models’ parameters to the target model precisely. We also leverage a weight-regular loss to minimize the difference between the source models’ and the target models’ parameters. Our quantitatively empirical evaluation shows that our CCAP module plus the weight-regular loss achieves lower model complexity without having serious performance decreasing.
Li-Yun Wang, Zahid Akhtar
ISM1
2020 FID: Frame Interpolation and DCT-based Video Compression
abstract
In this paper, we present a hybrid video compression technique that combines the advantages of residual coding techniques found in traditional DCT-based video compression and learning-based video frame interpolation to reduce the amount of residual data that needs to be compressed. Learning-based frame interpolation techniques use machine learning algorithms to predict frames but have difficulty with uncovered areas and non-linear motion. This approach uses DCT-based residual coding only on areas that are difficult for video interpolation and provides tunable compression for such areas through an adaptive selection of data to be encoded. Experimental data for both PSNR and the newer video multi-method assessment fusion (VMAF) metrics are provided. Our results show that we can reduce the amount of data required to represent a video stream compared with traditional video coding while outperforming video frame interpolation techniques in quality.
Yeganeh Jalalpour, Li-Yun Wang, Wu-chi Feng, Feng Liu 0015
ISM2
2019 Leveraging Image Processing Techniques to Thwart Adversarial Attacks in Image Classification
abstract
Deep Convolutional Neural Networks (DCNNs) are vulnerable to images that have been altered with well-engineered and imperceptible perturbations. We propose three color quantization pre-processing techniques to make DCNNs more robust to adversarial perturbation including Gaussian smoothing and PNM color reduction (GPCR), color quantization using Gaussian smoothing and K-means (GK-means), and fast GK-means. We evaluate the approaches on a subset of the ImageNet dataset. Our evaluation reveals that our GK-means-based algorithms have the best top-1 accuracy. We also present the trade-off between GK-means-based algorithms and GPCR with respect to computational time.
Yeganeh Jalalpour, Li-Yun Wang, Ryan Feng, Wu-chi Feng
ISM2
2014 Human Detection in Surveillance Video
abstract
In this paper, we propose an integrated approach for human detection in surveillance video. In our approach, the moving object is extracted by background subtraction; and the background model is updated by the first-order recurrence filter. Then, two complementary features are extracted for moving object classification. They are contour-based description: Fourier descriptor and region-based description: histogram of oriented gradient. As the binary classifier (support vector machine) is able to provide the posterior probability, we effectively integrate two types of features to achieve better performance. Experimental results show that the proposed approach is effective and outperforms some existing technique.
Liang-Hua Chen, Li-Yun Wang, Chih-Wen Su
Int. J. Pattern Recognit. Artif. Intell.2