Yung-Hui Li

dblp:148/9890 · DBLP profile ↗
← Back
41ranked-venue papers
0as first author
32since 2021 · last 2026
0000-0002-0475-3689ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 21 since 2021Artificial intelligence and machine learning · 17 · 15 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
abstract
Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on the critical challenge of how to restore surrogate-driven protected data in diverse MLLM scenarios. We first bridge this research gap by contributing the SPPE (Surrogate Privacy Protected Editable) dataset, which includes a wide range of privacy categories and user instructions to simulate real MLLM applications. This dataset offers protected surrogates alongside their various MLLM-edited versions, thus enabling the direct assessment of privacy recovery quality. By formulating privacy recovery as a guided generation task conditioned on complementary multimodal signals, we further introduce a unified approach that reliably reconstructs private content while preserving the fidelity of MLLM-generated edits. The experiments on both SPPE and InstructPix2Pix further show that our approach generalizes well across diverse visual content and editing tasks, achieving a strong balance between privacy protection and MLLM usability.
Yibing Liu, Peilin Chen 0001, Yung-Hui Li, Shiqi Wang 0001, Sam Kwong
AAAI4
2026 Knowledge-aware replay for multi-label class-incremental learning
Chengtai Cao, Xinhong Chen 0003, Qun Song 0001, Rui Tan 0001, Yung-Hui Li, Jianping Wang 0001
Expert Syst. Appl.5
2026 Defect image generation using diffusion model for defect detection augmentation
abstract
Automated visual inspection is critical for quality control in industrial manufacturing, yet training robust defect detection models is often hindered by the scarcity of defective samples. To address this challenge, we introduce a diffusion-based data augmentation framework capable of generating diverse and realistic synthetic defect images. Our method personalizes a text-to-image Latent Diffusion Model using DreamBooth, introducing a spatially weighted loss that forces the model to prioritize learning specific defect characteristics without attempting to reconstruct the non-defective background. To synthesize new data, we utilize Differential Diffusion to generate the learned defects onto defect-free images, eliminating boundary artifacts. We systematically evaluate the efficacy of our generated data on downstream supervised defect detection and localization tasks in two practical scenarios: as a complete substitute for real defect data and as a complement to it. Extensive experiments on the MVTec-AD dataset and a real-world industrial dataset demonstrate that our synthetic defects significantly boost downstream task performance. Furthermore, ablation studies confirm that these performance gains are robust and architecture-agnostic across various detectors.
Herman Prawiro, Ching-Yeh Chiang, Nien-Yi Jan, Kai-Lin Yang, Yi-Rong Lin, Yung-Hui Li, Tse-Yu Pan, Min-Chun Hu 0001
Multim. Tools Appl.6
2026 A Unified Experience Replay Framework for Spiking Deep Reinforcement Learning
abstract
Deep Reinforcement Learning (DRL) methods have shown remarkable success in many applications, yet their high energy consumption limits their practicability. Recent studies incorporated energy-efficient Spiking Neural Networks (SNNs) to build Spiking DRL methods and lower energy consumption by setting a shorter simulation duration for SNNs to compute fewer gradients. However, these existing Spiking DRL methods fail to sample sufficient high-quality samples within a fixed-size replay buffer and perform poorly when the simulation duration is small, introducing the challenging tradeoff between energy consumption and model performance. Motivated by such observations, we develop a generic resilient experience replay method that can be seamlessly integrated into existing spiking DRL methods to effectively address the above tradeoff. Specifically, we allow the replay buffer to dynamically expand as the number of training samples increases, thereby accommodating more potentially valuable candidate samples for policy training. Meanwhile, we introduce an adaptive approach to manage the buffer size by determining when to shrink the replay buffer and removing redundant samples automatically. This strategy prevents the buffer from expanding unnecessarily, thereby mitigating the potential negative impact on model performance. Extensive experimental results demonstrate that our approach significantly enhances the performance of five state-of-the-art (SOTA) spiking DRL methods across various simulation durations in sixteen tasks, in terms of return, without compromising their energy efficiency.
Meng Xu 0009, Xinhong Chen 0003, Bingyi Liu, Yi-Rong Lin, Yung-Hui Li, Jianping Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Optimizing Fidelity-Perception Tradeoff via Large Vision-Language Model Prior for Image Compression
abstract
Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.
Yudong Mao, Peilin Chen 0001, Lingyu Zhu 0006, Yung-Hui Li, Shiqi Wang 0001
IEEE Trans. Image Process.5
2025 Memory-Augmented Re-Completion for 3D Semantic Scene Completion
abstract
Semantic Scene Completion (SSC) aims to reconstruct a 3D voxel representation occupied by semantic classes based on ordinary inputs such as 2D RGB images, depth maps, or point clouds. Given the cost-effective and promising applications in autonomous driving, camera-based SSC has attracted considerable attention to developing various approaches. However, current methods mainly focus on precise 2D-to-3D projection while overlooking the challenge of completing invisible regions, leading to numerous false negatives and suboptimal SSC performance. To address this issue, we propose a novel architecture, Memory-augmented Re-completion (MARE), designed to enhance completion capability. Our MARE model encapsulates regional relationships by incorporating a memory bank that stores vital region-tokens while two protocols concerning diversity and age are adopted to optimize the bank adversarially. Additionally, we introduce a Re-completion pipeline incorporated with an Information Spreading module to progressively complete the invisible regions while bridging the scale gap between region-level and voxel-level information. Extensive experiments conducted on the SSCBench-KITTI-360 and SemanticKITTI datasets validate the effectiveness of our approach.
Yu-Wen Tseng, Sheng-Ping Yang, Jhih-Ciang Wu, I-Bin Liao, Yung-Hui Li, Hong-Han Shuai, Wen-Huang Cheng
AAAI5
2025 From Diffusion to Decision: A Diffusion-ReRanking in Scene Text Detection
abstract
Diffusion models have recently shown great potential in object detection and instance segmentation, yet their application to scene text detection, with its unique challenges such as instance variability and subjective human annotations, remains unexplored. In this paper, we propose DRR (Diffusion ReRanking), a method that adapts diffusion-based instance segmentation for scene text detection. Traditional instance segmentation often relies on classification scores for ranking, potentially overlooking the accuracy of bounding boxes and mask quality. DRR addresses this by incorporating two networks: a diffusion network, trained with a combination of projection loss and pairwise loss in the mask branch to produce more precise and tightly-bound segmentations, and a reranking network, which refines the results by evaluating bounding box accuracy and mask quality. Extensive experiments demonstrate the effectiveness of DRR, achieving a precision of 86.7%, recall of 81.7%, and an F-measure of 84.1% on CTW1500, highlighting DRR’s potential to advance scene text detection.
Jia-Ying Yong, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
AVSS4
2025 ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
abstract
Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this challenge, yet still suffer from limited trajectory diversity and uncalibrated mode confidence. While some approaches address these limitations by generating excessive trajectory candidates, they necessitate a postprocessing stage to identify the most representative modes, a process lacking universal principles and compromising trajectory accuracy. We are thus motivated to introduce ModeSeq, a new multimodal prediction paradigm that models modes as sequences. Unlike the common practice of decoding multiple plausible trajectories in one shot, ModeSeq requires motion decoders to infer the next mode step by step, thereby more explicitly capturing the correlation between modes and significantly enhancing the ability to reason about multimodality. Leveraging the inductive bias of sequential mode prediction, we also propose the EarlyMatch-Take-All (EMTA) training strategy to diversify the trajectories further. Without relying on dense mode prediction or heuristic post-processing, ModeSeq considerably improves the diversity of multimodal output while attaining satisfactory trajectory accuracy, resulting in balanced performance on motion prediction benchmarks. Moreover, ModeSeq naturally emerges with the capability of mode extrapolation, which supports forecasting more behavior modes when the future is highly uncertain.
Zikang Zhou, Hengjian Zhou, Jianping Wang 0001, Yung-Hui Li, Yu-Kai Huang 0001
CVPR6
2025 DAMO: Dual-Attention with Multi-Objective Optimization for Explainable Autonomous Driving
abstract
Deep learning has revolutionized autonomous driving; nevertheless, its inherent opacity hinders explainability, an essential requirement for public trust and regulatory approval. Existing explainable autonomous driving research typically employs a multi-task framework, simultaneously generating driving actions and their corresponding explanations (collectively called categories). Most methods use a two-stage approach: extracting category-related features and modeling category correlations separately. This separation overlooks the potential synergy between these two processes. Moreover, existing approaches often rely on simple linear combinations of task-specific losses, which may fail to optimally balance action and explanation objectives. To address these limitations, we propose Dual-Attention with Multi-Objective optimization (DAMO). DAMO introduces a dual-attention mechanism that alternates between cross-attention for category representation learning and self-attention for category correlation modeling, fostering mutual enhancement. Additionally, we devise a multi-objective optimization algorithm that dynamically balances tasks and achieves Pareto optimality with theoretical guarantees. Extensive evaluations on two benchmarks show that DAMO surpasses state-of-the-art baselines and a large vision-language model, delivering up to 13.9% performance improvement and enhanced generalization across diverse driving scenarios.
Chengtai Cao, Shenglin Wang, Xinhong Chen 0003, Yung-Hui Li, Jianping Wang 0001
ECAI4
2025 Global Regulation and Excitation via Attention Tuning for Stereo Matching
abstract
Stereo matching achieves significant progress with iterative algorithms like RAFT-Stereo and IGEV-Stereo. However, these methods struggle in ill-posed regions with occlusions, textureless, or repetitive patterns, due to a lack of global context and geometric information for effective iterative refinement. To enable the existing iterative approaches to incorporate global context, we propose the Global Regulation and Excitation via Attention Tuning (GREAT) framework which encompasses three attention modules. Specifically, Spatial Attention (SA) captures the global context within the spatial dimension, Matching Attention (MA) extracts global context along epipolar lines, and Volume Attention (VA) works in conjunction with SA and MA to construct a more robust cost-volume excited by global context and geometric details. To verify the universality and effectiveness of this framework, we integrate it into several representative iterative stereo-matching methods and validate it through extensive experiments, collectively denoted as GREAT-Stereo. This framework demonstrates superior performance in challenging ill-posed regions. Applied to IGEV-Stereo, among all published methods, our GREAT-IGEV ranks first on the Scene Flow test set, KITTI 2015, and ETH3D leaderboards, and achieves second on the Middlebury benchmark. Code is available at https://github.com/JarvisLee0423/GREAT-Stereo.
Xinhong Chen 0003, Zhengmin Jiang, Qian Zhou 0008, Yung-Hui Li, Jianping Wang 0001
ICCV5
2025 Unraveling Vanishing Point And Calibrating Tiny Objects For Semantic Scene Completion
abstract
Semantic Scene Completion (SSC) aims to jointly predict semantic categories and 3D occupancy of a scene from coarse inputs, which is crucial for providing reliable perception in autonomous driving. In this paper, we enhance existing SSC models by unveiling the vanishing point region, specifically addressing challenges posed by tiny objects and voxels distant from the monocular camera. At the core of our method, we propose the Vanishing Point Aggregator (VPA) to prior-itize features in high-density central areas. The proposed VPA seamlessly integrates the Vanishing Point Query (VPQ) with the vanilla instance query via a cross-attention fusion mechanism to refine feature representation. To evaluate the effectiveness of our method, we conduct comprehensive experiments on two standard SSC benchmarks and demonstrate that our method achieves SOTA performance. Our approach significantly improves the performance across various semantic classes, including a notable gain of 0.37 mIoU on SemanticKITTI and 0.5 mIoU on SSCBench-KITTI-360 for tiny objects. Ablation studies further validate the efficacy of our innovative query fusion strategy, showcasing its capability in long-range predictions for SSC tasks.
Sheng-Ping Yang, Yu-Wen Tseng, Yung-Chieh Yang, I-Bin Liao, Chi-En Huang, Shen-Hsuan Liu, Yung-Hui Li, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng
ICIP7
2025 PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
abstract
Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in high-level reasoning and long-horizon planning for robotic manipulation, they still fall short in grasping the nuanced physical properties required for effective human-robot interaction. In this paper, we introduce PAVLM (Point cloud Affordance Vision-Language Model), an innovative framework that utilizes the extensive multimodal knowledge embedded in pre-trained language models to enhance 3D affordance understanding of point cloud. PAVLM is an approach to integrates a geometric-guided propagation module with hidden embeddings from large language models (LLMs) to enrich visual semantics. On the language side, we prompt Llama-3.1 models to generate refined context-aware text, augmenting the instructional input with deeper semantic cues. Experimental results on the 3D-AffordanceNet benchmark demonstrate that PAVLM outperforms baseline methods for both full and partial point clouds, particularly excelling in its generalization to novel open-world affordance tasks of 3D objects. For more information, visit our project site: pavlm-source.github.io.
Shang-Ching Liu, Van Nhiem Tran, Wei-Lun Cheng, Yen-Lin Huang, I-Bin Liao, Yung-Hui Li
IROS7
2025 Fine-grained Stroke Recognition in Broadcast Table Tennis Videos with ATDT
abstract
This study introduces an automated system for fine-grained stroke recognition in broadcast table tennis videos, designed to address challenges in manual annotation and tactical analysis during international competitions. The proposed framework integrates an Adaptive Temporal Difference Model with a Transformer Encoder (ATDT), leveraging a combination of Temporal Difference Networks (TDN) and Temporal Adaptive Modules (TAM) to enhance spatial and temporal feature extraction. To enhance feature discriminability, we employ supervised contrastive learning, which promotes better representation learning for fine-grained action recognition. The system is divided into two primary modules: the Action Segmentation Module (ASM) and the Action Recognition Module (ARM). ASM precisely identifies the start and end times of each stroke action by incorporating ball trajectory analysis to identify precise hit timings and placements. The precise segmentation facilitates the subsequent ARM to implement a three-stage recognition process: forehand and backhand classification, group-based classification, and intra-group action classification. This hierarchical approach improves the system’s ability to differentiate between subtle stroke variations, even under the constraints of low-resolution broadcast footage. To validate the framework, the MISTT dataset was collected, comprising 3,618 stroke action clips from 18 international matches, with professional player annotations. The proposed ATDT model outperformed existing methods, achieving a top-1 accuracy improvement of 18% for forehand strokes and 25.58% for backhand strokes compared to baseline models. Moreover, our automatic annotation system takes only 1/30 of the time compared to the manual annotation process, demonstrating its efficiency.
Tang-Chen Chang, Duen-Chian Jheng, Hsuan-Ya Liang, Bill Louis Harchan, Pu Ching, Tsung-Hsun Tsai, Chih-Yi Chang, Te-Cheng Wu, Yung-Hui Li, Tse-Yu Pan, Hung-Kuo Chu, Min-Chun Hu 0001
ACM Trans. Multim. Comput. Commun. Appl.9
2025 A Unified Sparse Training Framework for Lightweight Lifelong Deep Reinforcement Learning
abstract
Lifelong deep reinforcement learning (DRL) methods enable continuous adaptation to new tasks and retention of old knowledge. However, these methods often necessitate large model sizes, leading to substantial computational and storage resource requirements during training and inference. Unfortunately, existing research has not yet provided a lightweight solution to address this issue. This work aims to develop a generic method that can be seamlessly integrated into existing lifelong DRL methods to facilitate their achievement of lightweight models while also yielding higher returns. While sparse training (ST) methods have been extensively used in the DRL community to achieve lightweight models, they exacerbate the issue of catastrophic forgetting and compromise generalization when applied in lifelong DRL. To improve generalization, we develop a gradient optimization method that leverages sharpness-aware minimization (SAM) to smooth the gradient surface of the model without introducing excessive computational complexity. In addition, to alleviate catastrophic forgetting and promote model convergence, we introduce a priority-based approach that samples effective past experiences from the replay buffer. Extensive experiments demonstrate that our approach achieves 90% sparsity in five representative lifelong DRL methods while achieving higher episode return and average return (up to 34% improvement) across all episodes compared to the dense models.
Meng Xu 0009, Xinhong Chen 0003, Yi-Rong Lin, Yung-Hui Li, Jianping Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2024 CCTR: Calibrating Trajectory Prediction for Uncertainty-Aware Motion Planning in Autonomous Driving
abstract
Autonomous driving systems rely on precise trajectory prediction for safe and efficient motion planning. Despite considerable efforts to enhance prediction accuracy, inherent uncertainties persist due to data noise and incomplete observations. Many strategies entail formalizing prediction outcomes into distributions and utilizing variance to represent uncertainty. However, our experimental investigation reveals that existing trajectory prediction models yield unreliable uncertainty estimates, necessitating additional customized calibration processes. On the other hand, directly applying current calibration techniques to prediction outputs may yield sub-optimal results due to using a universal scaler for all predictions and neglecting informative data cues. In this paper, we propose Customized Calibration Temperature with Regularizer (CCTR), a generic framework that calibrates the output distribution. Specifically, CCTR 1) employs a calibration-based regularizer to align output variance with the discrepancy between prediction and ground truth and 2) generates a tailor-made temperature scaler for each prediction using a post-processing network guided by context and historical information. Extensive evaluation involving multiple prediction and planning methods demonstrates the superiority of CCTR over existing calibration algorithms and uncertainty-aware methods, with significant improvements of 11%-22% in calibration quality and 17%-46% in motion planning.
Chengtai Cao, Xinhong Chen 0003, Jianping Wang 0001, Qun Song 0001, Rui Tan 0001, Yung-Hui Li
AAAI6
2024 DA2: Degree-Accumulated Data Augmentation on Point Clouds with Curriculum Dynamic Threshold Selection
Ta Chun Tai, Nhat-Tuong Do-Tran, Ngoc-Hoang-Lam Le, Yung-Hui Li
ACCV (9)4
2024 Learning Efficient Interaction Anchor for HOI Detection
abstract
Human-object interaction (HOI) detection seeks complicated relationships between humans and objects, yet struggles persist in correctly associating multiple objects with a single human in complex interaction scenarios. In this paper, we tackle such an issue by introducing a novel interaction anchor that employs flexible strategies across different decoder layers with Barlow constraint and Interactivity-Instance Fusion. The proposed modules are both additive and easily implementable in existing approaches, offering computational efficiency within transformer-based models to compact cross-interactivity. Extensive experiments validate the effectiveness of our method, demonstrating comparable performance on HICO-DET and V-COCO for HOI detection.
Lirong Xue, Kang-Yang Huang, Rong Chao, Jhih-Ciang Wu, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ICME6
2024 SGDCL: Semantic-Guided Dynamic Correlation Learning for Explainable Autonomous Driving
Chengtai Cao, Xinhong Chen 0003, Jianping Wang 0001, Qun Song 0001, Rui Tan 0001, Yung-Hui Li
IJCAI6
2024 TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views
abstract
We present TimeNeRF, a generalizable neural rendering approach for rendering novel views at arbitrary viewpoints and at arbitrary times, even with few input views. For real-world applications, it is expensive to collect multiple views and inefficient to re-optimize for unseen scenes. Moreover, as the digital realm, particularly the metaverse, strives for increasingly immersive experiences, the ability to model 3D environments that naturally transition between day and night becomes paramount. While current techniques based on Neural Radiance Fields (NeRF) have shown remarkable proficiency in synthesizing novel views, the exploration of NeRF's potential for temporal 3D scene modeling remains limited, with no dedicated datasets available for this purpose. To this end, our approach harnesses the strengths of multi-view stereo, neural radiance fields, and disentanglement strategies across diverse datasets. This equips our model with the capability for generalizability in a few-shot setting, allows us to construct an implicit content radiance field for scene representation, and further enables the building of neural radiance fields at any arbitrary time. Finally, we synthesize novel views of that time via volume rendering. Experiments show that TimeNeRF can render novel views in a few-shot setting without per-scene optimization. Most notably, it excels in creating realistic novel views that transition smoothly across different times, adeptly capturing intricate natural scene changes from dawn to dusk.
Hsiang-Hui Hung, Huu-Phu Do, Yung-Hui Li
ACM Multimedia3
2024 BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
abstract
Simulating realistic behaviors of traffic agents is pivotal for efficiently validating the safety of autonomous driving systems. Existing data-driven simulators primarily use an encoder-decoder architecture to encode the historical trajectories before decoding the future. However, the heterogeneity between encoders and decoders complicates the models, and the manual separation of historical and future trajectories leads to low data utilization. Given these limitations, we propose BehaviorGPT, a homogeneous and fully autoregressive Transformer designed to simulate the sequential behavior of multiple agents. Crucially, our approach discards the traditional separation between "history" and "future" by modeling each time step as the "current" one for motion generation, leading to a simpler, more parameter- and data-efficient agent simulator. We further introduce the Next-Patch Prediction Paradigm (NP3) to mitigate the negative effects of autoregressive modeling, in which models are trained to reason at the patch level of trajectories and capture long-range spatial-temporal interactions. Despite having merely 3M model parameters, BehaviorGPT won first place in the 2024 Waymo Open Sim Agents Challenge with a realism score of 0.7473 and a minADE score of 1.4147, demonstrating its exceptional performance in traffic agent simulation.
Zikang Zhou, Xinhong Chen 0003, Jianping Wang 0001, Nan Guan, Kui Wu 0001, Yung-Hui Li, Yu-Kai Huang 0001, Chun Jason Xue
NeurIPS7
2024 Natural Light Can Also be Dangerous: Traffic Sign Misinterpretation Under Adversarial Natural Light Attacks
abstract
Common illumination sources like sunlight or artificial light may introduce hidden vulnerabilities to AI systems. Our paper delves into these potential threats, offering a novel approach to simulate varying light conditions, including sunlight, headlights, and flashlight illuminations. Moreover, unlike typical physical adversarial attacks requiring conspicuous alterations, our method utilizes a model-agnostic black-box attack integrated with the Zeroth Order Optimization (ZOO) algorithm to identify deceptive patterns in a physically-applicable space. Consequently, attackers can recreate these simulated conditions, deceiving machine learning models with seemingly natural light. Empirical results demonstrate the efficacy of our method, misleading models trained on the GTSRB and LISA datasets under natural-like physical environments with an attack success rate exceeding 70% across all digital datasets, and remaining effective against all evaluated real-world traffic signs. Importantly, after adversarial training using samples generated from our approach, models showcase enhanced robustness, underscoring the dual value of our work in both identifying and mitigating potential threats.1
Teng-Fang Hsiao, Bo-Lun Huang, Zi-Xiang Ni, Yan-Ting Lin, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
WACV6
2024 Multi-view and multi-augmentation for self-supervised visual representation learning
Van-Nhiem Tran, Chi-En Huang, Shen-Hsuan Liu, Muhammad Saqlain Aslam, Kai-Lin Yang, Yung-Hui Li, Jia-Ching Wang
Appl. Intell.6
2024 Temporal Difference-Aware Graph Convolutional Reinforcement Learning for Multi-Intersection Traffic Signal Control
abstract
Traffic light control plays a crucial role in intelligent transportation systems. This paper introduces Temporal Difference-Aware Graph Convolutional Reinforcement Learning (TeDA-GCRL), a decentralized RL-based method for efficient multi-intersection traffic signal control. Specifically, we put forward a new graph architecture using each lane as a node for considering intersection relations. Additionally, we propose two new rewards by considering temporal information, namely Temporal-Aware Pressure on Incoming Lanes (TAPIL) and Temporal-Aware Action Consistency (TAAC), which enhance learning efficiency and time-interval sensitivity. Experimental results on five datasets show the superiority of TeDA-GCRL over state-of-the-art methods by at least 9.5% in average travel time.
Wei-Yu Lin, Yun-Zhu Song, Bo-Kai Ruan, Hong-Han Shuai, Li-Chun Wang 0001, Yung-Hui Li
IEEE Trans. Intell. Transp. Syst.7
2024 Self-Supervise Reinforcement Learning Method for Vacant Parking Space Detection Based on Task Consistency and Corrupted Rewards
abstract
This paper proposes a novel task-consistency learning method that enables us to train a vacant space detection network (target task) based on the logic consistency with the semantic outcomes from a flow-based motion behavior classifier (source task) in a parking lot. Note that the source task can introduce false detection during task-consistency learning, which implies noisy rewards or supervision. The target network can be trained in a reinforcement learning setting by appropriately designing the reward mechanism upon semantic consistency. We also introduce a novel symmetric constraint to detect corrupted samples and reduce the effect of noisy rewards. Unlike conventional corrupted learning methods that use only training losses to identify corrupted samples, our symmetric constraint also explores the relationship among training samples to improve performance. Compared with conventional supervised detection methods, the main contribution of our work is the ability to learn a vacant space detector via semantic consistency rather than supervised labels. The dynamic learning property allows the proposed detector to be easily deployed and updated in various lots without heavy human loads. Experiments demonstrate that our noisy task consistency mechanism can be successfully applied to train a vacant space detector from scratch.
Manh-Hung Nguyen 0002, Tzu-Yin Chao, Ching-Chun Hsiao, Yung-Hui Li
IEEE Trans. Intell. Transp. Syst.4
2024 Language-guided Residual Graph Attention Network and Data Augmentation for Visual Grounding
abstract
Visual grounding is an essential task in understanding the semantic relationship between the given text description and the target object in an image. Due to the innate complexity of language and the rich semantic context of the image, it is still a challenging problem to infer the underlying relationship and to perform reasoning between the objects in an image and the given expression. Although existing visual grounding methods have achieved promising progress, cross-modal mapping across different domains for the task is still not well handled, especially when the expressions are complex and long. To address the issue, we propose a language-guided residual graph attention network for visual grounding (LRGAT-VG), which enables us to apply deeper graph convolution layers with the assistance of residual connections between them. This allows us to better handle long and complex expressions than other graph-based methods. Furthermore, we perform a Language-guided Data Augmentation (LGDA), which is based on copy-paste operations on pairs of source and target images to increase the diversity of training data while maintaining the relationship between the objects in the image and the expression. With extensive experiments on three visual grounding benchmarks, including RefCOCO, RefCOCO+, and RefCOCOg, LRGAT-VG with LGDA achieves competitive performance with other state-of-the-art graph network-based referring expression approaches and demonstrates its effectiveness.
Jia Wang 0020, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ACM Trans. Multim. Comput. Commun. Appl.3
2024 HAPiCLR: heuristic attention pixel-level contrastive loss representation learning for self-supervised pretraining
Van-Nhiem Tran, Shen-Hsuan Liu, Chi-En Huang, Muhammad Saqlain Aslam, Kai-Lin Yang, Yung-Hui Li, Jia-Ching Wang
Vis. Comput.6
2023 Query-Centric Trajectory Prediction
abstract
Predicting the future trajectories of surrounding agents is essential for autonomous vehicles to operate safely. This paper presents QCNet, a modeling framework toward pushing the boundaries of trajectory prediction. First, we identify that the agent-centric modeling scheme used by existing approaches requires re-normalizing and re-encoding the input whenever the observation window slides forward, leading to redundant computations during online prediction. To overcome this limitation and achieve faster inference, we introduce a query-centric paradigm for scene encoding, which enables the reuse of past computations by learning representations independent of the global spacetime coordinate system. Sharing the invariant scene features among all target agents further allows the parallelism of multi-agent trajectory decoding. Second, even given rich encodings of the scene, existing decoding strategies struggle to capture the multimodality inherent in agents' future behavior, especially when the prediction horizon is long. To tackle this challenge, we first employ anchor-free queries to generate trajectory proposals in a recurrent fashion, which allows the model to utilize different scene contexts when decoding waypoints at different horizons. A refinement module then takes the trajectory proposals as anchors and leverages anchor-based queries to refine the trajectories further. By supplying adaptive and high-quality anchors to the refinement module, our query-based decoder can better deal with the multimodality in the output of trajectory prediction. Our approach ranks 1ston Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming all methods on all main metrics by a large margin. Meanwhile, our model can achieve streaming scene encoding and parallel multi-agent decoding thanks to the query-centric design ethos.
Zikang Zhou, Jianping Wang 0001, Yung-Hui Li, Yu-Kai Huang 0001
CVPR3
2023 Location-Aware Visual Question Generation with Lightweight Models
abstract
This work introduces a novel task, locationaware visual question generation (LocaVQG), which aims to generate engaging questions from data relevant to a particular geographical location.Specifically, we represent such location-aware information with surrounding images and a GPS coordinate.To tackle this task, we present a dataset generation pipeline that leverages GPT-4 to produce diverse and sophisticated questions.Then, we aim to learn a lightweight model that can address the Lo-caVQG task and fit on an edge device, such as a mobile phone.To this end, we propose a method which can reliably generate engaging questions from location-aware information.Our proposed method outperforms baselines regarding human evaluation (e.g., engagement, grounding, coherence) and automatic evaluation metrics (e.g., BERTScore, ROUGE-2).Moreover, we conduct extensive ablation studies to justify our proposed techniques for generating the dataset and solving the task.
Nicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang, I-Bin Liao, Yung-Hui Li, Lun-Wei Ku, Shao-Hua Sun
EMNLP6
2023 SVDnet: Singular Value Control and Distance Alignment Network for 3D Object Detection
abstract
The SOTA methods proposed voxelization or pillarization to regularize unordered point clouds, improving computing efficiency for LiDAR-based 3D object detection. However, they usually trade partial accuracy for speed. Thus, we bring up a new problem setting: “Is it possible to keep high detection accuracy while point-cloud quantization is applied?”. To this end, we found that the inconsistent sparsity of the point cloud over the depth distance, which is still an open question, might be the main reason. To address the inconsistency effect, we first proposed a new pillar-based vehicle detection model, named SVDnet, in which novel plug-ins are introduced in its backbone and neck. Specifically, a novel low-rank objective is designed to force the backbone to extract distance/sparsity-aware features and suppress the other feature variations among vehicle samples. Next, we alleviated the remaining feature inconsistency resulting from distance/sparsity in the neck by dynamic feature selection and adaptive feature fusion. Here, feature selection is realized by a position attention network, while feature fusion is achieved by a Distance Alignment Ratio-generation Network (DARN). Later, the selected and fused features, less sensitive to sparsity, are concatenated and fed to an SSD-like detection head. Besides, we also integrate the proposed plug-ins with multiple pillar/voxel-based methods for performance boosting. Our evaluation shows that SVDnet improves the average precision of the distant cases by 8.11% with only 0.23 milliseconds speed drop compared with PointPillars. Furthermore, the extensional results validate that our plug-ins can help SOTA pillar/voxel-based methods to gain noticeable improvement, especially for far-range objects.
Ming-Jen Chang, Chih-Jen Cheng, Ching-Chun Hsiao, Yung-Hui Li
IEEE Trans. Intell. Transp. Syst.4
2023 Referring Expression Comprehension Via Enhanced Cross-modal Graph Attention Networks
abstract
Referring expression comprehension aims to localize a specific object in an image according to a given language description. It is still challenging to comprehend and mitigate the gap between various types of information in the visual and textual domains. Generally, it needs to extract the salient features from a given expression and match the features of expression to an image. One challenge in referring expression comprehension is the number of region proposals generated by object detection methods is far more than the number of entities in the corresponding language description. Remarkably, the candidate regions without described by the expression will bring a severe impact on referring expression comprehension. To tackle this problem, we first propose a novel Enhanced Cross-modal Graph Attention Networks (ECMGANs) that boosts the matching between the expression and the entity position of an image. Then, an effective strategy named Graph Node Erase (GNE) is proposed to assist ECMGANs in eliminating the effect of irrelevant objects on the target object. Experiments on three public referring expression comprehension datasets show unambiguously that our ECMGANs framework achieves better performance than other state-of-the-art methods. Moreover, GNE is able to obtain higher accuracies of visual-expression matching effectively.
Jia Wang 0020, Jingcheng Ke, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation
abstract
Mutual learning, the related idea to knowledge distillation, is a group of untrained lightweight networks, which simultaneously learn and share knowledge to perform tasks together during training. In this paper, we propose a novel mutual learning approach, namely selective mutual learning. This is the simple yet effective approach to boost the performance of the networks for speech separation. There are two networks in the selective mutual learning method, they are like a pair of friends learning and sharing knowledge with each other. Especially, the high-confidence predictions are used to guide the remaining network while the low-confidence predictions are ignored. This helps to remove poor predictions of the two networks during sharing knowledge. The experimental results have shown that our proposed selective mutual learning method significantly improves the separation performance compared to existing training strategies including independently training, knowledge distillation, and mutual learning with the same network architecture.
Ha Minh Tan, Duc-Quang Vu, Chung-Ting Lee, Yung-Hui Li, Jia-Ching Wang
ICASSP4
2021 Enhanced (n, n)-threshold QR code secret sharing scheme based on error correction mechanism
Pengcheng Huang 0005, Chin-Chen Chang 0001, Yung-Hui Li, Yanjun Liu 0002
J. Inf. Secur. Appl.3
2020 Effectiveness evaluation of iris segmentation by using geodesic active contour (GAC)
Yuan-Tsung Chang, Timothy K. Shih, Yung-Hui Li, W. G. C. W. Kumara
J. Supercomput.3
2019 High-payload secret hiding mechanism for QR codes
Pengcheng Huang 0005, Chin-Chen Chang 0001, Yung-Hui Li, Yanjun Liu 0002
Multim. Tools Appl.3
2019 Efficient QR code authentication mechanism based on Sudoku
Pengcheng Huang 0005, Yung-Hui Li, Chin-Chen Chang 0001, Yanjun Liu 0002
Multim. Tools Appl.2
2018 Image Representation Using Supervised and Unsupervised Learning Methods on Complex Domain
abstract
Matrix factorization (MF) and its extensions have been intensively studied in computer vision and machine learning. In this paper, unsupervised and supervised learning methods based on MF technique on complex domain are introduced. Projective complex matrix factorization (PCMF) and discriminant projective complex matrix factorization (DPCMF) present two frameworks of projecting complex data to a lower dimension space. The optimization problems are formulated as the minimization of the real-valued functions of complex variables. Motivated by independence among extracted features, Fisher linear discriminant is used as hard constraint on supervised model. Experimental results on facial expression recognition (FER) show improved classification performance in comparison to real-valued features of both unsupervised and supervised NMFs.
Manh-Quan Bui, Viet-Hang Duong, Yung-Hui Li, Tzu-Chiang Tai, Jia-Ching Wang
ICASSP3
2018 Sudoku-based secret sharing approach with cheater prevention using QR code
Pengcheng Huang 0005, Chin-Chen Chang 0001, Yung-Hui Li
Multim. Tools Appl.3
2018 Efficient access control system based on aesthetic QR code
Pengcheng Huang 0005, Chin-Chen Chang 0001, Yung-Hui Li, Yanjun Liu 0002
Pers. Ubiquitous Comput.3
2017 Single channel source separation using graph sparse NMF and adaptive dictionary learning
abstract
The aim of single channel source separation is to accurately recover signals from mixtures. Non-negative matrix factorization (NMF) is a popular method to separate mixed signals using learned dictionaries. These dictionaries can be produced efficiently by sparse NMF to approximate the input signal as closely as possible. However, the literature does not consider the structure of the data in terms of the similarity among vertices of the input signal. Furthermore, state-of-art variants of NMF that are more efficient than conventional ones have not been utilized, and the learned dictionary is typically fixed in the separating phase. This strategy is not favorable because the training data and the testing data totally differ. To deal with these issues, our work proposes a method that incorporates the graph regularization into group sparsity β-NMF to improve the performance of source separation. The proposed algorithms differ from those in the literature by using an adaptive dictionary in which particular characteristics of the testing data are updated to produce newer dictionaries. Experimental results demonstrate that our proposed method is outstandingly effective in speech separation in various scenarios, relative to the baseline.
Yuan-Shan Lee, Yan-Bo Lin, Yung-Hui Li, Tzu-Chiang Tai, Jia-Ching Wang
Intell. Data Anal.4
2016 Extending the Capture Volume of an Iris Recognition System Using Wavefront Coding and Super-Resolution
abstract
Iris recognition has gained increasing popularity over the last few decades; however, the stand-off distance in a conventional iris recognition system is too short, which limits its application. In this paper, we propose a novel hardware-software hybrid method to increase the stand-off distance in an iris recognition system. When designing the system hardware, we use an optimized wavefront coding technique to extend the depth of field. To compensate for the blurring of the image caused by wavefront coding, on the software side, the proposed system uses a local patch-based super-resolution method to restore the blurred image to its clear version. The collaborative effect of the new hardware design and software post-processing showed great potential in our experiment. The experimental results showed that such improvement cannot be achieved by using a hardware-or software-only design. The proposed system can increase the capture volume of a conventional iris recognition system by three times and maintain the system's high recognition rate.
Sheng-Hsun Hsieh, Yung-Hui Li, Chung-Hao Tien, Chin-Chen Chang 0001
IEEE Trans. Cybern.2
2014 Heterogeneous IRIS recognition using heterogeneous eigeniris and sparse representation
abstract
When the iris images for training and testing are acquired by different iris image sensors, the recognition rate will be degraded and not as good as the one when both sets of images are acquired by the same image sensors. Such problem is called “heterogeneous iris recognition”. In this paper, we propose two novel patch-based heterogeneous dictionary learning methods using heterogeneous eigeniris and sparse representation which learn the basic atoms in iris textures across different image sensors and build connections between them. After such connections are built, at testing stage, it is possible to hallucinate (synthesize) iris images across different sensors. By matching training images with hallucinated images, the recognition rate can be successfully enhanced. Experimenting with an iris database consisting of 3015 images, we show that the EER is decreased 23.9% relatively by the proposed method using sparse representation, which proves the effectiveness of the proposed image hallucination method.
Bo-Ren Zheng, Dai-Yan Ji, Yung-Hui Li
ICASSP3