EDBT 2026 Demo / reviewers in the wild / expert
He-Yen Hsieh
dblp:209/1822
· DBLP profile ↗
17ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0001-7657-6549ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alfa: Attentive Low-Rank Filter Adaptation for Structure-Aware Cross-Domain Personalized Gaze EstimationabstractPre-trained gaze models learn to identify useful patterns commonly found across users, but subtle user-specific variations (i.e., eyelid shape or facial structure) can degrade model performance. Test-time personalization (TTP) adapts pre-trained models to these user-specific domain shifts using only a few unlabeled samples. Efficient fine-tuning is critical in performing this domain adaptation: data and computation resources can be limited-especially for on-device customization. While popular parameter-efficient fine-tuning (PEFT) methods address adaptation costs by updating only a small set of weights, they may not be taking full advantage of structures encoded in pre-trained filters. To more effectively leverage existing structures learned during pre-training, we reframe personalization as a process to reweight existing features rather than learning entirely new ones. We present Attentive Low-Rank Filter Adaptation (Alfa) to adapt gaze models by reweighting semantic patterns in pre-trained filters. With Alfa, singular value decomposition (SVD) extracts dominant spatial components that capture eye and facial characteristics across users. Via an attention mechanism, we need only a few unlabeled samples to adjust and reweight pre-trained structures, selectively amplifying those relevant to a target user. Alfa achieves the lowest average gaze errors across four cross-dataset gaze benchmarks, outperforming existing TTP methods and low-rank adaptation (LoRA)-based variants. We also show that Alfa's attentive low-rank methods can be applied to applications beyond vision, such as diffusion-based language models. He-Yen Hsieh, Wei-Te Mark Ting, H. T. Kung 0001 |
AAAI | 1 |
| 2025 | DFT Gaze: Distilled and Fine-Tuned Gaze Estimation for Personalization on Tiny DevicesabstractReal-time personalized gaze estimation on AR/VR devices requires both accuracy and efficiency, especially when adapting to individual users with limited personal data. This task is challenging due to low-latency requirements, the presence of dataset biases from dominant gaze directions, and risk of catastrophic forgetting during adaptation. We present Distilled and Fine-Tuned (DFT) Gaze, a lightweight model for personalized gaze estimation. Distilled from a larger teacher model, DFT Gaze reduces model size while retaining essential visual features through knowledge distillation, without relying on gaze-specific supervision. During fine-tuning, it integrates gaze-specific supervision with Adapters, reaching 281K parameters for efficient adaptation and online updates on edge devices. To mitigate dataset biases and reduce catastrophic forgetting, we introduce a clustering-based sampling that balances gaze distribution for better generalization and improves adaptation to individual gaze patterns, even with only 5 personal images. DFT Gaze outperforms state-of-the-art methods on the MPIIFaceGaze dataset for personalized gaze estimation. Despite having the smallest model size at 281K parameters, it maintains low gaze errors across other datasets, including MPIIGaze, OpenEDS2020, and AEA. At 10× smaller than its teacher model, DFT Gaze achieves fast inference, a low parameter count, and effective adaptation, making it well-suited for real-time applications in resource-constrained environments. He-Yen Hsieh, Ziyun Li 0001, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung 0001 |
ICIP | 1 |
| 2025 | FrameVoting: A Robust and Fast Method of Using Gaze Estimations to Identify Objects of InterestabstractWe introduce FrameVoting, a voting-based method for real-time, gaze-driven object identification. It is a training-free method that incurs small computation and low processing latency, making the method ideal for wearable devices. In FrameVoting, the Point of Gaze (PoG) in each frame is used to define a potential region of interest. Regions across multiple frames are compared using the Sum of Absolute Differences (SAD) as a similarity measure. Each frame votes for the region from each of the other frames that is most similar to the region in the current frame, and only the region receiving the most votes is considered as the user’s region of interest and sent to a classifier for inference. FrameVoting thus eliminates the need for frame-by-frame bounding box retrieval and object detection required by traditional methods, thereby reducing computation overhead and latency. The method is robust, as it eliminates the need for threshold-tuning to determine whether gaze estimations are focused on a specific object. Further, the method is efficient and fast, as inference is only performed on the most-voted region, and the SAD computation is highly parallelizable. Our experiments on the AEA Dataset demonstrate that FrameVoting reduces the frequency of inferences by 95.6% compared to frame-by-frame object detection, while still accurately identifying the user’s objects of interest in real-time at 30 fps on a Raspberry Pi 5. Kao-Den Chang, He-Yen Hsieh, H. T. Kung 0001, Ziyun Li 0001, Sai Qian Zhang |
ISCAS | 2 |
| 2023 | One-Shot Action Detection via Attention Zooming InabstractHinted by a modest support set, few-shot action detection (FSAD) aims at localizing the action instances of unseen classes within an untrimmed query video. Existing FSAD techniques mostly rely on generating a set of class-agnostic action proposals from the query video and then finding the most plausible ones by assessing their correlation to the support set. Such two-stage approaches are feasible but not efficient, largely due to neglecting the support information in generating the proposals. This work focuses on the one-shot image scenario and introduces the attention zooming in strategy to effectively and progressively carry out support-query cross-attention while generating proposals. The resulting one-stage model yields high-quality action proposals for boosting one-shot action detection (OSAD) performance. Our extensive experiments on the ActivityNet-1.3 and THUMOS-14 datasets demonstrate that the proposed framework can achieve state-of-the-art performance in tack-ling challenging image-based OSAD tasks. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
ICASSP | 1 |
| 2023 | Aggregating Bilateral Attention for Few-Shot Instance LocalizationabstractAttention filtering under various learning scenarios has proven advantageous in enhancing the performance of many neural network architectures. The mainstream attention mechanism is established upon the non-local block, also known as an essential component of the prominent Transformer networks, to catch long-range correlations. However, such unilateral attention is often hampered by sparse and obscure responses, revealing insufficient dependencies across images/patches, and high computational cost, especially for those employing the multi-head design. To overcome these issues, we introduce a novel mechanism of aggregating bilateral attention (ABA) and validate its usefulness in tackling the task of few-shot instance localization, reflecting the underlying query-support dependency. Specifically, our method facilitates uncovering informative features via assessing: i) an embedding norm for exploring the semantically-related cues; ii) context awareness for correlating the query data and support regions. ABA is then carried out by integrating the affinity relations derived from the two measurements to serve as a lightweight but effective query-support attention mechanism with high localization recall. We evaluate ABA on two localization tasks, namely, few-shot action localization and one-shot object detection. Extensive experiments demonstrate that the proposed ABA achieves superior performances over existing methods. He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu |
WACV | 1 |
| 2022 | Self-supervised Sparse Representation for Video Anomaly Detection
Jhih-Ciang Wu, He-Yen Hsieh, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu |
ECCV (13) | 2 |
| 2022 | Contextual Proposal Network for Action LocalizationabstractThis paper investigates the problem of Temporal Action Proposal (TAP) generation, which aims to provide a set of high-quality video segments that potentially contain actions events locating in long untrimmed videos. Based on the goal to distill available contextual information, we introduce a Contextual Proposal Network (CPN) composing of two context-aware mechanisms. The first mechanism, i.e., feature enhancing, integrates the inception-like module with long-range attention to capture the multi-scale temporal contexts for yielding a robust video segment representation. The second mechanism, i.e., boundary scoring, employs the bi-directional recurrent neural networks (RNN) to capture bi-directional temporal contexts that explicitly model actionness, background, and confidence of proposals. While generating and scoring proposals, such bi-directional temporal contexts are helpful to retrieve high-quality proposals of low false positives for covering the video action instances. We conduct experiments on two challenging datasets of ActivityNet-1.3 and THUMOS-14 to demonstrate the effectiveness of the proposed Contextual Proposal Network (CPN). In particular, our method respectively surpasses state-of-the-art TAP methods by 1.54% AUC on ActivityNet-1.3 test split and by 0.61% AR@200 on THUMOS-14 dataset. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
WACV | 1 |
| 2021 | Progressive Contextual Excitation for Smart Farming Application
Chia-Hung Bai, Setya Widyawan Prakosa, He-Yen Hsieh, Jenq-Shiou Leu, Wen-Hsien Fang |
CAIP (1) | 3 |
| 2021 | Adaptive Image Transformer for One-Shot Object DetectionabstractOne-shot object detection tackles a challenging task that aims at identifying within a target image all object instances of the same class, implied by a query image patch. The main difficulty lies in the situation that the class label of the query patch and its respective examples are not available in the training data. Our main idea leverages the concept of language translation to boost metric-learning-based detection methods. Specifically, we emulate the language translation process to adaptively translate the feature of each object proposal to better correlate the given query feature for discriminating the class-similarity among the proposal-query pairs. To this end, we propose the Adaptive Image Transformer (AIT) module that deploys an attention-based encoder-decoder architecture to simultaneously explore intra-coder and inter-coder (i.e., each proposal-query pair) attention. The adaptive nature of our design turns out to be flexible and effective in addressing the one-shot learning scenario. With the informative attention cues, the proposed model excels in predicting the class-similarity between the target image proposals and the query image patch. Though conceptually simple, our model significantly outperforms a state-of-the-art technique, improving the unseen-class object classification from 63.8 mAP and 22.0 AP50 to 72.2 mAP and 24.3 AP50 on the PASCAL-VOC and MS-COCO benchmark datasets, respectively. Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
CVPR | 2 |
| 2021 | Referring Image Segmentation via Language-Driven AttentionabstractThis paper aims to tackle the problem of referring image segmentation, which is targeted at reasoning the region of interest referred by a query natural language sentence. One key issue to address the referring image segmentation is how to establish the cross-modal representation for encoding the two modalities, namely, the query sentence and the input image. Most existing methods are designed to concatenate the features from each modality or to gradually encode the cross-modal representation concerning each word’s effect. In contrast, our approach leverages the correlation between the two modalities for constructing the cross-modal representation. To make the resulting cross-modal representation more discriminative for the segmentation task, we propose a novel mechanism of language-driven attention to encode the cross-modal representation for reflecting the attention between every single visual element and the entire query sentence. The proposed mechanism, named as Language-Driven Attention (LDA), first decouples the cross-modal correlation to channel-attention and spatial-attention and then integrates the two attentions for obtaining the cross-modal representation. The channel attention and the spatial attention respectively reveal how sensitive each channel or each pixel of a particular feature map is with respect to the query sentence. With a proper fusion of the two kinds of feature attention, the proposed LDA model can effectively guide the generation of the final cross-modal representation. The resulting representation is further strengthened for capturing the multi-receptive-field and multi-level-semantic for the intended segmentation. We assess our referring image segmentation model on four public benchmark datasets, and the experimental results show that our model achieves state-of-the-art performance Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu |
ICRA | 2 |
| 2021 | Robust Network Intrusion Detection Scheme Using Long-Short Term Memory Based Convolutional Neural Networks
Chia-Ming Hsu, Muhammad Zulfan Azhari, He-Yen Hsieh, Setya Widyawan Prakosa, Jenq-Shiou Leu |
Mob. Networks Appl. | 3 |
| 2021 | Implementing a real-time image captioning service for scene identification using embedded system
He-Yen Hsieh, Sheng-An Huang, Jenq-Shiou Leu |
Multim. Tools Appl. | 1 |
| 2021 | Correction to: Implementing a real-time image captioning service for scene identification using embedded system
He-Yen Hsieh, Sheng-An Huang, Jenq-Shiou Leu |
Multim. Tools Appl. | 1 |
| 2020 | Temporal Action Proposal Generation Via Deep Feature EnhancementabstractTemporal action proposal generation (TAPG) is a challenging problem for analyzing video content. It aims to localize the video segments which are likely to contain actions or events. Intuitively, making a satisfying prediction of these video segments is directly relies on their representation quality. A typical representation of a video segment is applying a two-stream feature, which comprises appearance and motion information. Rather than directly concatenating the two-stream features as the previous methods, we illustrate a feature-aggregation network (FA-Net) concerning the feature-relation among neighboring video segments for obtaining the high-quality representation that better characterizing the actions or events. Further, we design a feature-expansion network (FE-Net) to extract multi-granularity features for retrieving the proposals of high action-instance covering confidence. We evaluate our approach on two challenging datasets: ActivityNet-1.3 and THUMOS-14. The experiments showed that the proposed approach consistently outperforms the existing state-of-the-art TAPG methods. He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu |
ICIP | 1 |
| 2019 | Implementing a Real-Time Image Captioning Service for Scene Identification Using Embedded SystemabstractThis work aims to implement a real-time scene identification system using an image captioning model on an embedded system. The image captioning model can translate the image captured by a webcam installed on the embedded system into a human-readable sentence immediately. Users can get the information quickly by reading only the sentences. There are two stages in the image captioning model. First, a deep neural network extracts features from images captured from the webcam. Second, a long-short term memory generates the corresponding sentence. Due to the portability of the embedded system, our scene identification system can be placed anywhere at home or in the company. We evaluate the execution time in different aspects on several embedded systems and demonstrate the generated sentences from the captured images by our scene identification system. He-Yen Hsieh, Jenq-Shiou Leu, Sheng-An Huang |
SECON | 1 |
| 2018 | Towards the Implementation of Recurrent Neural Network Schemes for WiFi Fingerprint-Based Indoor PositioningabstractThe rapid development of Indoor Positioning System has attracted researcher to develop a robust scheme to predict the location based on Received Signal Strength Indicator (RSSI) signal. A lot of research topics presented in many journals and conferences by many researchers concern indoor positioning system as a main topic [1], [2]. Currently, the study related to find the robust algorithm for indoor positioning system becomes a high demand topic in several conferences. Our work intents to evaluate the effectiveness of Recurrent Neural Network (RNN) as a deep learning technique to be implemented in this field. In addition, LSTM as a variant of RNN scheme is also implemented. The purpose of this implementation is to explore both LSTM and original RNN to be utilized for localization in indoor positioning scheme, especially for Wifi Fingerprinting Dataset. From all evaluations, our proposed approach could get 99.7% accuracy for predicting which floor the sensor belongs to. In addition, the distance errors of our scheme are around 2.5-2.7 meters. He-Yen Hsieh, Setya Widyawan Prakosa, Jenq-Shiou Leu |
VTC Fall | 1 |
| 2017 | Prediction of Station Level Demand in a Bike Sharing System Using Recurrent Neural NetworksabstractBike sharing systems have been widely applied to many cities and brought convenience to local citizens for short-ranged transportation. The bike shortage problem due to uneven bikes distribution is one of the biggest challenges in bike sharing systems. In this paper, we focus on station level prediction for each bike station. The proposed architecture is based on Recurrent Neural Network (RNN) and we use only one model to predict both rental and return demand for every station at once which is efficient for online balancing strategies. Without considering the global level bike distribution, the MAPE/RMSLE of the sum over the demand of each station may be too high for rebalancing strategies but the MAE/RMSE are satisficing at station level. Our evaluation shows that the proposed methods meet satisfied results at station level and global level in New York Citi Bike dataset. Po-Chuan Chen, He-Yen Hsieh, Xanno Kharis Sigalingging, Jenq-Shiou Leu |
VTC Spring | 2 |