Hung-Min Hsu

dblp:139/5774 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-7180-1396ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Adaptive Edge Intelligence for Intersection Safety: Real-Time Dilemma Zone Management via DV-EISOS
abstract
Dilemma zone (DZ) protection is vital for intersection safety, yet legacy detectors often lack the spatial resolution and responsiveness needed to track fast-changing vehicle dynamics. This paper presents DV-EISOS, an edge-deployed, vision-driven system for real-time DZ detection and adaptive signal control. Running on Jetson AGX Orin, DV-EISOS integrates YOLOv8 for vehicle detection, ByteTrack for multi-object tracking, homography-based mapping, and time-to-intersection logic to convert RTSP camera feeds into actionable kinematic metrics with millisecond-level latency. Through an NTCIP/SNMP interface, the system issues yellow-phase extensions only when approaching vehicles are assessed as high risk. A field deployment at a signalized intersection in Bellevue, WA demonstrated accurate speed estimation (RMSE < 2 mph), timely phase adaptation, and responses to ~10% of potential DZ events. These results indicate that edge AI can deliver scalable, low-latency safety control using existing cameras in real-world urban environments.
Luyang Gong, Hung-Min Hsu, Yinhai Wang
SEC2
2024 2D-to-3D Mutual Iterative Optimization for 3D Multi-camera Multiple People Tracking
abstract
Multi-camera Multiple People Tracking (MMPT) is a challenging task in advanced visual monitoring systems. The main challenge of MMPT is how to accurately match the single-camera trajectories generated from different viewpoints and establish one global and complete cross-camera trajectory for each target, i.e., the multi-camera trajectory matching problem. In this paper, we propose a novel framework to solve this problem using a scene-aware multiple object tracking. Furthermore, unlike most existing methods that purely use single-camera trajectories for multiple object tracking, we introduce a new multiple camera compensation mechanism 2D-to-3D Mutual Iterative Optimization for MMPT (MIO-MMPT) to enhance the person tracking results, which exploits the crucial multi-camera relationships among the human trajectories appearing in different cameras both robustly and automatically. Based on the camera calibration, we can project the 2D coordinate into 3D coordinate to achieve more reliable tracking results for each person. Once we have the 3D tracking results, we can re-project to 2D coordinate of each camera to solve the missing detection issues from occlusion or the blind spot of the camera. According to our experimental results, the proposed method achieves a new state-of-the-art on ICCV 2021 MMPT dataset with MOTA of 95% and IDF1 of 96%.
Hung-Min Hsu, Zhongwei Cheng, Xinyu Yuan
AVSS1
2024 Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
abstract
The diffusion model, a new generative modeling paradigm, has achieved significant success in generating images, audio, video, and text. It has been adapted for sequence-to-sequence text generation (Seq2Seq) through DiffuSeq, termed the S2S-Diffusion model. Existing S2S-Diffusion models predominantly rely on fixed or hand-crafted rules to schedule noise during the diffusion and denoising processes. However, these models are limited by non-contextualized noise, which fails to fully consider the characteristics of Seq2Seq tasks. In this paper, we propose the Meta-Diffu$B$ framework—a novel scheduler-exploiter S2S-Diffusion paradigm designed to overcome the limitations of existing S2S-Diffusion models. We employ Meta-Exploration to train an additional scheduler model dedicated to scheduling contextualized noise for each sentence. Our exploiter model, an S2S-Diffusion model, leverages the noise scheduled by our scheduler model for updating and generation. Meta-Diffu$B$ achieves state-of-the-art performance compared to previous S2S-Diffusion models and fine-tuned pre-trained language models (PLMs) across four Seq2Seq benchmark datasets. We further investigate and visualize the impact of Meta-Diffu$B$'s noise scheduling on the generation of sentences with varying difficulties. Additionally, our scheduler model can function as a "plug-and-play" model to enhance DiffuSeq without the need for fine-tuning during the inference stage.
Yun-Yen Chuang, Hung-Min Hsu, Chen-Sheng Gu, Ling Zhen Li, Ray-I Chang, Hung-yi Lee
NeurIPS2
2023 MetaEx-GAN: Meta Exploration to Improve Natural Language Generation via Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) have been popularly researched in natural language generation, so-called Language GANs. Existing works adopt reinforcement learning (RL) based methods such as policy gradients for training Language GANs. The previous research of Language GANs usually focuses on stabilizing policy gradients or applying robust architectures (such as the large-scale pre-trained GPT-2) to achieve better performance. However, the quality and diversity of sampling are not guaranteed simultaneously. In this article, we propose a novel meta-learning-based generative adversarial network, Meta Exploration GAN (MetaEx-GAN), for ensuring the quality and diversity of sampling (sampling efficiency). In the proposed MetaEx-GAN, we develop an explorer trained by Meta Exploration to sample from the generated data to achieve better sampling efficiency. MetaEx-GAN employs MetaEx first applied to Language GANs to achieve better performance. We also propose a critical training method for MetaEx-GAN on the NLG task. According to our experimental results, MetaEx-GAN achieves state-of-the-art performance compared with existing Language GANs methods. Our experiments also demonstrate the generality of MetaEx-GAN with different architectures (involving GPT-2) and how MetaEx-GAN operates to improve Language GANs.
Yun-Yen Chuang, Hung-Min Hsu, Ray-I Chang, Hung-yi Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 LUNA: Localizing Unfamiliarity Near Acquaintance for Open-Set Long-Tailed Recognition
abstract
The predefined artificially-balanced training classes in object recognition have limited capability in modeling real-world scenarios where objects are imbalanced-distributed with unknown classes. In this paper, we discuss a promising solution to the Open-set Long-Tailed Recognition (OLTR) task utilizing metric learning. Firstly, we propose a distribution-sensitive loss, which weighs more on the tail classes to decrease the intra-class distance in the feature space. Building upon these concentrated feature clusters, a local-density-based metric is introduced, called Localizing Unfamiliarity Near Acquaintance (LUNA), to measure the novelty of a testing sample. LUNA is flexible with different cluster sizes and is reliable on the cluster boundary by considering neighbors of different properties. Moreover, contrary to most of the existing works that alleviate the open-set detection as a simple binary decision, LUNA is a quantitative measurement with interpretable meanings. Our proposed method exceeds the state-of-the-art algorithm by 4-6% in the closed-set recognition accuracy and 4% in F-measure under the open-set on the public benchmark datasets, including our own newly introduced fine-grained OLTR dataset about marine species (MS-LT), which is the first naturally-distributed OLTR dataset revealing the genuine genetic relationships of the classes.
Jiarui Cai, Yizhou Wang 0005, Hung-Min Hsu, Jenq-Neng Hwang, Kelsey Magrane, Craig S. Rose
AAAI3
2022 GAITTAKE: Gait Recognition by Temporal Attention and Keypoint-Guided Embedding
abstract
Gait recognition, which refers to the recognition or identification of a person based on their body shape and walking styles, derived from video data captured from a distance, is widely used in crime prevention, forensic identification, and social security. However, to the best of our knowledge, most of the existing methods use appearance, posture and temporal feautures without considering a learned temporal attention mechanism for global and local information fusion. In this paper, we propose a novel gait recognition framework, called Temporal Attention and Keypoint-guided Embedding (GaitTAKE), which effectively fuses temporal-attention-based global and local appearance feature and temporal aggregated human pose feature. Experimental results show that our proposed method achieves a new SOTA in gait recognition with rank-1 accuracy of 98.0% (normal), 97.5% (bag) and 92.2% (coat) on the CASIA-B gait dataset; 90.4% accuracy on the OU-MVLP gait dataset.
Hung-Min Hsu, Yizhou Wang 0005, Cheng-Yen Yang, Jenq-Neng Hwang, Le Uyen Thuc Hoang, Kwang-Ju Kim
ICIP1
2021 ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving Applications
abstract
The Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community.
Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu
ICMR6
2021 Multi-Target Multi-Camera Tracking of Vehicles Using Metadata-Aided Re-ID and Trajectory-Based Camera Link Model
abstract
In this paper, we propose a novel framework for multi-target multi-camera tracking (MTMCT) of vehicles based on metadata-aided re-identification (MA-ReID) and the trajectory-based camera link model (TCLM). Given a video sequence and the corresponding frame-by-frame vehicle detections, we first address the isolated tracklets issue from single camera tracking (SCT) by the proposed traffic-aware single-camera tracking (TSCT). Then, after automatically constructing the TCLM, we solve MTMCT by the MA-ReID. The TCLM is generated from camera topological configuration to obtain the spatial and temporal information to improve the performance of MTMCT by reducing the candidate search of ReID. We also use the temporal attention model to create more discriminative embeddings of trajectories from each camera to achieve robust distance measures for vehicle ReID. Moreover, we train a metadata classifier for MTMCT to obtain the metadata feature, which is concatenated with the temporal attention based embeddings. Finally, the TCLM and hierarchical clustering are jointly applied for global ID assignment. The proposed method is evaluated on the CityFlow dataset, achieving IDF1 76.77%, which outperforms the state-of-the-art MTMCT methods.
Hung-Min Hsu, Jiarui Cai, Yizhou Wang 0005, Jenq-Neng Hwang, Kwang-Ju Kim
IEEE Trans. Image Process.1
2020 Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link Model
abstract
Multi-target multi-camera tracking (MTMCT), i.e., tracking multiple targets across multiple cameras, is a crucial technique for smart city applications. In this paper, we propose an effective and reliable MTMCT framework for vehicles, which consists of a traffic-aware single camera tracking (TSCT) algorithm, a trajectory-based camera link model (CLM) for vehicle re-identification (ReID), and a hierarchical clustering algorithm to obtain the cross camera vehicle trajectories. First, the TSCT, which jointly considers vehicle appearance, geometric features, and some common traffic scenarios, is proposed to track the vehicles in each camera separately. Second, the trajectory-based CLM is adopted to facilitate the relationship between each pair of adjacently connected cameras and add spatio-temporal constraints for the subsequent vehicle ReID with temporal attention. Third, the hierarchical clustering algorithm is used to merge the vehicle trajectories among all the cameras to obtain the final MTMCT results. Our proposed MTMCT is evaluated on the CityFlow dataset and achieves a new state-of-the-art performance with IDF1 of 74.93%.
Hung-Min Hsu, Yizhou Wang 0005, Jenq-Neng Hwang
ACM Multimedia1
2017 Rearrange Social Overloaded Posts to Prevent Social Overload
abstract
According to the latest investigation, there are 1.7 million active social network users in Taiwan. Previous researches indicated social network posts have a great impact on users, and mostly, the negative impact is from the rising demands of social support, which further lead to heavier social overload. In this study, we propose social overloaded posts detection model (SODM) by deploying the latest text mining and deep learning techniques to detect the social overloaded posts and, then with the developed social overload prevention system (SOS), the social overload posts and non-social overload ones are rearranged with different sorting methods to prevent readers from excessive demands of social support or social overload. The empirical results show that our SOS helps readers to alleviate social overload when reading via social media.
Yun-Yen Chuang, Hung-Min Hsu, Tsui-Ying Lin, Ray-I Chang
ASONAM2
2016 Frame Dispatcher: A Multi-frame Classification System for Social Movement by Using Microblogging Data
abstract
Framing is a phenomenon that is studied and debated widely in sociology and political science. It refers to the manner in which audiences interpret information and justify their claims or activities. The subconscious influence of framing might lead to opinion changes and social movements. However, multi-frame classification on microblogging data has not yet been investigated. In this study, we aim to classify a large number of posts into frames. We describe in detail the implementation of a new algorithm for multi-frame classification tasks called Frame Dispatcher, which aims to classify microblogging data into frames. In our experiments, we extracted over 15,000 posts from approximately 200 Facebook fan pages concerning an anti-curriculum student movement. The experimental results show that Frame Dispatcher can classify microblogging data into frames efficiently and effectively.
Hung-Min Hsu, Wei-Sheng Zeng, Chen-Shuo Hung, Dung-Sheng Chen, Ray-I Chang, Shian-Hua Lin, Jan-Ming Ho
WI1
2013 Constructing mobile-oriented catalog in m-commerce using LDA-based self-adaptive genetic algorithm
abstract
The purpose of this paper is to develop a method to recommend products to customer via mobile devices. Collaborative recommendation is known as an effective way to recommend products. In this paper, we use the concept of collaborative recommendation to develop Mobile-Oriented Catalog (MOC). The proposed method is made from aggregating similar purchasing records to optimize combination of goods on mobile devices. This paper illustrates how to design attractive and collaborative catalog to recommend items by using Latent Dirichlet Allocation (LDA) based self-adaptive genetic algorithm (LDA-SAGA). LDA-SAGA is consisted of topic modeling concept and self-adaptive genetic algorithm. We use LDA as our topic modeling algorithm to construct MOC as a result that it is the simplest topic model. Our experimental evaluation on synthetic and real data shows that using preference as topic concept is effective. LDA-SAGA is especially outstanding with large number of customers and products. Finally, we compare the MOC which is used on mobile application (APP) of Amazon with the one used on Taobao and discuss the characteristics of their design. Different design of user interface on APP can lead to different scope of fitness value which is capable of explaining different market strategies of Taobao and Amazon.
Hung-Min Hsu, Ray-I Chang, Jan-Ming Ho
IJCNN1