EDBT 2026 Demo / reviewers in the wild / expert
Kuan-Chih Huang
dblp:58/9427
· DBLP profile ↗
14ranked-venue papers
8as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reason3D: Searching and Reasoning 3D Segmentation via Large Language ModelabstractRecent advancements in multimodal large language models (LLMs) have demonstrated significant potential across various domains, particularly in concept reasoning. However, their applications in understanding 3D environments remain limited, primarily offering textual or numerical outputs without generating dense, informative segmentation masks. This paper introduces Reason3D, a novel LLM designed for comprehensive 3D understanding. Reason3D processes point cloud data and text prompts to produce textual responses and segmentation masks, enabling advanced tasks such as 3D reasoning segmentation, hierarchical searching, express referring, and question answering with detailed mask outputs. We propose a hierarchical mask decoder that employs a coarse-to-fine approach to segment objects within expansive scenes. It begins with a coarse location estimation, followed by object mask estimation, using two unique tokens predicted by LLMs based on the textual query. Experimental results on large-scale ScanNet and Matterport3D datasets validate the effectiveness of our Reason3D across various tasks. Kuan-Chih Huang, Xiangtai Li, Lu Qi 0001, Shuicheng Yan, Ming-Hsuan Yang 0001 |
3DV | 1 |
| 2024 | PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionabstractRecent temporal LiDAR-based 3D object detectors achieve promising performance based on the two-stage proposal-based approach. They generate 3D box candidates from the first-stage dense detector, followed by different temporal aggregation methods. However, these approaches require per-frame objects or whole point clouds, posing challenges related to memory bank utilization. Moreover, point clouds and trajectory features are combined solely based on concatenation, which may neglect effective interactions between them. In this paper, we propose a point-trajectory transformer with long short-term memory for efficient temporal 3D object detection. To this end, we only utilize point clouds of current-frame objects and their historical trajectories as input to minimize the memory bank storage requirement. Furthermore, we introduce modules to encode trajectory features, focusing on long short-term and future-aware perspectives, and then effectively aggregate them with point cloud features. We conduct extensive experiments on the large-scale Waymo dataset to demon-strate that our approach performs well against state-of-the-art methods. Code and models will be made publicly available at https://github.com/kuanchihhuang/Ptt. Kuan-Chih Huang, Weijie Lyu, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
CVPR | 1 |
| 2024 | Weakly Supervised 3D Object Detection via Multi-level Visual Guidance
Kuan-Chih Huang, Yi-Hsuan Tsai, Ming-Hsuan Yang 0001 |
ECCV (1) | 1 |
| 2024 | Real-time coronary artery segmentation in CAG images: A semi-supervised deep learning strategy
Chih-Kuo Lee, Jhen-Wei Hong, Chia-Ling Wu, Jia-Ming Hou, Yen-An Lin, Kuan-Chih Huang, Po-Hsuan Tseng |
Artif. Intell. Medicine | 6 |
| 2023 | Delving into Motion-Aware Matching for Monocular 3D Object TrackingabstractRecent advances of monocular 3D object detection facilitate the 3D multi-object tracking task based on lowcost camera sensors. In this paper, we find that the motion cue of objects along different time frames is critical in 3D multi-object tracking, which is less explored in existing monocular-based approaches. To this end, we propose MoMA-M3T, a framework that mainly consists of three motion-aware components. First, we represent the possible movement of an object related to all object tracklets in the feature space as its motion features. Then, we further model the historical object tracklet along the time frame in a spatial-temporal perspective via a motion transformer. Finally, we propose a motion-aware matching module to associate historical object tracklets and current observations as final tracking results. We conduct extensive experiments on the nuScenes and KITTI datasets to demonstrate that our MoMA-M3T achieves competitive performance against state-of-the-art methods. Moreover, the proposed tracker is flexible and can be easily plugged into existing image-based 3D object detectors without re-training. Code and models are available at https://github.com/kuanchihhuang/MoMA-M3T. Kuan-Chih Huang, Ming-Hsuan Yang 0001, Yi-Hsuan Tsai |
ICCV | 1 |
| 2022 | MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerabstractMonocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the additional computational burden and achieve limited performance caused by inaccurate depth priors. To alleviate this, we propose MonoDTR, a novel end-to-end depth-aware transformer network for monocular 3D object detection. It mainly consists of two components: (1) the Depth-Aware Feature Enhancement (DFE) module that implicitly learns depth-aware features with auxiliary supervision without requiring extra computation, and (2) the Depth-Aware Transformer (DTR) module that globally integrates context- and depth-aware features. Moreover, different from conventional pixel-wise positional encodings, we introduce a novel depth positional encoding (DPE) to inject depth positional hints into transformers. Our proposed depth-aware modules can be easily plugged into existing image-only monocular 3D object detectors to improve the performance. Extensive experiments on the KITTI dataset demonstrate that our approach outperforms previous state-of-the-art monocular-based methods and achieves real-time detection. Code is available at https://github.com/kuanchihhuang/.MonoDTR. Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. Hsu |
CVPR | 1 |
| 2022 | $\mathrm {D^2ADA}$: Dynamic Density-Aware Active Domain Adaptation for Semantic Segmentation
Tsung-Han Wu, Yi-Syuan Liou, Shao-Ji Yuan, Hsin-Ying Lee 0002, Tung-I Chen, Kuan-Chih Huang, Winston H. Hsu |
ECCV (29) | 6 |
| 2022 | Effects of Multisensory Distractor Interference on Attentional DrivingabstractDistracted driving refers to multisensory integration and attention shifts between attentional driving and different interferences from different modalities, including visual and auditory stimuli. Here, we compared the behavioral performance with interacting multisensory distractors during attentional driving. Then, the independent component analysis (ICA) and event-related spectral perturbation (ERSP) were applied to investigate the neural oscillation changes. The behavioral results showed that the response times (RTs) increased when distractors appeared in response to attentional driving. Moreover, the RTs were longer when the distractor interference was presented in the auditory modality compared with the visual modality. Eye movement intervals showed shorter tracking saccades under distractor interference. These results may indicate that attentional driving performance was impaired under the exposure to multisensory distractor interference. The ERSPs under visual and auditory distraction exposure showed decreased beta power in the frontal area, increased theta and delta power in the central area, and decreased alpha power in the parietal area. During this process, distracted driving under cross-modal sensory interference required more neural oscillation involvement. Moreover, the visual modality showed increased gamma power in the frontal, central, parietal and occipital areas, while the auditory modality showed decreased gamma power in the frontal area, indicating that auditory interference could intervene in top-down attentional processing. Chin-Teng Lin, Yanqiu Tian, Yu-Kai Wang, Tien-Thong Nguyen Do, Yao-Lung Chang, Jung-Tai King, Kuan-Chih Huang, Lun-De Liao |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2021 | Multi-Stream Attention Learning for Monocular Vehicle Velocity and Inter-Vehicle Distance Estimation
Kuan-Chih Huang, Winston H. Hsu |
BMVC | 1 |
| 2021 | LAFFNet: A Lightweight Adaptive Feature Fusion Network for Underwater Image EnhancementabstractUnderwater image enhancement is an important low-level computer vision task for autonomous underwater vehicles and remotely operated vehicles to explore and understand the underwater environments. Recently, deep convolutional neural networks (CNNs) have been successfully used in many computer vision problems, and so does underwater image enhancement. There are many deep-learning-based methods with impressive performance for underwater image enhancement, but their memory and model parameter costs are hindrances in practical application. To address this issue, we propose a lightweight adaptive feature fusion network (LAFFNet). The model is the encoder-decoder model with multiple adaptive feature fusion (AAF) modules. AAF subsumes multiple branches with different kernel sizes to generate multi-scale feature maps. Furthermore, channel attention is used to merge these feature maps adaptively. Our method reduces the number of parameters from 2.5M to 0.15M (around 94% reduction) but outperforms state-of-the-art algorithms by extensive experiments. Furthermore, we demonstrate our LAFFNet effectively improves high-level vision tasks like salience object detection and single image depth estimation. Hao-Hsiang Yang, Kuan-Chih Huang |
ICRA | 2 |
| 2021 | Multi-Scale Aggregation with Self-Attention Network for Modeling Electrical Motor DynamicsabstractModeling induction motor dynamics is a crucial problem in the industry. The previous works mainly model the dynamics based on the physical model assumption and state equation. However, due to the complex internal structure of motors, the traditional methods cannot estimate dynamics precisely. To address this issue, we adopt a deep learning-based approach that takes the time-series motor data measured from the sensor to estimate the dynamics without making any assumptions about the motor’s interior. In this paper, we propose a multi-scale feature aggregation with self-attention network (MASNet) to deal with modeling motor dynamics. First, our model extracts multi-scale features from motor signals by various convolutional kernel sizes. Then, the proposed adaptive feature aggregation module fuses different receptive signals effectively. To further refine these high-level motor features, two-stream components, containing Bidirectional LSTM and self-attention module, are applied to improve motor contextual information. Moreover, we present a novel temporal relative loss to enhance consecutive signal consistency, which improves the performance of modeling dynamics. To deploy the service in the real-world scenario, our network is very lightweight and reduces the number of parameters from 0.62M to 0.09M (around 85% reduction) but outperforms state-of-the-art algorithms by extensive experiments on simulation and real-world motor datasets. Kuan-Chih Huang, Hao-Hsiang Yang |
IROS | 1 |
| 2021 | Leveraging Auxiliary Information from EMR for Weakly Supervised Pulmonary Nodule Detection
Hao-Hsiang Yang, Fu-En Wang, Cheng Sun 0004, Kuan-Chih Huang, Hung-Wei Chen, Hung-Chih Chen, Chun-Yu Liao, Shih-Hsuan Kao, Yu-Chiang Frank Wang, Chou-Chin Lan |
MICCAI (7) | 4 |
| 2016 | An EEG-Based Fatigue Detection and Mitigation SystemabstractResearch has indicated that fatigue is a critical factor in cognitive lapses because it negatively affects an individual's internal state, which is then manifested physiologically. This study explores neurophysiological changes, measured by electroencephalogram (EEG), due to fatigue. This study further demonstrates the feasibility of an online closed-loop EEG-based fatigue detection and mitigation system that detects physiological change and can thereby prevent fatigue-related cognitive lapses. More importantly, this work compares the efficacy of fatigue detection and mitigation between the EEG-based and a nonEEG-based random method. Twelve healthy subjects participated in a sustained-attention driving experiment. Each participant's EEG signal was monitored continuously and a warning was delivered in real-time to participants once the EEG signature of fatigue was detected. Study results indicate suppression of the alpha- and theta-power of an occipital component and improved behavioral performance following a warning signal; these findings are in line with those in previous studies. However, study results also showed reduced warning efficacy (i.e. increased response times (RTs) to lane deviations) accompanied by increased alpha-power due to the fluctuation of warnings over time. Furthermore, a comparison of EEG-based and nonEEG-based random approaches clearly demonstrated the necessity of adaptive fatigue-mitigation systems, based on a subject's cognitive level, to deliver warnings. Analytical results clearly demonstrate and validate the efficacy of this online closed-loop EEG-based fatigue detection and mitigation mechanism to identify cognitive lapses that may lead to catastrophic incidents in countless operational environments. Kuan-Chih Huang, Teng-Yi Huang, Chun-Hsiang Chuang, Jung-Tai King, Yu-Kai Wang, Chin-Teng Lin, Tzyy-Ping Jung |
Int. J. Neural Syst. | 1 |
| 2010 | The performance of visuo-motor coordination changes under force feedback assistance systemabstractIn this study, a system with force feedback assistance was adopted to improve human's motor learning. Subjects tried to perform a trajectory tracking task of visuo-motor coordination by using a joystick which involved force feedback assistance. Force feedback was proportionally generated (maximal output 9N) by the joystick and warned subjects of the deviation of the tracking task. In the experiments, motor behavior performances of assistance group and non-assistance group were taken into comparison. The results demonstrated that a joystick with force feedback assistance brought significant influence. Motor performance was improved in the motion of flexion and extension while adding the force feedback assistance. This study provided important new light on the force feedback assistance. It could be an advanced function to improve and deepen people's motor learning capabilities. Chin-Teng Lin, Chun-Ling Lin, Kuan-Chih Huang, Shi-An Chen, Jui-Hsin Tung |
ISCAS | 3 |