Cheng-Yen Yang

dblp:03/10433 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
14since 2021 · last 2025
0009-0004-2631-6756ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2025 A Depth-Aware Robust Multi-Object Tracker for Crowded Scene by Re-Prioritizing Association Order
abstract
Occlusion remains a major challenge in online Multi-Object Tracking (MOT), where existing multi-stage association methods often rely on detection confidence scores despite their weak correlation with occlusion, leading to frequent errors. We propose DARUMA, a depth-aware MOT framework that prioritizes non-occluded objects using occlusion-aware association by re-prioritizing the matching order and refines association with a depth-weighted cost metric for improved robustness in occluded and depth-varying environments. Additionally, we introduce Generic Observation-Centric Momentum (GOCM), which integrates depth-aware velocity estimation and confidence-weighted historical observations to enhance motion modeling. Our method can integrates into existing MOT frameworks, improving association robustness without additional supervision.Extensive evaluations on DanceTrack demonstrate that DARUMA achieves state-of-the-art performance, particularly in complex, occlusion-heavy scenarios.
Cheng-Yen Yang, Hsiang-Wei Huang, Kuang-Ming Chen, Kunjun Li, Farron Wallace, Chung-I Huang, Jenq-Neng Hwang
AVSS1
2025 Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression
abstract
Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and diversity of available 2D image data in contrast to the more limited 3D data. Multi-frame VLMs, although achieving superior performance, suffer from the difficulty of retaining all the detailed visual information in the 3D scene while limiting the number of visual tokens. Common methods such as token pooling, reduce visual token usage but often lead to information loss, impairing the model’s ability to preserve visual details essential for 3D question answering tasks. To address this, we propose voxel-based Dynamic Token Compression (DTC), which combines 3D spatial priors and visual semantics to achieve over 90% reduction in visual tokens usage for current multi-frame VLMs. Our method maintains performance comparable to state-of-the-art models on 3D question answering benchmarks including OpenEQA and ScanQA, demonstrating its effectiveness.
Hsiang-Wei Huang, Fu-Chen Chen, Wenhao Chai, Che-Chun Su, Sanghun Jung, Cheng-Yen Yang, Jenq-Neng Hwang, Min Sun 0001, Cheng-Hao Kuo
CVPR7
2025 MambaMOT: State-Space Model as Motion Predictor for Multi-Object Tracking
abstract
In the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusions prevalent in dynamic environments like sports and dance. This paper explores the possibilities of replacing the Kalman filter with a learning-based motion model that effectively enhances tracking accuracy and adaptability beyond the constraints of Kalman filter-based tracker. In this paper, our proposed method MambaMOT and MambaMOT+, demonstrate advanced performance on challenging MOT datasets such as DanceTrack and SportsMOT, showcasing their ability to handle intricate, nonlinear motion patterns and frequent occlusions more effectively than traditional methods.
Hsiang-Wei Huang, Cheng-Yen Yang, Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang
ICASSP2
2025 ToSA: Token Merging with Spatial Awareness
abstract
Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual token’s feature similarity for token merging, overlooking the potential of integrating spatial information, which can serve as a reliable criterion for token merging in the early layers of ViT, where the visual tokens only possess weak visual information. In this paper, we propose ToSA, a novel token merging method that combines both semantic and spatial awareness to guide the token merging process. ToSA leverages the depth image as input to generate pseudo spatial tokens, which serve as auxiliary spatial information for the visual token merging process. With the introduced spatial awareness, ToSA achieves a more informed merging strategy that better preserves critical scene structure. Experimental results demonstrate that ToSA outperforms previous token merging methods across multiple benchmarks on visual and embodied question answering while largely reducing the runtime of the ViT, making it an efficient solution for ViT acceleration. The code will be available at: https://github.com/hsiangwei0903/ToSA.
Hsiang-Wei Huang, Wenhao Chai, Kuang-Ming Chen, Cheng-Yen Yang, Jenq-Neng Hwang
IROS4
2025 Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
abstract
Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. Nevertheless, the coarse-to-fine methodology inherent in VAR results in exponential growth of the KV cache during inference, causing considerable memory consumption and computational redundancy. To address these bottlenecks, we introduce ScaleKV, a novel KV cache compression framework tailored for VAR architectures. ScaleKV leverages two critical observations: varying cache demands across transformer layers and distinct attention patterns at different scales. Based on these insights, ScaleKV categorizes transformer layers into two functional groups: drafters and refiners. Drafters exhibit dispersed attention across multiple scales, thereby requiring greater cache capacity. Conversely, refiners focus attention on the current token map to process local details, consequently necessitating substantially reduced cache capacity. ScaleKV optimizes the multi-scale inference pipeline by identifying scale-specific drafters and refiners, facilitating differentiated cache management tailored to each scale. Evaluation on the state-of-the-art text-to-image VAR model family, Infinity, demonstrates that our approach effectively reduces the required KV cache memory to 10% while preserving pixel-level fidelity.
Kunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng Hwang
NeurIPS3
2024 APTPose: Anatomy-aware Pre-Training for 3D Human Pose Estimation
Qing-Wen Yang, Kai-Wen Duan, Ting-Yi Lu, Cheng-Yen Yang, Jenq-Neng Hwang, Shang-Hong Lai
BMVC5
2024 2D Human Pose Estimation Calibration and Keypoint Visibility Classification
abstract
The confidence scores of 2D pose estimation are widely utilized in various fields, including multi-view 3D human pose estimation, skeleton-based human tracking, human action recognition, human re-identification, etc. Despite widespread use, confidence scores from 2D pose estimation methods are unreliable in indicating the accuracy of estimation results, particularly in occlusion situations, i.e., keypoints with high confidence scores may have low accuracy and vice versa. To address this issue, we propose a new 2D human pose estimation calibration method in this paper. Our method not only enhances the accuracy of 2D pose estimation but also aligns the confidence scores with the quality and visibility of keypoints. We achieve 77.6 mAP in the COCO val dataset, compared with 76.5 mAP of the original HRNet. For key-point visibility prediction, we can reach 89.4% accuracy, 87.6% precision, and 97.0% recall in the COCO val dataset.
Zhongyu Jiang, Haorui Ji, Cheng-Yen Yang, Jenq-Neng Hwang
ICASSP3
2024 A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Videos
abstract
Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counting datasets have been the mainstream of the current publicly available datasets. Therefore, we propose a large-scale dataset called YoutubeFish-35, which contains a total of 35 sequences of high-definition videos with high frame-per-second and more than 159,000 annotated center points across a selected variety of scenes. For bench-marking purposes, we select three mainstream methods for dense object counting and carefully evaluate them on the newly collected dataset. We propose TransVidCount, a new strong baseline that combines density and regression branches along the temporal domain in a unified framework and can effectively tackle indiscernible object counting with state-of-the-art performance on YoutubeFish-35 dataset.
Cheng-Yen Yang, Hsiang-Wei Huang, Zhongyu Jiang, Farron Wallace, Jenq-Neng Hwang
ICASSP1
2024 Boosting Online 3D Multi-Object Tracking through Camera-Radar Cross Check
abstract
In the domain of autonomous driving, the integration of multi-modal perception techniques based on data from diverse sensors has demonstrated substantial progress. Effectively surpassing the capabilities of state-of-the-art single-modality detectors through sensor fusion remains an active challenge. This work leverages the respective advantages of cameras in perspective view and radars in Bird’s Eye View (BEV) to greatly enhance overall detection and tracking performance. Our approach, Camera-Radar Associated Fusion Tracking Booster (CRAFTBooster) represents a pioneering effort to enhance radar-camera fusion in the tracking stage, contributing to improved 3D MOT accuracy. The superior experimental results on K-Radaar dataset, which exhibit 5-6% on IDF1 tracking performance gain, validate the potential of effective sensor fusion in advancing autonomous driving.
Sheng-Yao Kuan, Jen-Hao Cheng, Hsiang-Wei Huang, Wenhao Chai, Cheng-Yen Yang, Hugo Latapie, Gaowen Liu, Bing-Fei Wu, Jenq-Neng Hwang
IV5
2024 Back to Optimization: Diffusion-based Zero-Shot 3D Human Pose Estimation
abstract
Learning-based methods have dominated the 3D human pose estimation (HPE) tasks with significantly better performance in most benchmarks than traditional optimization-based methods. Nonetheless, 3D HPE in the wild is still the biggest challenge for learning-based models, whether with 2D-3D lifting, image-to-3D, or diffusion-based methods, since the trained networks implicitly learn camera intrinsic parameters and domain-based 3D human pose distributions and estimate poses by statistical average. On the other hand, the optimization-based methods estimate results case-by-case, which can predict more diverse and sophisticated human poses in the wild. By combining the advantages of optimization-based and learning-based methods, we propose the Zero-shot Diffusion-based Optimization (ZeDO) pipeline for 3D HPE to solve the problem of cross-domain and in-the-wild 3D HPE. Our multi-hypothesis ZeDO achieves state-of-the-art (SOTA) performance on Human3.6M, with minMPJPE 51.4mm, without training with any 2D-3D or image-3D pairs. Moreover, our single-hypothesis ZeDO achieves SOTA performance on 3DPW dataset with PA-MPJPE 40.3mm on cross-dataset evaluation, which even outperforms learning-based methods trained on 3DPW. Our code is available here: https://github.com/ipl-uw/ZeDO-Release.
Zhongyu Jiang, Zhuoran Zhou, Lei Li 0050, Wenhao Chai, Cheng-Yen Yang, Jenq-Neng Hwang
WACV5
2023 Multi-Object Tracking by Iteratively Associating Detections with Uniform Appearance for Trawl-Based Fishing Bycatch Monitoring
abstract
The aim of in-trawl catch monitoring for use in fishing operations is to detect, track and classify fish targets in real-time from video footage. Information gathered could be used to release unwanted bycatch in real-time. However, traditional multi-object tracking (MOT) methods have limitations, as they are developed for tracking vehicles or pedestrians with linear motions and diverse appearances, which are different from the scenarios such as livestock monitoring. Therefore, we propose a novel MOT method, built upon an existing observation-centric tracking algorithm, by adopting a new iterative association step to significantly boost the performance of tracking targets with a uniform appearance. The iterative association module is an extendable component that can be merged into most existing tracking methods. Our method offers improved performance in tracking targets with uniform appearance and outperforms state-of-the-art techniques on our underwater fish datasets as well as the MOT17 dataset, without increasing latency nor sacrificing accuracy as measured by HOTA, MOTA, and IDF1 performance metrics.
Cheng-Yen Yang, Yu Shyang Tan, Melanie J. Underwood, Charlotte Bodie, Zhongyu Jiang, Steve George, Karl Warr, Jenq-Neng Hwang, Emma Jones
ICIP1
2023 CameraPose: Weakly-Supervised Monocular 3D Human Pose Estimation by Leveraging In-the-wild 2D Annotations
abstract
To improve the generalization of 3D human pose estimators, many existing deep learning based models focus on adding different augmentations to training poses. However, data augmentation techniques are limited to the "seen" pose combinations and hard to infer poses with rare "unseen" joint positions. To address this problem, we present CameraPose, a weakly-supervised framework for 3D human pose estimation from a single image, which can not only be applied on 2D-3D pose pairs but also on 2D alone annotations. By adding a camera parameter branch, any in-the-wild 2D annotations can be fed into our pipeline to boost the training diversity and the 3D poses can be implicitly learned by reprojecting back to 2D. Moreover, CameraPose introduces a refinement network module with confidence-guided loss to further improve the quality of noisy 2D keypoints extracted by 2D pose estimators. Experimental results demonstrate that the CameraPose brings in clear improvements on cross-scenario datasets. Notably, it outperforms the baseline method by 3mm on the most challenging dataset 3DPW. In addition, by combining our proposed refinement network module with existing 3D pose estimators, their performance can be improved in cross-scenario evaluation.
Cheng-Yen Yang, Jiajia Luo, Yuyin Sun, Nan Qiao 0009, Ke Zhang 0028, Zhongyu Jiang, Jenq-Neng Hwang, Cheng-Hao Kuo
WACV1
2022 GAITTAKE: Gait Recognition by Temporal Attention and Keypoint-Guided Embedding
abstract
Gait recognition, which refers to the recognition or identification of a person based on their body shape and walking styles, derived from video data captured from a distance, is widely used in crime prevention, forensic identification, and social security. However, to the best of our knowledge, most of the existing methods use appearance, posture and temporal feautures without considering a learned temporal attention mechanism for global and local information fusion. In this paper, we propose a novel gait recognition framework, called Temporal Attention and Keypoint-guided Embedding (GaitTAKE), which effectively fuses temporal-attention-based global and local appearance feature and temporal aggregated human pose feature. Experimental results show that our proposed method achieves a new SOTA in gait recognition with rank-1 accuracy of 98.0% (normal), 97.5% (bag) and 92.2% (coat) on the CASIA-B gait dataset; 90.4% accuracy on the OU-MVLP gait dataset.
Hung-Min Hsu, Yizhou Wang 0005, Cheng-Yen Yang, Jenq-Neng Hwang, Le Uyen Thuc Hoang, Kwang-Ju Kim
ICIP3
2022 Unsupervised Domain Adaptation Learning for Hierarchical Infant Pose Recognition with Synthetic Data
abstract
The Alberta Infant Motor Scale (AIMS) is a well-known assessment scheme that evaluates the gross motor development of infants by recording the number of specific poses achieved. With the aid of the image-based pose recognition model, the AIMS evaluation procedure can be shortened and automated, providing early diagnosis or indicator of potential developmental disorder. Due to limited public infant-related datasets, many works use the SMIL-based method to generate synthetic infant images for training. However, this domain mismatch between real and synthetic training samples often leads to performance degradation during inference. In this paper, we present a CNN-based model which takes any infant image as input and predicts the coarse and fine-level pose labels. The model consists of an image branch and a pose branch, which respectively generates the coarse-level logits facilitated by the unsupervised domain adaptation and the 3D keypoints using the HRNet with SMPLify optimization. Then the outputs of these branches will be sent into the hierarchical pose recognition module to estimate the fine-level pose labels. We also collect and label a new AIMS dataset, which co—tains 750 real and 4000 synthetic infants images with AIMS pose labels. Our experimental results show that the proposed method can significantly align the distribution of synthetic and real-world datasets, thus achieving accurate performance on fine-grained infant pose recognition.
Cheng-Yen Yang, Zhongyu Jiang, Shih-Yu Gu, Jenq-Neng Hwang, Jang-Hee Yoo
ICME1
2020 /TPlace: Machine Learning-Based Delay-Aware Transistor Placement for Standard Cell Synthesis
abstract
Cell layout synthesis is a critical stage in modern digital IC design. In previous automatic synthesis solutions, algorithms always consider only cell area and routability. This is the first work to propose a method of delay-aware transistor placement for cell library synthesis at the sign-off level. We consider the delay and area of a cell in the transistor placement stage. Our methodology consists of three major steps. First, a search tree finds the candidate placement list that has the smallest area in a large search space. Then, a neural network filters out the unroutable candidates. Finally, a comparative convolutional neural network model, trained by sign-off level data, sorts the delays during the early placement stage. The experimental results show that the proposed CNN-based routable classifier can achieve up to 98% accuracy, and the proposed CNN-based delay ranker also can achieve up to 94.6% accuracy. The work obtains a 1.77% average sequential component delay improvement over the traditional cell synthesis method. Our method also has a 0.97% better delay performance than the human-level design.
Tai-Cheng Lee, Cheng-Yen Yang, Yih-Lang Li
ICCAD2
2019 Weakly-Supervised Learning for Attention-Guided Skull Fracture Classification In Computed Tomography Imaging
abstract
We propose a novel attention-guided deep learning framework for image classification, with the goal of predicting both image and pixel-level labels for input images. While the training images are with either positive or negative labels, we do not assume that each training image is with annotated pixel-level ground truth information, and thus our method can be realized in such a weakly supervised setting. Our proposed module can be easily combined with standard CNN architectures with no extra parameter needed. With the above advantages, we evaluate our model on a skull CT dataset, and the experimental results confirm the effectiveness and robustness of our approach over recent popular CNN architectures.
Cheng-Yen Yang, Chi-Hsin Lo, Huan-Chih Wang, Jen-Hai Chou, Yu-Chiang Frank Wang
ICIP1
2016 A Systematic ANSI S1.11 Filter Bank Specification Relaxation and Its Efficient Multirate Architecture for Hearing-Aid Systems
abstract
Recently, emerging mobile computing requires the high integration of hearing aids into a single system-on-chip. Modern hearing aid systems include a frequency decomposer, noise reduction, feedback cancellation, auditory compensation, and intelligent adaptation. The majority of existing works concentrated on improving the performance and efficiency on a single-signal processing block or one-sided effect. These works lacked comprehensive discussions on system-wide aspects regarding the overall impacts. To design an optimal hearing aid system, frequency decomposers, or the filter banks that dominate in hearing aid systems, are the first priority. We propose a systematic relaxation of the ANSI S1.11 specification and its design procedure for filter banks. The proposed design procedure overcomes the drawbacks of previous works and changes the five performance indices of the filter bank: delay, complexity, sub-band rate reduction, ripples of synthesized output, and prescription matching errors. These performance indices help system or algorithm designers in selecting a beneficial system-relaxed filter bank to achieve optimal hearing aids. The proposed multirate filter bank using the resampling method provides an efficient, low complexity, and delay-constrained computing architecture. Finally, seven design cases are used to demonstrate the proposed method, and comprehensive discussions of the five performance indices are presented.
Cheng-Yen Yang, Chih-Wei Liu, Shyh-Jye Jou
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 An efficient 18-band quasi-ANSI 1/3-octave filter bank using re-sampling method for digital hearing aids
abstract
This paper presents the multirate and re-sampling techniques to realize a low-delay, 18-band quasi-ANSI filter bank for digital hearing aids, which not only achieves a rather low computation complexity without a significant increase in the latency, but reduces greatly the total computation complexity for sub-band signal processing followed by the filter bank, such as noise reduction as well as wide dynamic range compression (WDRC). Researches done in the literature all focused on how to reduce the computation complexity of the filter bank. In particular, with the efficient multirate and interpolated FIR (IFIR) approaches for a 10-ms, 18-band quasi-ANSI filter bank, approximately 93% of the multiplications are saved, compared that with a straightforward parallel FIR filters architecture. However, they did not consider the computation complexity of the sub-band signal processing. In this paper, we first investigate realizing the FIR filter bank efficiently by using the multirate re-sampling techniques. To reduce the complexity, the optimized re-sampling factor for each filter band is explored carefully. Then, with the resampling technique, an efficient multirate quasi-ANSI FIR filter bank architecture is proposed. Compare to the state-of-the-art quasi-ANSI filter bank, approximately 17.7% of multiplicative complexity is reduced further and, up to 25% of the total computation complexity for sub-band signal processing followed by the filter bank is saved, but with only a slight increase in latency, i.e. 13.6 ms.
Cheng-Yen Yang, Chih-Wei Liu, Shyh-Jye Jou
ICASSP1