Shih-Chieh Huang

dblp:85/6552 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Face, body and person analysis · 50% Kernel, tree and ensemble methods · 25% Image recognition and object detection · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
coarse-to-fine detection
0.212015
Regressive Tree Structured Model for Facial Landmark Localization · ICCV 2015
Computer vision › Face, body and person analysis
face alignment
0.212015
Regressive Tree Structured Model for Facial Landmark Localization · ICCV 2015
Computer vision › Face, body and person analysis
face detection
0.212015
Regressive Tree Structured Model for Facial Landmark Localization · ICCV 2015
Machine learning › Kernel, tree and ensemble methods
tree-based models
0.212015
Regressive Tree Structured Model for Facial Landmark Localization · ICCV 2015

Methods — techniques the papers use, named apart from their topics

tree-structured model · 0.2support vector regression · 0.2
YearPublicationVenuePosition
2025 SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
abstract
Voice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses are only loosely coupled with the evaluation metric of VAD. To address these challenges, we propose SincQDR-VAD, a compact and robust framework that combines a Sinc-extractor front-end with a novel quadratic disparity ranking loss. The Sinc-extractor uses learnable bandpass filters to capture noise-resistant spectral features, while the ranking loss optimizes the pairwise score order between speech and non-speech frames to improve the area under the receiver operating characteristic curve (AUROC). A series of experiments conducted on representative benchmark datasets show that our framework considerably improves both AUROC and $F_{2}$-Score, while using only $69 \%$ of the parameters compared to prior arts, confirming its efficiency and practical viability.
Chien-Chun Wang, En-Lun Yu, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen
ASRU4
2025 Flexible VAD-PVAD Transition: A Detachable PVAD Module for Dynamic Encoder RNN VAD
En-Lun Yu, Chien-Chun Wang, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen
INTERSPEECH4
2024 Speaker Conditional Sinc-Extractor for Personal VAD
En-Lun Yu, Kuan-Hsun Ho, Jeih-Weih Hung, Shih-Chieh Huang, Berlin Chen
INTERSPEECH4
2015 Regressive Tree Structured Model for Facial Landmark Localization
abstract
Although the Tree Structured Model (TSM) is proven effective for solving face detection, pose estimation and landmark localization in an unified model, its sluggish run time makes it unfavorable in practical applications, especially when dealing with cases of multiple faces. We propose the Regressive Tree Structure Model (RTSM) to improve the run-time speed and localization accuracy. The RTSM is composed of two component TSMs, the coarse TSM (c-TSM) and the refined TSM (r-TSM), and a Bilateral Support Vector Regressor (BSVR). The c-TSM is built on the low-resolution octaves of samples so that it provides coarse but fast face detection. The r-TSM is built on the mid-resolution octaves so that it can locate the landmarks on the face candidates given by the c-TSM and improve precision. The r-TSM based landmarks are used in the forward BSVR as references to locate the dense set of landmarks, which are then used in the backward BSVR to relocate the landmarks with large localization errors. The forward and backward regression goes on iteratively until convergence. The performance of the RTSM is validated on three benchmark databases, the Multi-PIE, LFPW and AFW, and compared with the latest TSM to demonstrate its efficacy.
Gee-Sern Hsu, Kai-Hsiang Chang, Shih-Chieh Huang
ICCV3
2015 Face detection and landmark localization using Bilayer Tree Structured Model
abstract
Although the Tree Structured Model (TSM) is proven effective for face detection, pose estimation and landmark localization, its sluggish runtime makes it unfavorable in practical applications. We propose the Bilayer Tree Structure Model (BTSM) to improve the run-time speed while keeping the performance unchanged or slightly better. The BTSM is composed of two component TSMs, the coarse c-TSM and the refined r-TSM. The c-TSM is trained on low-resolution samples so that it can provide coarse but fast detection, The r-TSM is trained on mid-resolution samples so that it can locate precise part locations. The performance of the BTSM is validated on three benchmark databases, the Multi-PIE, LFPW and AFW, and compared with the latest TSM to demonstrate its efficacy.
Gee-Sern Hsu, Kai-Hsiang Chang, Shih-Chieh Huang, Sheng-Luen Chung
ICIP3
2014 Ensuring the integrity and non-repudiation of remitting e-invoices in conventional channels with commercially available NFC devices
abstract
Despite the globally recognized advantages of e-invoicing and various efforts to implement such systems, retailers and stores may still have difficulties in promoting purely paperless e-invoices due to the lack of a convenient and secure way for consumers to receive and retrieve the e-invoices. As such, paper-based invoices may still be issued along with e-invoices, contradicting an important benefit of e-invoicing — paper consumption reduction. Thanks to the advances in smart phones and Near Field Communication (NFC) technologies, e-invoices can be delivered via NFC-enabled smartpones, allowing consumers to examine the content immediately after transactions and to easily retrieve them later on. Still, an extra security mechanism is needed to ensure the integrity and non-repudiation of the content, as invoices may bear some value and thus become the target of a security attack. In this paper, we propose a secure NFC-based e-invoice remitting scheme using standard NFC P2P communications, and discuss how it fulfills major security requirements, including authenticity, integrity, and non-repudiation. The proposed system is also implemented and tested in Taiwan's e-invoicing system.
Shi-Cho Cha 0001, Yuh-Jzer Joung, Yen-Chung Tseng, Shih-Chieh Huang, Guan-Heng Chen, Chih-Teng Tseng
SNPD4
2011 The Effects of LEGO Robotics and Embodiment in Elementary Science Learning
Carol M. Lu, John B. Black, Seokmin Kang, Shih-Chieh Huang
CogSci4
2007 A Dynamic E-learning System for the Collaborative Business Environment
abstract
Employees should continue their learning activities to face the changing and more competitive business environment. Nowadays, information technology plays an important role of building an e-learning system. In particular, employees work in a collaborative business environment can take advantage of IT-based e- learning system to enhance their learning and share knowledge with others among the business alliance. In this environment, traditional on-job training programs require specific space and time for the tutors and learners that cause extra expenses and reduce learning motivation. Therefore, building a dynamic e-learning system through IT becomes an essential part of the employee training programs. It should provide timely knowledge for employees to complete their assigned tasks individually or collaboratively with partners. A structure with five modules of dynamically configuring appropriate learning objects for learners is proposed. An interactive mode of retrieving learning materials based on individual requirement is presented. Personalized information is then issued to employees and partners to improve problem-solving capability.
Hsien-Jung Wu, Shih-Chieh Huang
ICALT2