Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shih-Po Lee

dblp:288/0210 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 60% Information extraction and text analysis · 22% Graph learning · 11%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
error detection
0.912025
Error Recognition in Procedural Videos Using Generalized Task Graph · ICCV 2025
Computer vision › Video understanding and tracking › activity recognition › procedural activity understanding
procedural video understanding
0.912025
Error Recognition in Procedural Videos Using Generalized Task Graph · ICCV 2025
Computer vision › Video understanding and tracking
action segmentation
0.812024
Error Detection in Egocentric Procedural Task Videos · CVPR 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan representation
task graph
0.312025
Error Recognition in Procedural Videos Using Generalized Task Graph · ICCV 2025
Machine learning › Graph learning › graph neural network
graph convolutional network
0.212024
Error Detection in Egocentric Procedural Task Videos · CVPR 2024
Machine learning › Graph learning › graph representation › scene graph representation
object relation graph
0.212024
Error Detection in Egocentric Procedural Task Videos · CVPR 2024

Methods — techniques the papers use, named apart from their topics

generalized task graph · 0.9step prototype learning · 0.8graph convolutional network · 0.8contrastive learning · 0.8active object detection · 0.8
YearPublicationVenuePosition
2025 Error Recognition in Procedural Videos Using Generalized Task Graph
Shih-Po Lee, Ehsan Elhamifar
ICCV1
2024 Error Detection in Egocentric Procedural Task Videos
abstract
We present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step prototype learning module to segment actions and learn useful features for error detection. Based on the observation that interactions between hands and objects often inform action and error understanding, we propose to combine holistic frame features with relations features, which we learn by building a graph using active object detection followed by a Graph Convolutional Network. To handle errors, unseen during training, we use our contrastive step prototype learning to learn multiple prototypes for each step, capturing variations of error-free step executions. At inference time, we use feature-prototype similarities for error detection. By experiments on three datasets, we show that our proposed framework outperforms state-of-the-art video anomaly detection methods for error detection and provides smooth action and error predictions.11Code and data is available at https://github.com/robert80203/EgoPER_official
Shih-Po Lee, Zijia Lu, Minh Hoai, Ehsan Elhamifar
CVPR1
2023 HuPR: A Benchmark for Human Pose Estimation Using Millimeter Wave Radar
abstract
This paper introduces a novel human pose estimation benchmark, Human Pose with Millimeter Wave Radar (HuPR), that includes synchronized vision and radio signal components. This dataset is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation. There are two advantages of using mmWave radar to perform human pose estimation. First, it is robust to dark and low-light conditions. Second, it is not visually perceivable by humans and thus, can be widely applied to applications with privacy concerns, e.g., surveillance systems in patient rooms. In addition to the benchmark, we propose a cross-modality training framework that leverages the ground-truth 2D keypoints representing human body joints for training, which are systematically generated from the pre-trained 2D pose estimation network based on a monocular camera input image, avoiding laborious manual label annotation efforts. The framework consists of a new radar pre-processing method that better extracts the velocity information from radar data, Cross- and Self-Attention Module (CSAM), to fuse multi-scale radar features, and Pose Refinement Graph Convolutional Networks (PRGCN), to refine the predicted keypoint confidence heatmaps. Our intensive experiments on the HuPR benchmark show that the proposed scheme achieves better human pose estimation performance with only radar data, as compared to traditional pre-processing solutions and previous radiofrequency-based methods. Our code is available at here1
Shih-Po Lee, Niraj Prakash Kini, Wen-Hsiao Peng, Ching-Wen Ma, Jenq-Neng Hwang
WACV1
2021 GSVNET: Guided Spatially-Varying Convolution for Fast Semantic Segmentation on Video
abstract
This paper addresses fast semantic segmentation on video. Video segmentation often calls for real-time, or even faster than real-time, processing. One common recipe for conserving computation arising from feature extraction is to propagate features of few selected keyframes. However, recent advances in fast image segmentation make these solutions less attractive. To leverage fast image segmentation for furthering video segmentation, we propose a simple yet efficient propagation framework. Specifically, we perform lightweight flow estimation in 1/8-downscaled image space for temporal warping in segmentation outpace space. Moreover, we introduce a guided spatially-varying convolution for fusing segmentations derived from the previous and current frames, to mitigate propagation error and enable lightweight feature extraction on non-keyframes. Experimental results on Cityscapes and CamVid show that our scheme achieves the state-of-the-art accuracy-throughput trade-off on video segmentation.
Shih-Po Lee, Si-Cun Chen, Wen-Hsiao Peng
ICME1
2021 Weakly-Supervised Image Semantic Segmentation Using Graph Convolutional Networks
abstract
This work addresses weakly-supervised image semantic segmentation based on image-level class labels. One common approach to this task is to propagate the activation scores of Class Activation Maps (CAMs) using a random-walk mechanism in order to arrive at complete pseudo labels for training a semantic segmentation network in a fully-supervised manner. However, the feed-forward nature of the random walk imposes no regularization on the quality of the resulting complete pseudo labels. To overcome this issue, we propose a Graph Convolutional Network (GCN)-based feature propagation framework. We formulate the generation of complete pseudo labels as a semi-supervised learning task and learn a 2-layer GCN separately for every training image by back-propagating a Laplacian and an entropy regularization loss. Experimental results on the PASCAL VOC 2012 dataset confirm the superiority of our scheme to several state-of-the-art baselines. Our code is available at https: //github.com/Xavier-Pan/WSGCN.
Shun-Yi Pan, Cheng-You Lu, Shih-Po Lee, Wen-Hsiao Peng
ICME3