Yuhan Shen

dblp:228/7784 · also Yu-Han Shen · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-3443-1701ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 6 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Understanding Multi-Task Activities from Single-Task Videos
abstract
We introduce and develop a framework for Multi-Task Temporal Action Segmentation (MT-TAS), a novel paradigm that addresses the challenges of interleaved actions when performing multiple tasks simultaneously. Traditional action segmentation models, trained on single-task videos, struggle to handle task switches and complex scenes inherent in multi-task scenarios. To overcome these challenges, our MT-TAS approach synthesizes multi-task video data from single-task sources using our Multi-Task Sequence Blending and Segment Boundary Learning modules. Additionally, we propose to dynamically isolate foreground and background elements within video frames, addressing the intricacies of object layouts in multi-task scenarios and enabling a new two-stage temporal action segmentation framework with Foreground-Aware Action Refinement. Also, we introduce the Multi-Task Egocentric Kitchen Activities (MEKA) dataset, containing 12 hours of egocentric multi-task videos, to rigorously benchmark MT-TAS models. Extensive experiments demonstrate that our framework effectively bridges the gap between single-task training and multi-task testing, advancing temporal action segmentation with state-of-the-art performance in complex environments.1
Yuhan Shen, Ehsan Elhamifar
CVPR1
2025 MOSCATO: Predicting Multiple Object State Change through Actions
Parnian Zameni, Yuhan Shen, Ehsan Elhamifar
ICCV2
2025 Augmented Reality-Based Interactive Scheme for Robot-Assisted Percutaneous Renal Puncture Navigation
abstract
ABSTRACT In this paper, we present an Augmented Reality (AR)‐based application combined with a robotic system for percutaneous renal puncture navigation interaction and demonstrate its technical feasibility. Our system provides an intuitive interaction scheme between the surgeon and the robot without the need for traditional external input devices, and applies an image‐target‐based 3D registration scheme to transform the coordinate system between Hololens2 and the robot without using additional tracking devices. Users can visualize the abdominal puncture phantom and obtain 3D depth information of the lesion site by wearing Hololens2 and control the robot directly using buttons or gestures. To investigate the accuracy and feasibility of the proposed interaction scheme, six subjects were recruited to complete 3D registration alignment accuracy experiments, and puncture positioning accuracy experiments using ultrasound unaided navigation, AR unaided navigation and AR robotic navigation. The results showed that the average alignment error of 3D registration was 3.61 ± 1.05 mm. The average positioning errors of ultrasound freehand navigation, AR freehand navigation and AR robotic navigation were 7.67 ± 2.00 mm, 6.13 ± 1.07 mm and 5.52 ± 0.37 mm, respectively; the average puncture times were 34.86 ± 1.67 s, 22.40 ± 2.07 s, and 29.41 ± 1.37 s.
Yiwei Zhuang, Hua Xie, Wei Qing, Haoliang Li, Yuhan Shen, Yichun Shen
Comput. Animat. Virtual Worlds6
2024 Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videos
abstract
We address the problem of online (streaming) action seg-mentation for egocentric procedural task videos. While pre-vious studies have mostly focused on offline action segmen-tation, where entire videos are available for both training and inference, the transition to online action segmentation is crucial for practical applications like AR/VR task assistants. Notably, applying an offline-trained model directly to online inference results in a significant performance drop due to the inconsistency between training and inference. We propose an online action segmentation framework by first modifying existing architectures to make them causal. Sec-ond, we develop a novel action progress prediction module to dynamically estimate the progress of ongoing actions and using them to refine the predictions of causal action segmen-tation. Third, we propose to learn task graphs from training videos and leverage them to obtain smooth and procedure-consistent segmentations. With the combination of progress and task graph with casual action segmentation, our frame-work effectively addresses prediction uncertainty and over-segmentation in online action segmentation and achieves significant improvement on three egocentric datasets.11Code is available at https://github.com/Yuhan-Shen/ProTAS.
Yuhan Shen, Ehsan Elhamifar
CVPR1
2024 Learning to Segment Referred Objects from Narrated Egocentric Videos
abstract
Egocentric videos provide a first-person perspective of the wearer's activities, involving simultaneous interactions with multiple objects. In this work, we propose the task of weakly-supervised Narration-based Video Object Segmentation (NVOS). Given an egocentric video clip and a narration of the wearer's activities, our aim is to segment object instances mentioned in the narration, with-out using any spatial annotations during training. Existing weakly-supervised video object grounding methods typ-ically yield bounding boxes for referred objects. In contrast, we propose ROSA, a weakly-supervised pixel-level grounding framework learning alignments between referred objects and segmentation mask proposals. Our model harnesses vision-language models pre-trained on image-text pairs to embed region masks and object phrases. During training, we combine (a) a video-narration contrastive loss that implicitly supervises the alignment between regions and phrases, and (b) a region-phrase contrastive loss based on inferred latent alignments. To address the lack of annotated NVOS datasets in egocentric videos, we create a new evaluation benchmark, VISOR-NVOS, leveraging existing annotations of segmentation masks from VISOR alongside 14.6k newly-collected, object-based video clip narrations. Our approach achieves state-of-the-art zero-shot pixel-level grounding performance compared to strong baselines under similar supervision. Additionally, we demonstrate generalization capabilities for zero-shot video object grounding on YouCook2, a third-person instructional video dataset.
Yuhan Shen, Xitong Yang, Matt Feiszli, Ehsan Elhamifar, Lorenzo Torresani, Effrosyni Mavroudi
CVPR1
2022 Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task Videos
abstract
We address the problem of action segmentation in instructional task videos with a small number of weakly-labeled training videos and a large number of unlabeled videos, which we refer to as Semi-Weakly-Supervised Learning (SWSL) of actions. We propose a general SWSL framework that can efficiently learn from both types of videos and can leverage any of the existing weakly-supervised action segmentation methods. Our key observation is that the distance between the transcript of an unlabeled video and those of the weakly-labeled videos from the same task is small yet often nonzero. Therefore, we develop a Soft Restricted Edit (SRE) loss to encourage small variations between the predicted transcripts of unlabeled videos and ground-truth transcripts of the weakly-labeled videos of the same task. To compute the SRE loss, we develop a flexible transcript prediction (FTP) method that uses the output of the action classifier to find both the length of the transcript and the sequence of actions occurring in an unlabeled video. We propose an efficient learning scheme in which we alternate between minimizing our proposed loss and generating pseudo-transcripts for unlabeled videos. By experiments on two benchmark datasets, we demonstrate that our approach can significantly improve the performance by using unlabeled videos, especially when the number of weakly-labeled videos is small.11Code available at https://github.com/Yuhan-Shen/SWSL..
Yuhan Shen, Ehsan Elhamifar
CVPR1
2021 Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence Alignment
abstract
We address the problem of unsupervised localization of task-relevant actions (key-steps) and feature learning in instructional videos using both visual and language instructions. Our key observation is that the sequences of visual and linguistic key-steps are weakly aligned: there is an ordered one-to-one correspondence between most visual and language key-steps, while some key-steps in one modality are absent in the other. To recover the two sequences, we develop an ordered prototype learning module, which extracts visual and linguistic prototypes representing key-steps. To find weak alignment and perform feature learning, we develop a differentiable weak sequence alignment (DWSA) method that finds ordered one-to-one matching between sequences while allowing some items in a sequence to stay unmatched. We develop an efficient forward and backward algorithm for computing the alignment and the loss derivative with respect to parameters of visual and language feature learning modules. By experiments on two instructional video datasets, we show that our method significantly improves the state of the art.
Yuhan Shen, Lu Wang 0008, Ehsan Elhamifar
CVPR1
2020 Staged Training Strategy and Multi-Activation for Audio Tagging with Noisy and Sparse Multi-Label Data
Kexin He, Yuhan Shen, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP2
2020 Energy-Efficient Activation and Uplink Transmission for Cellular IoT
abstract
Consider a large-scale cellular network in which base stations (BSs) serve massive Internet of Things (IoT) devices. Since IoT devices are powered by a capacity-limited battery, how to prolong their working lifetime is a paramount problem for the success of cellular IoT systems. This article proposes how to use BSs to manage the active and dormant operating modes of the IoT devices via downlink signaling in an energy-efficient fashion and how the IoT devices perform energy-efficient uplink power control to improve their uplink coverage. We first investigate the fundamental statistical properties of an activation signaling process induced by BSs that would like to activate the devices in their cells, which helps to derive the neat expressions of the true, false, and total activation probabilities that reveal joint downlink power control and BS coordination is an effective means to significantly improve the activation performance. We then propose an energy-efficient uplink power control for IoT devices which is shown to save power and ameliorate the uplink coverage probability at the same time. We also propose an energy-efficient downlink power control and BS coordination scheme, which is shown to remarkably improve the activation and uplink coverage performances at the same time.
Chun-Hung Liu, Yuhan Shen, Chia-han Lee
IEEE Internet Things J.2
2019 Hierarchical Pooling Structure for Weakly Labeled Sound Event Detection
abstract
Sound event detection with weakly labeled data is considered as a problem of multi-instance learning.And the choice of pooling function is the key to solving this problem.In this paper, we proposed a hierarchical pooling structure to improve the performance of weakly labeled sound event detection system.Proposed pooling structure has made remarkable improvements on three types of pooling function without adding any parameters.Moreover, our system has achieved competitive performance on Task 4 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge using hierarchical pooling structure.
Kexin He, Yuhan Shen, Weiqiang Zhang 0001
INTERSPEECH2
2019 Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
abstract
In this paper, we propose a temporal-frequential attention model for sound event detection (SED). Our network learns how to listen with two attention models: a temporal attention model and a frequential attention model. Proposed system learns when to listen using the temporal attention model while it learns where to listen on the frequency axis using the frequential attention model. With these two models, we attempt to make our system pay more attention to important frames or segments and important frequency components for sound event detection. Our proposed method is demonstrated on the task 2 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge and achieves competitive performance.
Yuhan Shen, Kexin He, Weiqiang Zhang 0001
INTERSPEECH1
2019 Recover Glacier Velocity Fields Derived From the SAR Speckle Tracking Technique Using Artificial Neural Network
abstract
The speckle tracking technique has demonstrated its great potential in glacier velocity field (GVF) mapping applications. It analyzes the cross correlation between two synthetic aperture radar (SAR) images, which is capable of providing 2-D glacier motion measurements with acceptable accuracy. Since the coherence between a SAR image pair illuminating glacier areas cannot be preserved everywhere in nearly all cases, inevitable no data areas will be presented in SAR speckle tracking products. In this letter, a relatively convenient method is proposed to recover a GVF generated by the speckle tracking technique. This method considers a GVF recovery problem to be a unitary supervised training problem and resolve it via an artificial neural network. A targeted architecture of the network for recovering a GVF is presented. Moreover, the parameters correlative with glacier motion mechanism are proposed to be introduced into the network, which is able to effectively improve the performance of GVF recovery. For the purpose of validation, the proposed method is compared with the well-known Kriging interpolation method based on real speckle tracking products. The experimental results demonstrate that the proposed method can effectively recover a GVF derived from the speckle tracking technique.
Faming Gong, Shujun Liu, Yuhan Shen
IEEE Geosci. Remote. Sens. Lett.5