Qianshi Pang

dblp:371/3952 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0009-7702-7897ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
3D vision · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 50% Visualization and visual analytics · 50%
Human-computer interaction and pervasive computing
1 paper
User interface design and tools · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.012026
OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding · AAAI 2026
Computer vision › 3D vision › 3d scene understanding
open-vocabulary 3d scene understanding
1.012026
OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding · AAAI 2026
Multimedia analysis and retrieval
video summarization
0.912025
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations · ACM Multimedia 2025
Visualization and visual analytics
visual augmentation
0.912025
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations · ACM Multimedia 2025
User interface design and tools
interactive systems
0.912025
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations · ACM Multimedia 2025
User interface design and tools › multimedia interface design
video browsing interface
0.912025
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

speech content analysis · 1.7automatic visual augmentation generation · 1.7benchmark construction · 1.0
YearPublicationVenuePosition
2026 OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding
abstract
Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the context of object classes, which is insufficient in providing a holistic evaluation to what extent a model understands the 3D scene. In this paper, we introduce a more challenging task called Generalized Open-Vocabulary 3D Scene Understanding (GOV-3D) to explore the open vocabulary problem beyond object classes. It encompasses an open and diverse set of generalized knowledge, expressed as linguistic queries of fine-grained and object-specific attributes. To this end, we contribute a new benchmark named OpenScan, which consists of 3D object attributes across eight representative linguistic aspects, including affordance, property, and material. We further evaluate state-of-the-art OV-3D methods on our OpenScan benchmark and discover that these methods struggle to comprehend the abstract vocabularies of the GOV-3D task, a challenge that cannot be addressed simply by scaling up object classes during training. We highlight the limitations of existing methodologies and explore promising directions to overcome the identified shortcomings.
Youjun Zhao, Jiaying Lin 0001, Shuquan Ye, Qianshi Pang, Rynson W. H. Lau
AAAI4
2025 VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
abstract
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 0001, Fan Zhou 0001, Shujin Lin
ACM Multimedia3