Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Keyu Pan

dblp:294/1953 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 42% Multi-agent systems · 26% Robot manipulation · 13%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › temporal grounding
temporal language grounding
1.422024
Temporally Language Grounding With Multi-Modal Multi-Prompt Tuning · IEEE Trans. Multim. 2024
RewardTLG: Learning to Temporally Language Grounding from Flexible Reward · SIGIR 2023
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
1.012026
How Foundational Skills Influence VLM-based Embodied Agents: A Native Perspective · AAAI 2026
Computer vision › Vision and language
vision-language model
1.012026
How Foundational Skills Influence VLM-based Embodied Agents: A Native Perspective · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems › multimodal agent
vision-language model agent
1.012026
How Foundational Skills Influence VLM-based Embodied Agents: A Native Perspective · AAAI 2026
Computer vision › Vision and language
multimodal prompt learning
0.812024
Temporally Language Grounding With Multi-Modal Multi-Prompt Tuning · IEEE Trans. Multim. 2024
Computer vision › Video understanding and tracking › temporal localization
video moment localization
0.812024
Temporally Language Grounding With Multi-Modal Multi-Prompt Tuning · IEEE Trans. Multim. 2024
Natural language and speech › Language models and text generation
prompt tuning
0.212024
Temporally Language Grounding With Multi-Modal Multi-Prompt Tuning · IEEE Trans. Multim. 2024
Machine learning › Reinforcement learning
reward design
0.212023
RewardTLG: Learning to Temporally Language Grounding from Flexible Reward · SIGIR 2023

Methods — techniques the papers use, named apart from their topics

language model · 1.5skill decoupling · 1.0benchmark construction · 1.0prompt tuning · 0.8multimodal transformer · 0.8multi-task learning · 0.8reward shaping · 0.7reinforcement learning · 0.7
YearPublicationVenuePosition
2026 How Foundational Skills Influence VLM-based Embodied Agents: A Native Perspective
abstract
Recent advances in vision–language models (VLMs) have shed light on human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents still rely on high-level commands or discretised action spaces—``non-native'' settings that diverge markedly from the real world. Moreover, current benchmarks focus exclusively on high-level tasks, while lacking joint evaluation and analysis on both low- and high-level. To bridge these gaps, we present \textbf{NativeEmbodied}, a challenging benchmark for VLM-driven embodied agents that adopts a unified, native low-level action space. Built upon diverse simulated scenes, NativeEmbodied first designs three representative high-level tasks in complex scenarios to evaluate overall performance. For more detailed and comprehensive performance analysis, we further decouple the entangled skills behind complex tasks and construct four types of low-level tasks, each corresponding to a key fundamental embodied skill. This joint evaluation across task and skill granularities enables a fine-grained assessment of embodied agent. Comprehensive experiments on the best VLMs reveal pronounced deficiencies in certain fundamental embodied skills. Further analysis shows that these bottlenecks severely constrain performance on high-level tasks. Our NativeEmbodied not only pinpoints the key challenges faced by current VLM-driven embodied agents, but also provides valuable insight for future development of this field.
Pi Bu, Keyu Pan, Xinrun Xu, Yingxiu Zhao, Tong Xu 0001
AAAI3
2026 mmWave Radar-Based Continuous Sign Language Recognition: Lightweight Modeling, Contextual Optimization, and Embedded Implementation
abstract
Communication barriers for the hearing-impaired represent a significant societal concern. Existing sign language recognition solutions are constrained by lighting sensitivity, wearable device burdens, and limited vocabulary coverage. This work presents the first FMCW mmWave radar-based continuous sign language recognition framework, with three key innovations: (1) the Radar Continuous Chinese Sign Language-108 (RCCSL-108) dataset with 108 isolated signs and 66 natural sentences to address the critical absence of sentence-level radar data; (2) a novel lightweight lightweight Continuous Sign Language Recognition Temporal Convolutional Network (CSLR-TCN) that incorporates dual-dilated convolutions for variable-length inputs, achieving 92.66% frame accuracy and robust cross-user generalization; and (3) a novel contextual reordering module that enhances semantic understanding by 6.33% on unseen data by mitigating emotion loss in translation. The implemented embedded system achieves 87.49% recognition accuracy in untrained user experiments with ≥15 FPS throughput, marking a pivotal advancement toward practical radar-based sign language interfaces.
Zhiyan Lin, Minming Gu, Kaiyu Chen, Keyu Pan
IEEE Internet Things J.4
2026 A multi-stage few-shot framework for extensible radar-based human activity recognition
Keyu Pan, Wei-Ping Zhu 0001, Bo Shi 0001
Signal Process.1
2025 Adaptive-Prompt-Driven Few-Shot Class-Incremental Learning for Human Activity Recognition With Radar Modality
abstract
Human activity recognition (HAR) has attracted growing interest due to its wide ranging applications in healthcare, security, and smart surveillance. In particular, radar-based HAR offers a robust, non-contact alternative to conventional sensor based methods. Despite its promise, existing approaches often rely on substantial labeled data and exhibit limited adaptability in small-shot incremental learning scenarios. To address these challenges, we propose a novel few shot class incremental learning (FSCIL) framework for radar-based HAR. Our framework exploits vision transformers (ViTs) augmented with adaptive prompt mechanisms, hybrid attention blocks (HAB), and dynamic feature refinement strategies, which can substantially improve incremental learning performance. We validate the proposed approach on a fused radar dataset containing four sets of diverse radar datasets collected under varying environmental conditions. Our validation shows a harmonic accuracy (HAcc) up to 69.54 % and a Top-1 average accuracy (Top-1 Avg.) up to 91.20 %, outperforming state-of-the-art methods, with a session performance degradation as low as 10.80%. These results underscore the robustness of our approach in recognizing novel activities while preserving previously acquired knowledge.
Keyu Pan, Wei-Ping Zhu 0001
IEEE Internet Things J.1
2024 Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning
Yi Zhai 0005, Keyu Pan, Renwei Zhang, Shuo Liu 0019, Zichun Ye, Jianmin Ji, Jie Zhao 0002, Yu Zhang 0086, Yanyong Zhang
OSDI3
2024 Temporally Language Grounding With Multi-Modal Multi-Prompt Tuning
abstract
The task of temporally language grounding (TLG), aiming to locate a video moment within an untrimmed video that matches a given textual query, has attracted considerable research attention in recent years. Typical retrieval-based TLG methods are inefficient due to their reliance on a large number of pre-segmented candidate moments, while localization-based TLG solutions adopt reinforcement learning, resulting in unstable convergence. Meanwhile, the cutting-edge capabilities of multi-modal architecture, especially pre-training paradigm, have not been fully exploited. Therefore, how to perform TLG task efficiently and stably is a non-trivial task. In this work, we propose a novel TLG solution named Multi-modal Multi-Prompt Tuning (MMPT), which formulates the TLG task as a prompt-based multi-modal problem and integrates multiple sub-tasks to tune the performance. In this way, off-the-shelf pre-trained models can be directly leveraged to achieve more stable performance. Specifically, a flexible multi-prompt strategy is contributed to rewrite the query firstly, which contains the query, the start and end timestamps. Among them, various prompt templates are integrated to enhance robustness. Thereafter, a multi-modal Transformer is adopted to fully learn the multi-modal context. Moreover, we design various sub-tasks to optimize this novel framework including the matching task, localization task and joint learning task. Extensive experiments on two real-world datasets validate the effectiveness and rationality of our proposed solution.
Yawen Zeng, Ning Han 0005, Keyu Pan, Qin Jin
IEEE Trans. Multim.3
2023 RewardTLG: Learning to Temporally Language Grounding from Flexible Reward
abstract
Given a textual sentence provided by a user, the Temporal Language Grounding (TLG) task is defined as the process of finding a semantically relevant video moment or clip from an untrimmed video. In recent years, localization-based TLG methods have been explored, which adopt reinforcement learning to locate a clip from the video. However, these methods are not stable enough due to the stochastic exploration mechanism of reinforcement learning, which is sensitive to the reward. Therefore, providing a more flexible and reasonable reward has become a focus of attention for both academia and industry.
Yawen Zeng, Keyu Pan, Ning Han 0005
SIGIR2