Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shitong Sun

dblp:304/7923 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1825-655XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 33% Video understanding and tracking · 22% Question answering and dialogue systems · 17%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 44% Accessibility and assistive technology · 44% Interaction techniques and input · 13%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › multimodal question answering
egocentric question answering
1.012026
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI · WWW 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI · WWW 2026
Computer vision › Vision and language
multimodal grounding
1.012026
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation · AAAI 2026
Computer vision › Vision and language › visual grounding
referential grounding
1.012026
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation · AAAI 2026
Computer vision › Video understanding and tracking
action recognition
0.712023
FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation · ICCV 2023
Machine learning › Efficient and distributed learning
federated learning
0.712023
FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation · ICCV 2023
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.712023
FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation · ICCV 2023
Interaction techniques and input › spatial interaction
egocentric interaction
0.312026
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation · AAAI 2026
Privacy and data protection
privacy-preserving machine learning
0.212023
FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation · ICCV 2023

Methods — techniques the papers use, named apart from their topics

multimodal intent recognition · 3.0hierarchical context compression · 3.0WebRTC · 3.0zero-shot prompting · 2.0vision-language model · 2.0temporal chain-of-thought · 2.0dialogue-driven reasoning · 2.0knowledge distillation · 1.3graph neural network · 1.3adaptive topology structure · 1.3temporal chain of thought · 1.0
YearPublicationVenuePosition
2026 Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
abstract
The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4-8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer accuracy for referential grounding by 5%. Overall, our method provides a plug-and-play framework that effectively resolves multimodal ambiguity and significantly enhances user experience in egocentric interaction.
Weitong Cai, Shitong Sun, You He 0003, Jiankang Deng, Hang Zhang 0010, Jifei Song, Zhensong Zhang
AAAI4
2026 Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
abstract
What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overload, smart glasses coupled with AI agents could turn the web into an always-on assistive layer over daily life. We present Egocentric Co-Pilot, a web-native neuro-symbolic framework that runs on smart glasses and uses a Large Language Model (LLM) to orchestrate a toolbox of perception, reasoning, and web tools. An egocentric reasoning core combines Temporal Chain-of-Thought with Hierarchical Context Compression to support long-horizon question answering and decision support over continuous first-person video, far beyond a single model's context window. Additionally, a lightweight multimodal intent layer maps noisy speech and gaze into structured commands. We further implement and evaluate a cloud-native WebRTC pipeline integrating streaming speech, video, and control messages into a unified channel for smart glasses and browsers. In parallel, we deploy an on-premise WebSocket baseline, exposing concrete trade-offs between local inference and cloud offloading in terms of latency, mobility, and resource use. Experiments on Egolife and HD-EPIC demonstrate competitive or state-of-the-art egocentric QA performance, and a human-in-the-loop study on smart glasses shows higher task completion and user satisfaction than leading commercial baselines. Taken together, these results indicate that web-connected egocentric co-pilots can be a practical path toward more accessible, context-aware assistance in everyday life. By grounding operation in web-native communication primitives and modular, auditable tool use, Egocentric Co-Pilot offers a concrete blueprint for assistive, always-on web agents that support education, accessibility, and social inclusion for people who may benefit most from contextual, egocentric AI.
Weitong Cai, Shitong Sun, Fengyi Fang, You He 0003, Yiqiao Xie, Jiankang Deng, Hang Zhang 0010, Jifei Song, Zhensong Zhang
WWW4
2026 Leverage cross-domain variations for generalizable person ReID representation learning
Qilei Li, Shitong Sun, Weitong Cai, Shaogang Gong
Pattern Recognit.2
2026 Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts
Qilei Li, Shitong Sun, Da Li 0001, Timothy M. Hospedales, Shaogang Gong
Pattern Recognit.2
2026 GridCLIP: One-stage object detection by grid-level CLIP representation learning
abstract
• We exploit CLIP to supplement the missing knowledge of undersampled and unseen categories in training a one-stage detector, mitigating the poor performance due to the long-tail data distribution in most existing detection training data. • We propose a simple yet effective visual-to-visual knowledge distillation method for learning undersampled and unseen categories for constructing a one-stage CLIP-based detector, providing 2.4 AP gains on unseen categories compared to the baseline. • GridCLIP is capable of handling Open-Vocabulary Object Detection with considerable scalability and generalizability, reaching comparable performance to two-stage detectors with much higher training and inference speed, without using extra pretraining processes or additional fine-tuning datasets. CLIP provides a shared image-text representation space with rich and diverse vocabulary, enabling object detection in undersampled and unseen categories. Recent CLIP-based object detection works show two-stage detectors typically outperform one-stage designs, but with significantly higher computational costs. A fundamental limitation of a two-stage detector is region-level alignment (distillation), which requires hundreds of image encoder forward passes from both the detector and CLIP in each image. In this work, we propose GridCLIP, a one-stage detector that requires only a single image encoder inference per input image, achieving up to 43 × faster training and 5 × faster inference compared to its two-stage counterpart ViLD, while substantially narrowing the accuracy gap. GridCLIP introduces a dual alignment strategy to learn fine-grained, grid-level representations: (1) grid-level alignment: learning grid-level features aligned with CLIP text encoder using annotated category labels, and (2) image-level alignment: aggregating grid-level features into an image-level representation aligned with the CLIP image encoder, which allows GridCLIP to learn grid-level representations of a broad range of categories, especially undersampled and unseen categories. Experiments on the LVIS benchmark show that GridCLIP achieves competitive results, with strong generalization to COCO and VOC, demonstrating its efficiency and effectiveness as a CLIP-based detector.
Jiayi Lin 0002, Shitong Sun, Shaogang Gong
Pattern Recognit.2
2026 BenchCIR: Benchmarking robustness in composed image retrieval across modalities
abstract
Composed image retrieval aims to retrieve images based on a query that consists of a reference image and text describing desired modifications to that image. It has recently attracted attention for its ability to tailor image retrieval to user intentions by combining information-rich reference images with concise natural language instructions. Despite its current success, the robustness of composed image retrieval methods to either (1) common corruptions or (2) variations of the textual descriptions have never been systematically evaluated. In this paper, we perform the first robustness study of composed image retrieval, establishing three new benchmarks for a systematic evaluation of robustness to common corruption (in both the textual and visual domains) and robustness in text understanding. For analysis of natural image corruption, we introduce two new large-scale benchmark datasets, CIRR-C and FashionIQ-C, for the open domains and fashion domains respectively–both of which feature 75 visual corruptions and 35 textual corruptions. To facilitate robust evaluation of text understanding, we introduce a new diagnostic dataset CIRR-D by expanding the CIRR dataset with synthetic data, specifically probing text understanding across variations in: numerical, attribute, object removal, and background. We introduce BenchCIR, a testbed for evaluating composed image retrieval model robustness with standardized evaluation protocols. Through benchmarking ten published models in the testbed, we reveal insights into how the composition of visual and textual modalities affects model robustness. The code is in https://suntongtongtong.github.io/BenchCIR/
Shitong Sun, Qilei Li, Shaogang Gong, Weitong Cai, Philip Torr 0001, Jindong Gu
Pattern Recognit.1
2024 Federated zero-shot learning with mid-level semantic knowledge transfer
abstract
Conventional centralized deep learning paradigms are not feasible when data from different sources cannot be shared due to data privacy or transmission limitation. To resolve this problem, federated learning has been introduced to transfer knowledge across multiple sources (clients) with non-shared data while optimizing a globally generalized central model (server). Existing federated learning paradigms mostly focus on transmitting image encoders that take instance-sensitive images as input, making them less generalizable and vulnerable to privacy inference attacks. In contrast, in this work, we consider transferring mid-level semantic knowledge (such as attribute) which is not sensitive to specific objects of interest and therefore is more privacy-preserving and general. To this end, we formulate a new Federated Zero-Shot Learning (FZSL) paradigm to learn mid-level semantic knowledge at multiple local clients with non-shared local data and cumulatively aggregate a globally generalized central model for deployment. To improve model discriminative ability, we explore semantic knowledge available from either a language or a vision-language foundation model in order to enrich the mid-level semantic space in FZSL. Extensive experiments on five zero-shot learning benchmark datasets validate the effectiveness of our approach for optimizing a generalizable federated learning model with mid-level semantic knowledge transfer.
Shitong Sun, Chenyang Si, Guile Wu, Shaogang Gong
Pattern Recognit.1
2023 FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation
abstract
Existing skeleton-based action recognition methods typically follow a centralized learning paradigm, which can pose privacy concerns when exposing human-related videos. Federated Learning (FL) has attracted much attention due to its outstanding advantages in privacy-preserving. However, directly applying FL approaches to skeleton videos suffers from unstable training. In this paper, we investigate and discover that the heterogeneous human topology graph structure is the crucial factor hindering training stability. To address this limitation, we pioneer a novel Federated Skeleton-based Action Recognition (FSAR) paradigm, which enables the construction of a globally generalized model without accessing local sensitive data. Specifically, we introduce an Adaptive Topology Structure (ATS), separating generalization and personalization by learning a domain-invariant topology shared across clients and a domain-specific topology decoupled from global model aggregation. Furthermore, we explore Multi-grain Knowledge Distillation (MKD) to mitigate the discrepancy between clients and server caused by distinct updating patterns through aligning shallow block-wise motion features. Extensive experiments on multiple datasets demonstrate that FSAR outperforms state-of-the-art FL-based methods while inherently protecting privacy.
Jingwen Guo, Hong Liu 0008, Shitong Sun, Tianyu Guo 0001, Min Zhang 0005, Chenyang Si
ICCV3
2021 Decentralised Person Re-Identification with Selective Knowledge Aggregation
Shitong Sun, Guile Wu, Shaogang Gong
BMVC1