Dian Shao

dblp:196/6225 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-0862-9941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion
abstract
Recognizing fine-grained actions from temporally corrupted skeleton sequences remains a significant challenge, particularly in real-world scenarios where online pose estimation often yields substantial missing data. Existing methods often struggle to accurately recover temporal dynamics and fine-grained spatial structures, resulting in the loss of subtle motion cues crucial for distinguishing similar actions. To address this, we propose FineTec, a unified framework for Fine-grained action recognition under Temporal Corruption. FineTec first restores a base skeleton sequence from corrupted input using context-aware completion with diverse temporal masking. Next, a skeleton-based spatial decomposition module partitions the skeleton into five semantic regions, further divides them into dynamic and static subgroups based on motion variance, and generates two augmented skeleton sequences via targeted perturbation. These, along with the base sequence, are then processed by a physics-driven estimation module, which utilizes Lagrangian dynamics to estimate joint accelerations. Finally, both the fused skeleton position sequence and the fused acceleration sequence are jointly fed into a GCN-based action recognition head. Extensive experiments on both coarse-grained (NTU-60, NTU-120) and fine-grained (Gym99, Gym288) benchmarks show that FineTec significantly outperforms previous methods under various levels of temporal corruption. Specifically, FineTec achieves top-1 accuracies of 89.1% and 78.1% on the challenging Gym99-severe and Gym288-severe settings, respectively, demonstrating its robustness and generalizability.
Dian Shao, Mingfei Shi, Like Liu
AAAI1
2026 PraNet-V2: Dual-Supervised Reverse Attention for Medical Image Segmentation
Bo-Cheng Hu, Ge-Peng Ji, Dian Shao, Deng-Ping Fan
Comput. Vis. Media3
2026 Knowledge Rectification for Camouflaged Object Detection: Unlocking Insights From Low-Resolution Data
abstract
Camouflaged object detection (COD) relies on multi-granularity structural information and fine-grained details to distinguish objects from highly similar backgrounds. Whereas low-resolution data lacks high-frequency cues such as textures and sharp edges, retaining only coarse structures. These not only weaken discriminative features but also introduce resolution-induced camouflage beyond natural blending. Existing COD methods assume high-resolution data and fail to address this dual-source ambiguity, resulting in significant performance degradation and underscoring the need for approaches that explicitly explore essential spatial priors under low-resolution constraints. Therefore, we propose KRNet, the first framework explicitly designed for COD in low-resolution settings. KRNet presents a Leader-Follower framework where the Leader extracts dual gold-standard distributions: conditional and hybrid, from supporting data to drive the Follower in rectifying knowledge learned from low-resolution data. The framework further benefits from a cross-consistency strategy, and a stronger time-prompt conditional encoder that improve the rectification of these distributions. Extensive experiments on benchmark datasets demonstrate that KRNet outperforms state-of-the-art COD methods and SR-assisted COD approaches, highlighting its effectiveness in tackling the challenges of low-resolution data in COD. Code: https://github.com/whyandbecause/KRNet/tree/main.
Juwei Guan, Xiaolin Fang 0001, Dian Shao, Haotian Gong, Tongxin Zhu, Zhipeng Cai 0001, Junzhou Luo
IEEE Trans. Image Process.5
2025 SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization
abstract
Human action understanding is crucial for the advancement of multimodal systems. While recent developments, driven by powerful large language models (LLMs), aim to be general enough to cover a wide range of categories, they often overlook the need for more specific capabilities. In this work, we address the more challenging task of Fine-grained Action Recognition (FAR), which focuses on detailed semantic labels within shorter temporal duration (e.g., ``salto backward tucked with 1 turn"). Given the high costs of annotating fine-grained labels and the substantial data needed for fine-tuning LLMs, we propose to adopt semi-supervised learning (SSL). Our framework, SeFAR, incorporates several innovative designs to tackle these challenges. Specifically, to capture sufficient visual details, we construct Dual-level temporal elements as more effective representations, based on which we design a new strong augmentation strategy for the Teacher-Student learning paradigm through involving moderate temporal perturbation. Furthermore, to handle the high uncertainty within the teacher model's predictions for FAR, we propose the Adaptive Regulation to stabilize the learning process. Experiments show that SeFAR achieves state-of-the-art performance on two FAR datasets, FineGym and FineDiving, across various data scopes, as well as two classical coarse-grained datasets, UCF101 and HMDB51. Further analysis and ablation studies validate the effectiveness of our designs. Additionally, we show that the features extracted by SeFAR could largely promote the ability of multimodal models to understand fine-grained and domain-specific semantics.
Yongle Huang, Zhenbang Xu, Zihan Jia, Haozhou Sun, Dian Shao
AAAI6
2025 FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance
abstract
Despite significant advances in video generation, synthesizing physically plausible human actions remains a persistent challenge, particularly in modeling fine-grained semantics and complex temporal dynamics. For instance, generating gymnastics routines such as "switch leap with 0.5 turn" poses substantial difficulties for current methods, often yielding unsatisfactory results. To bridge this gap, we propose FinePhys1, a Fine-grained human action generation framework that incorporates Physics to obtain effective skeletal guidance. Specifically, FinePhys first estimates 2D poses in an online manner and then performs 2D-to-3D dimension lifting via in-context learning. To mitigate the instability and limited interpretability of purely data-driven 3D poses, we further introduce a physics-based motion re-estimation module governed by Euler-Lagrange equations, calculating joint accelerations via bidirectional temporal updating. The physically predicted 3D poses are then fused with data-driven ones, offering multi-scale 2D heatmap guidance for the diffusion process. Evaluated on three fine-grained action subsets from FineGym (FX-JUMP, FX-TURN, and FX-SALTO), FinePhys significantly outperforms competitive baselines. Comprehensive qualitative results further demonstrate FinePhys’s ability to generate more natural and plausible fine-grained human actions.
Dian Shao, Mingfei Shi, Shengda Xu, Yongle Huang, Binglu Wang
CVPR1
2025 FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
Haojian Huang, Xinxiang Yin, Dian Shao
ACM Multimedia4
2025 CM-YOLO: Context Modulated Representation Learning for Ship Detection
abstract
Ship detection is essential for both military and civilian applications. Existing ship detection methods focus on prominent offshore ships, paying less attention to complex nearshore ships, which are easily confused with the intricate background. Utilizing contextual information, such as location and shape, can enhance ship detection and classification in complex environments. In this article, we propose a context modulated representation learning-based detection method termed as CM-YOLO. It adopts the classical detector design framework, which includes the backbone, neck, and head. The input image is sequentially processed through these components to obtain the detection results. Our method specifically optimizes ship detection in complex scenarios. To achieve this, we propose a dual path context enhancement neck (DCEN) to extract contextual information for ship detection. The neck builds on the path augmentation feature pyramid network with the proposed dual path context enhancement (DCE) module, which is designed to enhance feature representations by incorporating high-level semantic information. It captures long-range dependencies across both channel and spatial dimensions while suppressing irrelevant features. Additionally, to enhance the scale-aware capability of the head for detecting multiscale ships in complex environments, we introduce the multicontext boosted (MCB) detection head. The MCB can flexibly adjust the receptive field and extracts relevant context for ships of various scales using multiple large-kernel convolutions. We conduct experiments on three commonly used ship datasets: Seaships7000, ShipRSImageNet, DIOR-ship, and HRSC2016. Experiment results demonstrate that CM-YOLO achieves excellent performance compared with other leading ship detection methods.
Lingtong Min, Feiyang Dou, Dian Shao, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.4
2025 Multilevel Distribution Alignment for Multisource Universal Domain Adaptation
abstract
The multisource universal domain adaptation (MSUDA) relaxes the constraints between the source and target domains, enabling the transfer of knowledge between domains without any restrictions on the number of source domains and the existence of unknown (private) categories. However, identifying the unknown samples in the target domain is extremely challenging since there are no available samples with the same label in source domains. Another immense challenge lies in extracting domain-invariant features for knowledge transfer since there are distribution discrepancies between each source and target domain. In this article, we propose the multirepresentation DA network (MRDAN) to classify the unlabeled targets by harnessing multiple source domains with nonidentical label sets. First, we propose a threshold-free conflict-based predictions with uncertainty (CPU) module, which comprehensively mines the complementary knowledge from different source domains to identify both known and unknown samples simultaneously. To accurately extract the domain-invariant features for recognizing known and unknown samples, a multilevel distribution alignment (MLDA) strategy is introduced to decrease the distribution discrepancy between multiple domains with nonidentical category spaces progressively. Finally, comprehensive experiments conducted on three commonly used datasets demonstrate the effectiveness of the proposed MRDAN in recognizing both known and unknown samples.
Liang-Bo Ning 0001, Zuowei Zhang 0001, Weiping Ding 0001, Dian Shao, Yining Zhu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Just a Hint: Point-Supervised Camouflaged Object Detection
Dian Shao, Guangqian Guo, Shan Gao 0003
ECCV (35)2
2024 P2P: Transforming from Point Supervision to Explicit Visual Prompt for Object Detection and Segmentation
Guangqian Guo, Dian Shao, Sha Meng, Shan Gao 0003
IJCAI2
2024 FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
Haojian Huang, Junhao Dong 0001, Mingzhe Zheng, Dian Shao
ACM Multimedia5
2024 Effective Rotate: Learning Rotation-Robust Prototype for Aerial Object Detection
abstract
Aerial images often depict objects with arbitrary orientations, which pose challenges for conventional object detectors to detect and classify. To address this issue, rotation-equivariant Convolutional Neural Networks (CNNs) have been proposed to extract rotation-equivariant features. However, the orientation encoding in these networks is often unstable and noisy, deteriorating detection performance. In this paper, we first analyze the rotation-equivariant network. Then, we propose a Rotation-robust Prototype Generation (RPG) method, which consists of two parts, stabilization module and enhancement module. In stabilization module, we generate rotation-robust prototypes to increase the stability of cyclic shifts. In enhancement module, we use the obtained prototype to improve the response of the features to object semantics. The RPG method can be used as a plug-and-play module in both one-stage and two-stage detectors. With only 30 lines of code, we achieve an average 1% improvement on four challenging datasets, including DOTA-V1.5, DOTA-v1.0, DIOR-R, and HRSC2016.
Chaowei Wang, Guangqian Guo, Chang Liu 0047, Dian Shao, Shan Gao 0003
IEEE Trans. Geosci. Remote. Sens.4
2023 Tracking without Label: Unsupervised Multiple Object Tracking via Contrastive Similarity Learning
abstract
Unsupervised learning is a challenging task due to the lack of labels. Multiple Object Tracking (MOT), which inevitably suffers from mutual object interference, occlusion, etc., is even more difficult without label supervision. In this paper, we explore the latent consistency of sample features across video frames and propose an Unsupervised Contrastive Similarity Learning method, named UCSL, including three contrast modules: self-contrast, cross-contrast, and ambiguity contrast. Specifically, i) self-contrast uses intra-frame direct and inter-frame indirect contrast to obtain discriminative representations by maximizing self-similarity. ii) Cross-contrast aligns cross- and continuous-frame matching results, mitigating the persistent negative effect caused by object occlusion. And iii) ambiguity contrast matches ambiguous objects with each other to further increase the certainty of subsequent object association through an implicit manner. On existing benchmarks, our method outperforms the existing unsupervised methods using only limited help from ReID head, and even provides higher accuracy than lots of fully supervised methods.
Sha Meng, Dian Shao, Jiacheng Guo, Shan Gao 0003
ICCV2
2021 DIRV: Dense Interaction Region Voting for End-to-End Human-Object Interaction Detection
abstract
Recent years, human-object interaction (HOI) detection has achieved impressive advances. However, conventional two-stage methods are usually slow in inference. On the other hand, existing one-stage methods mainly focus on the union regions of interactions, which introduce unnecessary visual information as disturbances to HOI detection. To tackle the problems above, we propose a novel one-stage HOI detection approach DIRV in this paper, based on a new concept called interaction region for the HOI problem. Unlike previous methods, our approach concentrates on the densely sampled interaction regions across different scales for each human-object pair, so as to capture the subtle visual features that is most essential to the interaction. Moreover, in order to compensate for the detection flaws of a single interaction region, we introduce a novel voting strategy that makes full use of those overlapped interaction regions in place of conventional Non-Maximal Suppression (NMS). Extensive experiments on two popular benchmarks: V-COCO and HICO-DET show that our approach outperforms existing state-of-the-arts by a large margin with the highest inference speed and lightest network architecture. Our code is publicly available at www.github.com/MVIG-SJTU/DIRV.
Haoshu Fang, Yichen Xie 0002, Dian Shao, Cewu Lu
AAAI3
2021 DecAug: Augmenting HOI Detection via Decomposition
abstract
Human-object interaction (HOI) detection requires a large amount of annotated data. Current algorithms suffer from insufficient training samples and category imbalance within datasets. To increase data efficiency, in this paper, we propose an efficient and effective data augmentation method called DecAug for HOI detection. Based on our proposed object state similarity metric, object patterns across different HOIs are shared to augment local object appearance features without changing their states. Further, we shift spatial correlation between humans and objects to other feasible configurations with the aid of a pose-guided Gaussian Mixture Model while preserving their interactions. Experiments show that our method brings up to 3.3 mAP and 1.6 mAP improvements on V-COCO and HICO-DET dataset for two advanced models. Specifically, interactions with fewer samples enjoy more notable improvement. Our method can be easily integrated into various HOI detection models with negligible extra computational consumption.
Haoshu Fang, Yichen Xie 0002, Dian Shao, Yong-Lu Li 0001, Cewu Lu
AAAI3
2020 Intra- and Inter-Action Understanding via Temporal Action Parsing
abstract
Current methods for action recognition primarily rely on deep convolutional networks to derive feature embeddings of visual and motion features. While these methods have demonstrated remarkable performance on standard benchmarks, we are still in need of a better understanding as to how the videos, in particular their internal structures, relate to high-level semantics, which may lead to benefits in multiple aspects, e.g. interpretable predictions and even new methods that can take the recognition performances to a next level. Towards this goal, we construct TAPOS, a new dataset developed on sport videos with manual annotations of sub-actions, and conduct a study on temporal action parsing on top. Our study shows that a sport activity usually consists of multiple sub-actions and that the awareness of such temporal structures is beneficial to action recognition. We also investigate a number of temporal parsing methods, and thereon devise an improved method that is capable of mining sub-actions from training data without knowing the labels of them. On the constructed TAPOS, the proposed method is shown to reveal intra-action information, i.e. how action instances are made of sub-actions, and inter-action information, i.e. one specific sub-action may commonly appear in various actions.
Dian Shao, Yue Zhao 0006, Bo Dai 0002, Dahua Lin
CVPR1
2020 FineGym: A Hierarchical Video Dataset for Fine-Grained Action Understanding
abstract
On public benchmarks, current action recognition techniques have achieved great success. However, when used in real-world applications, e.g. sport analysis, which requires the capability of parsing an activity into phases and differentiating between subtly different actions, their performances remain far from being satisfactory. To take action recognition to a new level, we develop FineGym, a new dataset built on top of gymnasium videos. Compared to existing action recognition datasets, FineGym is distinguished in richness, quality, and diversity. In particular, it provides temporal annotations at both action and sub-action levels with a three-level semantic hierarchy. For example, a “balance beam” activity will be annotated as a sequence of elementary sub-actions derived from five sets: “leap-jump-hop”, “beam-turns”, “flight-salto”, “flight-handspring”, and “dismount”, where the sub-action in each set will be further annotated with finely defined class labels. This new level of granularity presents significant challenges for action recognition, e.g. how to parse the temporal structures from a coherent action, and how to distinguish between subtly different action classes. We systematically investigates different methods on this dataset and obtains a number of interesting findings. We hope this dataset could advance research towards action understanding.
Dian Shao, Yue Zhao 0006, Bo Dai 0002, Dahua Lin
CVPR1
2018 Find and Focus: Retrieve and Localize Video Events with Natural Language Queries
Dian Shao, Yue Zhao 0006, Qingqiu Huang, Yu Qiao 0001, Dahua Lin
ECCV (9)1