EDBT 2026 Demo / reviewers in the wild / expert
Xiaoshan Yang
dblp:74/9989
· DBLP profile ↗
3ranked-venue papers in the field
0as first author
2since 2021 · last 2025
0000-0001-5453-9755ORCID · conflict
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VidCog: Empowering LLM with Long Video Understanding via Human-like Temporal Cognitive LoopabstractComprehending long-form videos, with their extensive temporal contexts and rich semantic complexities, remains a frontier challenge in video understanding. Recently, many existing methods offer promise for long video understanding yet often exhibit operational inefficiencies and suboptimal reasoning. These core challenges typically stem from fragmented multi-step reasoning, unreliable iterative control over information gathering, and visual retrieval strategies that inadequately adapt to query-aware granularities. In this paper, we propose a novel framework named VidCog that mimicks human cognitive processes to achieve robust and efficient long video understanding with Large Language Models (LLMs). VidCog features a Unified Reasoning Engine (URE) that transforms the discrete reasoning tasks into a single, cohesive LLM invocation, and a Contrastive Policy-Optimized Reasoning Gate (CPRG) that learns from relative preferences among contrastive query-exemplars to ensure reliable iterative decision-making. Furthermore, we propose Triadic Optimal Transport Visual Evidence Miner (TOT-VEM) to adaptively capture global-local temporal visual evidence by modeling it as a novel triadic optimal transport problem. Experiments on challenging long-video benchmarks demonstrate that VidCog consistently outperforms the strong baseline in both reasoning accuracy and efficiency, validating the superiority of the human-like cognitive loop. Xiaoshan Yang, Changsheng Xu |
MMAsia | 2 |
| 2021 | Few-shot Egocentric Multimodal Activity RecognitionabstractActivity recognition based on egocentric multimodal data collected by wearable devices has become increasingly popular recently. However, conventional activity recognition methods face the dilemma of the lack of large-scale labeled egocentric multimodal datasets due to the high cost of data collection. In this paper, we propose a new task of few-shot egocentric multimodal activity recognition, which has at least two significant challenges. On the one hand, it is difficult to extract effective features from the multimodal data sequences of video and sensor signals due to the scarcity of the samples. On the other hand, how to robustly recognize novel activity classes with very few labeled samples becomes another more critical challenge due to the complexity of the multimodal data. To resolve the challenges, we propose a two-stream graph network, which consists of a heterogeneous graph-based multimodal association module and a knowledge-aware activity classifier module. The former uses a heterogeneous graph network to comprehensively capture the dynamic and complementary information contained in the multimodal data stream. The latter learns robust activity classifiers through knowledge propagation among the classifier parameters of different classes. In addition, we adopt episodic training strategy to improve the generalization ability of the proposed few-shot activity recognition model. Experiments on two public datasets show that the proposed model achieves better performances than other baseline models. Jinxing Pan, Xiaoshan Yang, Yi Huang 0037, Changsheng Xu |
MMAsia | 2 |
| 2019 | Multimodal Attribute and Feature Embedding for Activity RecognitionabstractHuman Activity Recognition (HAR) automatically recognizes human activities such as daily life and work based on digital records, which is of great significance to medical and health fields. Egocentric video and human acceleration data comprehensively describe human activity patterns from different aspects, which have laid a foundation for activity recognition based on multimodal behavior data. However, on the one hand, the low-level multimodal signal structures differ greatly and the mapping to high-level activities is complicated. On the other hand, the activity labeling based on multimodal behavior data has high cost and limited data amount, which limits the technical development in this field. In this paper, an activity recognition model MAFE based on multimodal attribute feature embedding is proposed. Before the activity recognition, the middle-level attribute features are extracted from the low-level signals of different modes. On the one hand, the mapping complexity from the low-level signals to the high-level activities is reduced, and on the other hand, a large number of middle-level attribute labeling data can be used to reduce the dependency on the activity labeling data. We conducted experiments on Stanford-ECM datasets to verify the effectiveness of the proposed MAFE method. Yi Huang 0037, Wanting Yu, Xiaoshan Yang, Wei Wang 0354, Jitao Sang 0001 |
MMAsia | 4 |