EDBT 2026 Demo / reviewers in the wild / expert
Yangjun Ou
dblp:245/3872
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-4142-4833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RCAGFusion: Recursive cross-attention guided deep feature fusion for indoor scene recognition
Chen Wang 0058, Hebao Qiu, Xiong Pan, Yangjun Ou |
Image Vis. Comput. | 5 |
| 2026 | Fold-Llama: Data-efficient robotic fabric folding via geometry-to-text encoding and lightweight LLMs fine-tuning
Ruhan He, Lianqing Yu, Xianyi Zeng, Yangjun Ou |
Knowl. Based Syst. | 5 |
| 2026 | Class-aware prototype augmentation and decoupled feature distillation for class-incremental learning
Chengdong Wang, Yangjun Ou, Xianfang Tang, Yuan Wu 0007, Wuxuan Shi, Xueliang Liu |
Pattern Recognit. | 2 |
| 2026 | SemCo: Toward Semantic Coherent Visual Relationship ForecastingabstractVisual Relationship Forecasting (VRF) in video aims to anticipate relations among objects without observing future visual content. The task relies on capturing and modeling the semantic coherence in object interactions, as it underpins the evolution of events and scenes in videos. However, existing VRF datasets provide limited support for learning such coherence. Their noisy annotations fail to reflect distinct action or scene changes, and weak correlations exist between different actions and relational transitions in subject-object pairs. Furthermore, existing methods struggle to distinguish similar relationships and overfit to unchanging relationships in consecutive frames rather than reflecting the semantic coherence of object interactions. To address these challenges, we present SemCoBench, a benchmark that emphasizes semantic coherence for visual relationship forecasting in video. Based on action labels and short-term subject-object pairs, SemCoBench decomposes relationship categories and dynamics by cleaning and reorganizing video datasets to ensure predicting semantic coherence in object interactions. In addition, we propose the Semantic Coherent Transformer (SemCoFormer), which consists of a Relationship Augmented Module (RAM) and a Coherent Reasoning Module (CRM). The RAM bidirectionally enhances cross-modal features to better distinguish similar relationships, while the CRM performs sparse encoding of cross-frame relation transitions to model dynamic relational changes. The experimental results on SemCoBench demonstrate that modeling the semantic coherence is a key step toward reasonable, fine-grained, and diverse visual relationship forecasting, contributing to a more comprehensive understanding of video scenes. Our project page: https://lyao-61.github.io/ SemCo/. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | CPR-CLIP: Cross-modal Consistent and Prompt-diverse Regularized CLIP for Action RecognitionabstractPrompt learning has emerged as an effective strategy for adapting multimodal vision-language models to downstream tasks such as video action recognition, which has been widely applied in human-computer interaction, video surveillance and medical fields. However, existing approaches often suffer from overfitting to task-specific distributions and prompt collapse, especially under limited data or when class semantics are highly similar. In this work, we propose Cross-modal Consistent and Prompt-diverse Regularized CLIP (CPR-CLIP), a flexible framework that enhances the generalization of prompt learning in action recognition. Specifically, we introduce Cross-modal Consistency Regularization (CCR) strategy to preserve the original feature representations encoded by the frozen CLIP encoder, thereby reducing overfitting caused by prompt adaptation. Additionally, we design a Prompt Diversity Regularization (PDR) term to encourage category-wise separation among prompts, alleviating prompt collapse and enhancing discriminability. Experiments on HMDB51, UCF101, and SSv2 under base-to-novel and few-shot settings show that CPR-CLIP consistently outperforms existing methods, achieving strong generalization to novel classes. Yangjun Ou, Chen Wang 0058 |
MMAsia | 2 |
| 2025 | From Body Parts to Holistic Action: A Fine-Grained Teacher-Student CLIP for Action RecognitionabstractAction recognition in dynamic video remains challenging, particularly when distinguishing between visually similar actions. While existing methods often rely on holistic representations, they overlook the fine-grained details that are significant for accurate classification. We propose a novel Fine-grained Teacher-student CLIP (FT-CLIP) that integrates body part analysis with holistic action recognition through a teacher-student architecture, bridging the gap between fine-grained action parsing and overall action understanding. The teacher model processes individual body parts alongside specialized description to generate part-specific features, which are then aggregated and distilled into the student model. Through knowledge distillation with learnable prompts, the student model effectively learns to capture subtle action distinctions while maintaining efficient inference. FT-CLIP achieves a more nuanced understanding of complex actions by progressing from detailed body part analysis to comprehensive action recognition. Experiments on Kinetics-TPS under a fully-supervised setting and on HMDB51 and UCF101 under a zero-shot setting demonstrate the effectiveness of our method. Yangjun Ou, Ruhan He, Chi Liu 0004 |
IEEE Signal Process. Lett. | 1 |
| 2025 | HARG: Hierarchical Adaptive Reasoning Graph for Activity ParsingabstractAs a video understanding task, activity parsing aims at encompassing actions into multiple levels of activity components, including activity, sub-activity and atomic action, enabling understanding of complex video scenes within multimedia systems. Existing methods form activity parsing as a multi-task learning problem to predict multi-granular activity labels simultaneously, which ignores modeling the hierarchical structure and the fine-grained transitions of activity components at different levels. In this paper, we propose a Hierarchical Adaptive Reasoning Graph (HARG) to model the hierarchical structure (i.e., object level$\rightarrow$atomic action level$\rightarrow$activity level) dynamically and precisely. To achieve that, an object reasoning graph (ORG) and an atomic action reasoning graph (ARG) are designed to reason fine-grained information transitions between multiple actors at different levels. In addition, an adaptive segmentation module (ASM) is investigated for bridging the gap among different levels, permitting step-by-step reasoning from the object level to the atomic action level. Experimental results show our method outperforms state-of-the-art methods on two activity parsing datasets, achieving hierarchical modeling and fine-grained reasoning for activity understanding. The code is available on GitHub:https://github.com/whuoyj/HARG. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Open-Vocabulary RGB-Thermal Semantic Segmentation
Xiaoyun Yan, Zhaojing Wang, Junwei Tang, Yangjun Ou, Xinrong Hu, Tao Peng 0006 |
ECCV (74) | 6 |
| 2023 | MARANet: Multi-scale Adaptive Region Attention Network for Few-Shot Learning
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006 |
CGI (1) | 3 |
| 2023 | Graphormer-Based Contextual Reasoning Network for Small Object Detection
Jia Chen 0012, Xiyang Li, Yangjun Ou, Xinrong Hu, Tao Peng 0006 |
PRCV (9) | 3 |
| 2023 | Multiple visual relationship forecasting and arrangement in videos
Wanping Ouyang, Yaosi Hu, Yangjun Ou, Zhenzhong Chen 0001 |
Neurocomputing | 3 |
| 2023 | 3D Deformable Convolution Temporal Reasoning network for action recognition
Yangjun Ou |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Object-Relation Reasoning Graph for Action RecognitionabstractAction recognition is a challenging task since the attributes of objects as well as their relationships change constantly in the video. Existing methods mainly use object-level graphs or scene graphs to represent the dynamics of objects and relationships, but ignore modeling the fine-grained relationship transitions directly. In this paper, we propose an Object-Relation Reasoning Graph (OR2G) for reasoning about action in videos. By combining an object-level graph (OG) and a relation-level graph (RG), the proposed OR2G catches the attribute transitions of objects and reasons about the relationship transitions between objects simultaneously. In addition, a graph aggregating module (GAM) is investigated by applying the multi-head edge-to-node message passing operation. GAM feeds back the information from the relation node to the object node and enhances the coupling between the object-level graph and the relation-level graph. Experiments in video action recognition demonstrate the effectiveness of our approach when compared with the state-of-the-art methods. Yangjun Ou, Li Mi, Zhenzhong Chen 0001 |
CVPR | 1 |
| 2021 | Multimodal Local-Global Attention Network for Affective Video Content AnalysisabstractWith the rapid development of video distribution and broadcasting, affective video content analysis has attracted a lot of research and development activities recently. Predicting emotional responses of movie audiences is a challenging task in affective computing, since the induced emotions can be considered relatively subjective. In this article, we propose a multimodal local-global attention network (MMLGAN) for affective video content analysis. Inspired by the multimodal integration effect, we extend the attention mechanism to multi-level fusion and design a multimodal fusion unit to obtain a global representation of affective video. The multimodal fusion unit selects key parts from multimodal local streams in the local attention stage and captures the information distribution across time in the global attention stage. Experiments on the LIRIS-ACCEDE dataset, the MediaEval 2015 and 2016 datasets, the FilmStim dataset, the DEAP dataset and the VideoEmotion dataset demonstrate the effectiveness of our approach when compared with the state-of-the-art methods. Yangjun Ou, Zhenzhong Chen 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Band-Independent Encoder-Decoder Network for Pan-Sharpening of Remote Sensing ImagesabstractPan-sharpening is a fundamental task for remote sensing image processing. It aims at creating a high-resolution multispectral (HRMS) image from a multispectral (MS) image and a panchromatic (PAN) image. In this article, a new band-independent encoder-decoder network is proposed for pan-sharpening. The network takes a single band of the MS (BMS) image, the PAN image, and the low-resolution PAN (LRPAN) image as inputs. The output of the network is the corresponding band of high-resolution MS (HRBMS) image. In this way, the network can process MS images with any number of bands. The overall structure of the network consists of two encoder-decoder modules at low-resolution and high-resolution, respectively. An auxiliary LRPAN image is used to speed up the training and improve the performance. The partly shared network and hierarchical structure for low-resolution and high-resolution enable a better fusion of features extracted from different scales. With a fast fine-tuning strategy, the trained model can be applied to images from different sensors. Experiments performed on different data sets demonstrate that the proposed method outperforms several state-of-the-art pan-sharpening methods in both visual appearance and objective indexes, and the single-band evaluation results further verify the superiority of the proposed method. Chi Liu 0004, Yongjun Zhang 0002, Shugen Wang, Yangjun Ou, Yi Wan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Pan-Sharpening Using an Efficient Bidirectional Pyramid NetworkabstractPan-sharpening is an important preprocessing step for remote sensing image processing tasks; it fuses a low-resolution multispectral image and a high-resolution (HR) panchromatic (PAN) image to reconstruct a HR multispectral (MS) image. This paper introduces a new end-to-end bidirectional pyramid network for pan-sharpening. The overall structure of the proposed network is a bidirectional pyramid, which permits the network to process MS and PAN images in two separate branches level by level. At each level of the network, spatial details extracted from the PAN image are injected into the upsampled MS image to reconstruct the pan-sharpened image from coarse resolution to fine resolution. Subpixel convolutional layers and the enhanced residual blocks are used to make the network efficient. Comparison of the results obtained with our proposed method and the results using other widely used state-of-the-art approaches confirms that our proposed method outperforms the others in visual appearance and objective indexes. Yongjun Zhang 0002, Chi Liu 0004, Yangjun Ou |
IEEE Trans. Geosci. Remote. Sens. | 4 |