EDBT 2026 Demo / reviewers in the wild / expert
Jianjun Gao 0005
dblp:92/9182-5
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0004-9137-2869ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptSR: Cascade Prompting for Lightweight Image Super-ResolutionabstractAlthough the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-based self-attention modeling. The quadratic computational complexity relative to window size restricts its ability to use a large window size for expanding the receptive field while maintaining low computational costs. To address this challenge, we propose PromptSR, a novel prompt-empowered lightweight image SR method. The core component is the proposed cascade prompting block (CPB), which enhances global information access and local refinement via three cascaded prompting layers: a global anchor prompting layer (GAPL) and two local prompting layers (LPLs). The GAPL leverages downscaled features as anchors to construct low-dimensional anchor prompts (APs) through cross-scale attention, significantly reducing computational costs. These APs, with enhanced global perception, are then used to provide global prompts, efficiently facilitating long-range token connections. The two LPLs subsequently combine category-based self-attention and window-based self-attention to refine the representation in a coarse-to-fine manner. They leverage attention maps from the GAPL as additional global prompts, enabling them to perceive features globally at different granularities for adaptive local refinement. In this way, the proposed CPB effectively combines global priors and local details, significantly enlarging the receptive field while maintaining the low computational costs of our PromptSR. The experimental results demonstrate the superiority of our method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations. Our code will be released at https://github.com/wenyang001/PromptSR. Wenyang Liu, Jianjun Gao 0005, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 3 |
| 2025 | A Structure-Aware and Motion-Adaptive Framework for 3D Human Pose Estimation with Mamba
Jianjun Gao 0005, Kim-Hui Yap |
ICCV | 3 |
| 2025 | From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionabstractRecent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edge device deployment. Meanwhile, conventional GSR models often lack generalization ability, falling short in recognizing unseen and rare situations. In this paper, we exploit transferring knowledge from a teacher MLLM to a small GSR model to enhance its generalization and zero-shot abilities, thereby introducing the task of Open-vocabulary Grounded Situation Recognition (Ov-GSR). To achieve this, we propose Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations. Specifically, the MIPD framework first leverages the LLM-based Judgmental Rationales Generator (JRG) to construct positive and negative glimpse and gaze rationales enriched with contextual semantic information. The proposed scene-aware and instance-perception prompts are then introduced to align rationales with visual information from the MLLM teacher via the Negative-Guided Multimodal Prompting Alignment (NMPA) module, effectively capturing holistic and perceptual multimodal knowledge. Finally, the aligned multimodal knowledge is distilled into the student Ov-GSR model, providing a stronger foundation for generalization that enhances situation understanding, bridges the gap between seen and unseen scenarios, and mitigates prediction bias in rare cases. We evaluate MIPD on the refined Ov-SWiG dataset, achieving superior performance on seen, rare, and unseen situations, and further demonstrate improved unseen detection on the HICO-DET dataset. Jianjun Gao 0005, Wenyang Liu, Kejun Wu, Yi Wang 0068, Soo Chin Liew |
ACM Multimedia | 3 |
| 2025 | SSH-Net: A self-supervised and hybrid network for noisy image watermark removal
Wenyang Liu, Jianjun Gao 0005, Kim-Hui Yap |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | CL-HOI: Cross-level human-object interaction distillation from multimodal large language models
Jianjun Gao 0005, Wenyang Liu, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
Knowl. Based Syst. | 1 |
| 2025 | OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian TrackingabstractMultiple pedestrian tracking is crucial for enhancing safety and efficiency in intelligent transport and autonomous driving systems by predicting movements and enabling adaptive decision-making in dynamic environments. It optimizes traffic flow, facilitates human interaction, and ensures compliance with regulations. However, it faces the challenge of tracking pedestrians in the presence of occlusion. Existing methods overlook effects caused by abnormal detections during partial occlusion. Subsequently, these abnormal detections can lead to inaccurate motion estimation, unreliable appearance features, and unfair association. To address these issues, we propose an adaptive occlusion-aware multiple pedestrian tracker, OccluTrack, to mitigate the effects caused by partial occlusion. Specifically, we first introduce a plug-and-play abnormal motion suppression mechanism into the Kalman Filter to adaptively detect and suppress outlier motions caused by partial occlusion. Second, we develop a pose-guided re-identification (Re-ID) module to extract discriminative part features for partially occluded pedestrians. Last, we develop a new occlusion-aware association method towards fair Intersection over Union (IoU) and appearance embedding distance measurement for occluded pedestrians. Extensive evaluation results demonstrate that our method outperforms state-of-the-art methods on MOTChallenge and DanceTrack datasets. Particularly, the performance improvements on IDF1 and ID Switches, as well as visualized results, demonstrate the effectiveness of our method in multiple pedestrian tracking. Jianjun Gao 0005, Yi Wang 0068, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Empowering Large Language Model for Continual Video Question Answering with Collaborative PromptingabstractIn recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content.In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting.To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting.These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research.Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14% accuracy on NExT-QA and 71.24% accuracy on DramaQA, highlighting its practical relevance and effectiveness. Jianjun Gao 0005, Wenyang Liu, Runzhong Zhang, Kim-Hui Yap |
EMNLP | 3 |
| 2024 | Contextual Human Object Interaction Understanding from Pre-Trained Large Language ModelabstractExisting human object interaction (HOI) detection methods have introduced zero-shot learning techniques to recognize unseen interactions, but they still have limitations in understanding context information and comprehensive reasoning. To overcome these limitations, we propose a novel HOI learning framework, ContextHOI, which serves as an effective contextual HOI detector to enhance contextual understanding and zero-shot reasoning ability. The main contributions of the proposed ContextHOI are a novel context-mining decoder and a powerful interaction reasoning large language model (LLM). The context-mining decoder aims to extract linguistic contextual information from a pre-trained vision-language model. Based on the extracted context information, the proposed interaction reasoning LLM further enhances the zero-shot reasoning ability by leveraging rich linguistic knowledge. Extensive evaluation demonstrates that our proposed framework outperforms existing zero-shot methods on the HICO-DET and SWIG-HOI datasets, as high as 19.34% mAP on unseen interaction can be achieved. Jianjun Gao 0005, Kim-Hui Yap, Kejun Wu, Duc Tri Phan, Kratika Garg, Boon Siew Han |
ICASSP | 1 |
| 2024 | Hdplifter: Hierarchical Dynamics Perception For 2D-to-3D Human Pose LiftingabstractRecent 2D-to-3D pose lifting networks have achieved remarkable success in monocular 3D human pose estimation through learning joint dependencies. We observed that the extracted 2D pose sequences encountered spatial pose topology ambiguity and temporal movement patterns information loss. Existing methods overlook these intrinsic limitations, resulting in inferior inferences of the corresponding 3D poses. To address these, we introduce the Hierarchical Dynamics Pose Lifter (HDPLifter), which captures spatial human joint connections and subtle temporal movement patterns while maintaining global modeling through hierarchical perception. Specifically, we propose a Structure-aware Spatial Transformer using adaptive topology learning to efficiently integrate spatial joint connections. Moreover, a novel Hierarchical Temporal Transformer is utilized to comprehensively capture subtle joint movement patterns along with global movement patterns with a scaleable receptive field. In both modules, we utilize a 2D depth-wise convolution as a feedforward network to further gather local joint correlations in the spatial and temporal domains simultaneously. Our model, HDPLifter, surpasses the state-of-the-art approach (Motion-BERT) on Human3.6M and MPI-INF-3DHP datasets with P1 errors of 38.0mm and 14.4mm, respectively, while utilizing only 1/5 of the parameters compared to it. Jianjun Gao 0005, Duc Tri Phan, Kim-Hui Yap |
ICIP | 2 |
| 2024 | CM2-Net: Continual Cross-Modal Mapping Network For Driver Action RecognitionabstractDriver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and depth. Nevertheless, compared to RGB modality only, it is always laborious and costly to collect extensive data for all types of non-RGB modalities in car cabin environments. Therefore, previous works have suggested independently learning each non-RGB modality by fine-tuning a model pretrained on RGB videos, but these methods are less effective in extracting informative features when faced with newly-incoming modalities due to large domain gaps. In contrast, we propose a Continual Cross-Modal Mapping Network (CM2Net) to continually learn each newly-incoming modality with instructive prompts from the previously-learned modalities. Specifically, we have developed Accumulative Cross-modal Mapping Prompting (ACMP), to map the discriminative and informative features learned from previous modalities into the feature space of newly-incoming modalities. Then, when faced with newly-incoming modalities, these mapped features are able to provide effective prompts for which features should be extracted and prioritized. These prompts are accumulating throughout the continual learning process, thereby boosting further recognition performances. Extensive experiments conducted on the Drive&Act dataset demonstrate the performance superiority of $\mathrm{CM}^{2}-\mathrm{Net}$ on both uni- and multi-modal driver action recognition. Jianjun Gao 0005, Dan Lin 0008, Wenyang Liu, Kim-Hui Yap |
ICIP | 4 |
| 2024 | Temporal Sentence Grounding with Temporally Global Textual KnowledgeabstractTemporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multimodal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method. Runzhong Zhang, Jianjun Gao 0005, Kejun Wu, Kim-Hui Yap, Yi Wang 0068 |
ICME | 3 |
| 2024 | MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action RecognitionabstractDriver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging, being gloomy and dark, and with the development of sensors, various cameras such as IR and depth cameras have emerged for analyzing drivers' behaviors. Therefore, in this paper, we propose a novel multimodal fusion transformer, named Multi-Fuser, which identifies cross-modal interrelations and interactions among multimodal car cabin videos and adaptively integrates different modalities for improved representations. Specifically, MultiFuser comprises layers of Bi-decomposed Modules to model spatiotemporal features, with a modality synthesizer for multi-modal features integration. Each Bi-decomposed Module includes a Modal Expertise ViT block for extracting modality-specific features and a Patch-wise Adaptive Fusion block for efficient cross-modal fusion. Extensive experiments are conducted on Drive&Act dataset and the results demonstrate the efficacy of our proposed approach. Jianjun Gao 0005, Dan Lin 0008, Kim-Hui Yap |
MMSP | 3 |
| 2023 | METFormer: A Motion Enhanced Transformer for Multiple Object TrackingabstractMultiple object tracking (MOT) is an important task in computer vision, especially video analytics. Transformer-based methods are emerging approaches using both tracking and detection queries. However, motion modeling in existing transformer-based methods lacks effective association capability. Thus, this paper introduces a new METFormer model, a Motion Enhanced TransFormer-based tracker with a novel global-local motion context learning technique to mitigate the lack of motion information in existing transformer-based methods. The global-local motion context learning technique first centers on difference-guided global motion learning to obtain temporal information from adjacent frames. Based on global motion, we leverage context-aware local object motion modelling to study motion patterns and enhance the feature representation for individual objects. Experimental results on the benchmark MOT17 dataset show that our proposed method can surpass the state-of-the-art Trackformer [21] by 1.8% on IDF1 and 21.7% on ID Switches under public detection settings. Jianjun Gao 0005, Kim-Hui Yap, Yi Wang 0068, Kratika Garg, Boon Siew Han |
ISCAS | 1 |