EDBT 2026 Demo / reviewers in the wild / expert
Jialin Gao
dblp:32/10264
· DBLP profile ↗
23ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
Guibao Shen, Quande Liu, Jialin Gao, Lan Du 0002, Cunjian Chen, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 5 |
| 2026 | FreeEdit: Mask-Free Reference-Based Image Editing With Multi-Modal InstructionabstractIntroducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based image editing, which can accurately reproduce the visual concept from the reference image based on user-friendly language instructions. Our approach leverages the multi-modal instruction encoder to encode language instructions to guide the editing process. This implicit way of locating the editing area eliminates the need for manual editing masks. To enhance the reconstruction of reference details, we introduce the Decoupled Residual Refer-Attention (DRRA) module. This module is designed to integrate fine-grained reference features extracted by a detail extractor into the image editing process in a residual way without interfering with the original self-attention. Given that existing datasets are unsuitable for reference-based image editing tasks, particularly due to the difficulty in constructing image triplets that include a reference image, we curate a high-quality dataset, FreeBench, using a newly developed twice-repainting scheme. FreeBench comprises the images before and after editing, detailed editing instructions, as well as a reference image that maintains the identity of the edited object, encompassing tasks such as object addition, replacement, and deletion. By conducting phased training on FreeBench followed by quality tuning, FreeEdit achieves high-quality zero-shot editing through convenient language instructions. We conduct extensive experiments to evaluate the effectiveness of FreeEdit across multiple task types, demonstrating its superiority over existing methods. Runze He, Linjiang Huang, Shaofei Huang 0001, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision TasksabstractEnd-to-end backpropagation remains the dominant training paradigm in deep learning, yet it suffers from inherent drawbacks, including update locking, high GPU memory consumption, and limited biological plausibility. Supervised local learning alleviates these issues by dividing the network into multiple blocks and training each block independently with an auxiliary network. However, gradient isolation also weakens the influence of downstream representations on earlier blocks, often resulting in a clear accuracy gap to end-to-end training. We propose Momentum Auxiliary Network++ (MAN++), a scalable framework that improves supervised local learning via a lightweight parameter-space transfer between adjacent blocks. MAN++ employs the exponential moving average (EMA) of parameters from adjacent blocks to propagate contextual information across the network. To address feature mismatches arising from direct EMA parameter transfer, we introduce a learnable scaling bias, which compensates feature statistics mismatch and stabilizes the transfer. Extensive experiments on image classification, object detection, and semantic segmentation across multiple architectures illustrate that MAN++ achieves accuracy on par with end-to-end training while substantially reducing GPU memory usage. These results position MAN++ as a practical and effective alternative to conventional backpropagation, offering new insights into scalable supervised local learning for vision tasks. Junhao Su, Hengyu Shi, Tianyang Han, Yurui Qiu, Junfeng Luo, Xiaoming Wei, Jialin Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Learning Prediction-aware Prior in Transformer Network for Accurate Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to precisely locate a spatio-temporal tube in an untrimmed video corresponding to a given language description. Many existing methods decouple spatial and temporal grounding as separate tasks, missing the strong interdependencies between the two, which are crucial for accurately aligning spatial regions (such as objects) with their motion over time. Thus, to enhance spatio-temporal associations, we introduce a new Prior-Driven Transformer Network (PDTNet) with predicted temporal boundaries as priors to guide object bounding boxes for improved spatial grounding over time. Firstly, PDTNet employs a temporal prior, termed reference query, to enhance discriminability between language-related and language-irrelevant visual content, improving temporal boundary localization. Further, the context within predicted temporal boundaries serves as another prior knowledge to modulate spatial features. We also introduce a prediction-aware Gaussian prior to precise object localization. This ensures consistent tube construction and accurate object localization. Extensive experiments on STVG benchmarks validate the effectiveness of PDTNet. Code is available at https://github.com/tongzhang111/PDTNet . Yongshun Gong, Jialin Gao, Yanyu Xu 0001, Xiushan Nie, Li-Zhen Cui 0001, Chengqi Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingabstractRecent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge stems from two key issues: first, incorporating spatial-temporal localization introduces a vast number of coordinate combinations, complicating the alignment of linguistic and visual coordinate representations; second, encoding fine-grained temporal and spatial information during video feature compression is inherently difficult. To address these issues, we propose LLaVA-ST, a MLLM for fine-grained spatial-temporal multimodal understanding. In LLaVA-ST, we propose Language-Aligned Positional Embedding, which embeds the textual coordinate special token into the visual space, simplifying the alignment of fine-grained spatial-temporal correspondences. Additionally, we design the Spatial-Temporal Packer, which decouples the feature compression of temporal and spatial resolutions into two distinct point-to-region attention processing streams. Furthermore, we propose ST-Align dataset with 4.3M training samples for fine-grained spatial-temporal multimodal understanding. With ST-align, we present a progressive training pipeline that aligns the visual and textual feature through sequential coarse-to-fine stages. Additionally, we introduce an ST-Align benchmark to evaluate spatial-temporal interleaved fine-grained understanding tasks, which include Spatial-Temporal Video Grounding (STVG) , Event Localization and Captioning (ELC) and Spatial Video Grounding (SVG). LLaVA-ST achieves outstanding performance on 11 benchmarks requiring fine-grained temporal, spatial, or spatial-temporal interleaving multimodal understanding. Our code, data and benchmark will be released at https://github.com/appletea233/LLaVA-ST. Shaofei Huang 0001, Tianrui Hui, Jialin Gao, Xiaoming Wei, Si Liu 0001 |
CVPR | 6 |
| 2025 | DOMR: Establishing Cross-View Segmentation via Dense Object MatchingabstractCross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual understanding. In this work, we propose the Dense Object Matching and Refinement (DOMR) framework to establish dense object correspondences across views. The framework centers around the Dense Object Matcher (DOM) module, which jointly models multiple objects. Unlike methods that directly match individual object masks to image features, DOM leverages both positional and semantic relationships among objects to find correspondences. DOM integrates a proposal generation module with a dense matching module that jointly encodes visual, spatial, and semantic cues, explicitly constructing inter-object relationships to achieve dense matching among objects. Furthermore, we combine DOM with a mask refinement head designed to improve the completeness and accuracy of the predicted masks, forming the complete DOMR framework. Extensive evaluations on the Ego-Exo4D benchmark demonstrate that our approach achieves state-of-the-art performance with a mean IoU of 49.7% on Ego→Exo and 55.2% on Exo→Ego. These results outperform those of previous methods by 5.8% and 4.3%, respectively, validating the effectiveness of our integrated approach for cross-view understanding. Jitong Liao, Yulu Gao, Shaofei Huang 0001, Jialin Gao, Jie Lei 0002, Ronghua Liang, Si Liu 0001 |
ACM Multimedia | 4 |
| 2025 | Multi-scale feature enhanced detection of foreign object intrusions on railways
Jialin Gao, Liqiang Zhu, Baoqing Guo |
J. Supercomput. | 1 |
| 2024 | Boundary Denoising for Video Activity LocalizationabstractVideo activity localization aims at understanding the semantic content in long, untrimmed videos and retrieving actions of interest. The retrieved action with its start and end locations can be used for highlight generation, temporal action detection, etc. Unfortunately, learning the exact boundary location of activities is highly challenging because temporal activities are continuous in time, and there are often no clear-cut transitions between actions. Moreover, the definition of the start and end of events is subjective, which may confuse the model. To alleviate the boundary ambiguity, we propose to study the video activity localization problem from a denoising perspective. Specifically, we propose an encoder-decoder model named DenosieLoc. During training, a set of temporal spans is randomly generated from the ground truth with a controlled noise scale. Then, we attempt to reverse this process by boundary denoising, allowing the localizer to predict activities with precise boundaries and resulting in faster convergence speed. Experiments show that DenosieLoc advances
several video activity understanding tasks. For example, we observe a gain of +12.36% average mAP on the QV-Highlights dataset.
Moreover, DenosieLoc achieves state-of-the-art performance on the MAD dataset but with much fewer predictions than others. Mengmeng Xu 0006, Mattia Soldan, Jialin Gao, Shuming Liu 0001, Juan-Manuel Pérez-Rúa, Bernard Ghanem |
ICLR | 3 |
| 2024 | From 2D to 3D: AISG-SLA Visual Localization Challenge
Jialin Gao, Bill Ong, Darld Lwi, Zhen Hao Ng, Xun Wei Yee, Mun-Thye Mak, Wee Siong Ng, See-Kiong Ng, Hui Ying Teo, Victor Khoo, Georg Bökman, Johan Edstedt, Kirill Brodt, Clémentin Boittiaux, Maxime Ferrera, Stepan Konev |
IJCAI | 1 |
| 2024 | High-compressed deepfake video detection with contrastive spatiotemporal distillation
Yizhe Zhu, Chunhui Zhang 0001, Jialin Gao, Xin Sun 0020, Zihan Rui, Xi Zhou 0001 |
Neurocomputing | 3 |
| 2024 | Learning Feature Semantic Matching for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (STVG) aims to localize a spatio-temporal tube, including temporal boundaries and object bounding boxes, that semantically corresponds to a given language description in an untrimmed video. The existing onestage solutions in this task face two significant challenges, namely, vision-text semantic misalignment and spatial mislocalization, which limit their performance in grounding. These two limitations are mainly caused by neglect of fine-grained alignment in crossmodality fusion and the reliance on a text-agnostic query in sequentially spatial localization. To address these issues, we propose an effective model with a newly designed Feature Semantic Matching (FSM) module based on a Transformer architecture to address the above issues. Our method introduces a crossmodal feature matching module to achieve multi-granularity alignment between video and text while preventing the weakening of important features during the feature fusion stage. Additionally, we design a query-modulated matching module to facilitate text-relevant tube construction by multiple query generation and tubulet sequence matching. To ensure the quality of tube construction, we employ a novel mismatching rectify contrastive loss to rectify the mismatching between the learnable query and the objects corresponding to the text descriptions by restricting the generated spatial query. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on two challenging STVG benchmarks. Hao Fang 0010, Hao Zhang 0048, Jialin Gao, Xiankai Lu, Xiushan Nie, Yilong Yin |
IEEE Trans. Multim. | 4 |
| 2023 | AVForensics: Audio-driven Deepfake Video Detection with Masking Strategy in Self-supervisionabstractExisting cross-dataset deepfake detection approaches exploit mouth-related mismatches between the auditory and visual modalities in fake videos to enhance generalisation to unseen forgeries. However, such methods inevitably suffer performance degradation with limited or unaltered mouth motions, we argue that face forgery detection consistently benefits from using high-level cues across the whole face region. In this paper, we propose a two-phase audio-driven multi-modal transformer-based framework, termed AVForensics, to perform deepfake video content detection from an audio-visual matching view related to full face. In the first pre-training phase, we apply the novel uniform masking strategy to model global facial features and learn temporally dense video representations in a self-supervised cross-modal manner, by capturing the natural correspondence between the visual and auditory modalities regardless of large-scaled labelled data and heavy memory usage. Then we use these learned representations to fine-tune for the down-stream deepfake detection task in the second phase, which encourages the model to offer accurate predictions based on captured global facial movement features. Extensive experiments and visualizations on various public datasets demonstrate the superiority of our self-supervised pre-trained method for achieving generalisable and robust deepfake video detection. Yizhe Zhu, Jialin Gao, Xi Zhou 0001 |
ICMR | 2 |
| 2023 | Exploiting enhanced and robust RGB-D face representation via progressive multi-modal learning
Yizhe Zhu, Jialin Gao, Tianshu Wu |
Pattern Recognit. Lett. | 2 |
| 2023 | Video Moment Retrieval via Comprehensive Relation-Aware NetworkabstractVideo moment retrieval aims to retrieve a target moment from an untrimmed video that semantically corresponds to the given language query. Existing methods commonly treat it as a regression task or a ranking task from the perspective of computer vision. Most of these works neglect comprehensive relations between video content and language context at a multi-granularity level and fail to efficiently model temporal relations among different video moments. In this paper, we formulate video moment retrieval into video reading comprehension by treating the input video as a text passage and language query as a question. To tackle the above impediments, we propose a Comprehensive Relation-aware Network (CRNet) to perceive comprehensive relations from extensive aspects. Specifically, we unite visual and textual features simultaneously at both clip-level and moment-level to thoroughly exploit inter-modality information, leading to a coarse-and-fine cross-modal interaction. Moreover, a background suppression module is introduced to restrain irrelevant background clips, meanwhile, a novel IoU attention mechanism and graph attention layer are efficiently devised to focus on the dependencies among highly-correlated video moments for the best choice selection. In-depth experiments on three public datasets TACoS, ActivityNet Captions, and Charades-STA demonstrate the superiority of our solution. Xin Sun 0020, Jialin Gao, Yizhe Zhu, Xi Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Attention-guided Fine-grained Feature Learning For Robust Face Forgery DetectionabstractWith the swift development of deep learning, hyper-realistic images generated by advanced facial manipulation techniques have posed a serious threat to the trustworthiness of digital media information. Most existing approaches formulate face forgery detection as a coarse-grained classification problem. They still significantly focus on low-level semantic cues which are sensitive to common corruptions such as video compression and generalise poorly to unseen forgeries. To address this issue, we propose a novel fine-grained feature learning framework for face forgery detection. To be specific, we adopt fine-grained frequency decomposition via a patch-wise manner to extract more sufficient information hidden in the frequency domain. We also propose the depth-wise separable attention module to select more informative features and captures the fine-grained cues from different input spaces. Moreover, to further explore the essential discrepancies, attention-guided feature augment module is introduced to fully exploit the multi-domain relations and incorporate frequency features into spatial clues. Extensive experiments and visualizations on public datasets fully demonstrate the effectiveness and robustness of our method against the state-of-the-art competitors. Yizhe Zhu, Jialin Gao |
ICPR | 2 |
| 2022 | You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in VideosabstractMoment retrieval in videos is a challenging task that aims to retrieve the most relevant video moment in an untrimmed video given a sentence description. Previous methods tend to perform self-modal learning and cross-modal interaction in a coarse manner, which neglect fine-grained clues contained in video content, query context, and their alignment. To this end, we propose a novel Multi-Granularity Perception Network (MGPN) that perceives intra-modality and inter-modality information at a multi-granularity level. Specifically, we formulate moment retrieval as a multi-choice reading comprehension task and integrate human reading strategies into our framework. A coarse-grained feature encoder and a co-attention mechanism are utilized to obtain a preliminary perception of intra-modality and inter-modality information. Then a fine-grained feature encoder and a conditioned interaction module are introduced to enhance the initial perception inspired by how humans address reading comprehension problems. Moreover, to alleviate the huge computation burden of some existing methods, we further design an efficient choice comparison module and reduce the hidden size with imperceptible quality loss. Extensive experiments on Charades-STA, TACoS, and ActivityNet Captions datasets demonstrate that our solution outperforms existing state-of-the-art methods. Xin Sun 0020, Jialin Gao, Xi Zhou 0001 |
SIGIR | 3 |
| 2022 | Efficient Video Grounding With Which-Where Reading ComprehensionabstractVideo grounding aims at localizing the temporal moment related to the given language description, which is very helpful to many cross-modal content understanding applications like visual question answering and sentence-video search. Existing approaches usually directly regress the temporal boundaries of an event described by a query sentence in the video sequence. This direct regression manner often encounters a large decision space due to diverse target events and variable video durations, leading to inaccurate localization as well as inefficient grounding. This paper presents an efficient framework termed from which to where to facilitate video grounding. The core idea is imitating the reading comprehension process to gradually narrow the decision space, in what we decompose the direct regression into two steps. The “which” step first roughly selects a candidate area by evaluating which video segment in the predefined set is closest to the ground truth. To this end, we formulate this step into a multi-choice reading comprehension problem and propose a criterion to select the best-matched segment. In this way, the excessive decision space is effectively reduced. The “where” step aims to precisely regress the temporal boundary of the selected video segment from the shrunk decision space. We thus introduce a triple-span representation for each candidate video segment to use the regional context for better boundary regression. The “which” and “where” steps can be combined into a unified framework and learned end-to-end, leading to an efficient video grounding system. Extensive experiments on Charades-STA, ActivityNet-Captions, and TACoS benchmarks clearly demonstrate the effectiveness of our framework. Jialin Gao, Xin Sun 0020, Bernard Ghanem, Xi Zhou 0001, Shiming Ge |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Relation-aware Video Reading Comprehension for Temporal Language GroundingabstractTemporal language grounding in videos aims to localize the temporal span relevant to the given query sentence.Previous methods treat it either as a boundary regression task or a span extraction task.This paper will formulate temporal language grounding into video reading comprehension and propose a Relation-aware Network (RaNet) to address it.This framework aims to select a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction.A choicequery interactor is proposed to match the visual and textual information simultaneously in sentence-moment and token-moment levels, leading to a coarse-and-fine cross-modal interaction.Moreover, a novel multi-choice relation constructor is introduced by leveraging graph convolution to capture the dependencies among video moment choices for the best choice selection.Extensive experiments on ActivityNet-Captions, TACoS, and Charades-STA demonstrate the effectiveness of our solution.Codes will be available at https: //github.com/Huntersxsx/RaNet. Jialin Gao, Xin Sun 0020, Mengmeng Xu 0006, Xi Zhou 0001, Bernard Ghanem |
EMNLP (1) | 1 |
| 2021 | Skeleton-Based Action Recognition With Focusing-Diffusion Graph Convolutional NetworksabstractGraph Convolutional Networks have been successfully applied in skeleton-based action recognition. The key is fully exploring the spatial-temporal context. This letter proposes a Focusing-Diffusion Graph Convolutional Network (FDGCN) to address this issue. Each skeleton frame is first decomposed into two opposite-direction graphs for subsequent focusing and diffusion processes. Next, the focusing process generates a spatial-level representation for each frame individually by an attention module. This representation is regarded as a supernode to aggregate the feature from each joint node in each frame for spatial context extraction. After generating supernodes for the entire sequence, a transformer encoder layer is proposed to capture the temporal context further. Finally, these supernodes pass the embedded spatial-temporal context back to the spatial joints through the diffusion graph in the diffusing process. Extensive experiments on the NTU RGB+D and Skeleton-Kinetics benchmarks demonstrate the effectiveness of our approach. Jialin Gao, Tong He 0002, Xi Zhou 0001, Shiming Ge |
IEEE Signal Process. Lett. | 1 |
| 2021 | Self-Guided Body Part Alignment With Relation Transformers for Occluded Person Re-IdentificationabstractPerson re-identification in the wild is often challenged by occlusion. Existing methods mainly rely on learned external cues like pose or parsing to ease occlusion distraction. This knowledge highly related to body semantics may introduce alignment effects, leading to additional requirements for dedicated training data and inference computation. We propose the Self-guided Body Part Alignment method that learns cue-free semantic-aligned local prediction for feature representations to avoid high-cost dependence on external cues. First, scale-wise global spatial attention is utilized to determine essential body parts automatically. A relation transformer network is then employed to predict semantic-aligned local parts, guided with anchored global information by constraint loss. Similarity metrics for all parts are merged with threshold conditions to filter invisible body parts comprehensively. Experimental results on occluded and holistic person reID benchmarks show the proposed method outperforms other cue-relied and cue-free methods. As far as we know, this is the first method that applies transformer networks on local predictions for occluded reID tasks. Guanshuo Wang, Jialin Gao, Xi Zhou 0001, Shiming Ge |
IEEE Signal Process. Lett. | 3 |
| 2020 | Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid NetworkabstractAccurate temporal action proposals play an important role in detecting actions from untrimmed videos. The existing approaches have difficulties in capturing global contextual information and simultaneously localizing actions with different durations. To this end, we propose a Relation-aware pyramid Network (RapNet) to generate highly accurate temporal action proposals. In RapNet, a novel relation-aware module is introduced to exploit bi-directional long-range relations between local features for context distilling. This embedded module enhances the RapNet in terms of its multi-granularity temporal proposal generation ability, given predefined anchor boxes. We further introduce a two-stage adjustment scheme to refine the proposal boundaries and measure their confidence in containing an action with snippet-level actionness. Extensive experiments on the challenging ActivityNet and THUMOS14 benchmarks demonstrate our RapNet generates superior accurate proposals over the existing state-of-the-art methods. Jialin Gao, Zhixiang Shi, Guanshuo Wang, Yufeng Yuan, Shiming Ge, Xi Zhou 0001 |
AAAI | 1 |
| 2020 | Convolutional neural network with adaptive inferential framework for skeleton-based action recognition
Hong'en Huang, Hang Su 0006, Zhigang Chang, Mingyang Yu 0005, Jialin Gao, Xinzhe Li 0002, Shibao Zheng |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | General Interaction-Aware Neural Network for Action Recognition
Jialin Gao, Guanshuo Wang, Yufeng Yuan |
PRICAI (3) | 1 |