Xingning Dong

dblp:317/0005 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-0245-9064ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
abstract
The user base of short video apps has experienced unprecedented growth in recent years, resulting in a significant demand for video content analysis. In particular, text-video retrieval, which aims to find the top matching videos given text descriptions from a vast video corpus, is an essential function, the primary challenge of which is to bridge the modality gap. Nevertheless, most existing approaches treat texts merely as discrete tokens and neglect their syntax structures. Moreover, the abundant spatial and temporal clues in videos are often underutilized due to the lack of interaction with text. To address these issues, we argue that using texts as guidance to focus on relevant temporal frames and spatial regions within videos is beneficial. In this paper, we propose a novel Syntax-Hierarchy-Enhanced text-video retrieval method (SHE-Net) that exploits the inherent semantic and syntax hierarchy of texts to bridge the modality gap from two perspectives. First, to facilitate a more fine-grained integration of visual content, we employ the text syntax hierarchy, which reveals the grammatical structure of text descriptions, to guide the visual representations. Second, to further enhance the multi-modal interaction and alignment, we also utilize the syntax hierarchy to guide the similarity calculation. We evaluated our method on four public text-video retrieval datasets of MSR-VTT, MSVD, DiDeMo, and ActivityNet. The experimental results and ablation studies confirm the advantages of our proposed method.
Xuzheng Yu, Chen Jiang 0006, Xingning Dong, Tian Gan 0002, Ming Yang 0007, Qingpei Guo
IEEE Trans. Circuits Syst. Video Technol.3
2024 EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints
Xingning Dong, Tian Gan 0002, Chunluan Zhou, Ming Yang 0007, Qingpei Guo
IJCAI2
2024 MBC-ATA: Maximum Binary Classification and Anchor-based Triplet Augmentation for Unbiased Scene Graph Generation
Xingning Dong, Jinfei Gao, Pei Shen, Tian Gan 0002
MMAsia2
2024 M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
abstract
We present a Recipe for Effective and Efficient zero-shot video-text Retrieval, dubbed M2-RAAP. Upon popular image-text models like CLIP, most current adaptation-based video-text pre-training methods are confronted by three major issues, i.e., noisy data corpus, time-consuming pre-training, and limited performance gain. Towards this end, we conduct a comprehensive study including four critical steps in video-text pre-training. Specifically, we investigate 1) data filtering and refinement, 2) video input type selection, 3) temporal modeling, and 4) video feature enhancement. We then summarize this empirical study into the M2-RAAP recipe, where our technical contributions lie in 1) the data filtering and text re-writing pipeline resulting in 1M high-quality bilingual video-text pairs, 2) the promotion of video inputs with key-frames to accelerate pre-training, and 3) the Auxiliary-Caption-Guided (ACG) strategy to enhance video features. We conduct extensive experiments by adapting three image-text foundation models on two refined video-text datasets from different languages, validating the robustness and reproducibility of M2-RAAP for adaptation-based pre-training. Results demonstrate that M2-RAAP yields superior performance with significantly less data (-90%) and time consumption (-95%), establishing a new SOTA on four English zero-shot retrieval datasets and two Chinese ones. Codebase and refined bilingual data annotations are available at https://github.com/alipay/Ant-Multi-Modal-Framework/tree/main/prj/M2_RAAP.
Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang 0007, Qingpei Guo
SIGIR1
2024 SNP-S3: Shared Network Pre-Training and Significant Semantic Strengthening for Various Video-Text Tasks
abstract
We present a framework for learning cross-modal video representations by directly pre-training on raw data to facilitate various downstream video-text tasks. Our main contributions lie in the pre-training framework and proxy tasks. First, based on the shortcomings of two mainstream pixel-level pre-training architectures (limited applications or less efficient), we propose Shared Network Pre-training (SNP). By employing one shared BERT-type network to refine textual and cross-modal features simultaneously, SNP is lightweight and could support various downstream applications. Second, based on the intuition that people always pay attention to several “significant words” when understanding a sentence, we propose the Significant Semantic Strengthening (S3) strategy, which includes a novel masking and matching proxy task to promote the pre-training performance. Experiments conducted on three downstream video-text tasks and six datasets demonstrate that, we establish a new state-of-the-art in pixel-level video-text pre-training; we also achieve a satisfactory balance between the pre-training efficiency and the fine-tuning performance. The codebase and pre-trained models are available athttps://github.com/dongxingning/SNPS3.
Xingning Dong, Qingpei Guo, Tian Gan 0002, Qing Wang 0059, Jianlong Wu, Xiangyuan Ren
IEEE Trans. Circuits Syst. Video Technol.1
2023 CNVid-3.5M: Build, Filter, and Pre-Train the Large-Scale Public Chinese Video-Text Dataset
abstract
Owing to well-designed large-scale video-text datasets, recent years have witnessed tremendous progress in video-text Pre-training. However, existing large-scale video-text datasets are mostly English-only. Though there are certain methods studying the Chinese video-text Pre-training, they pre-train their models on private datasets whose videos and text are unavailable. This lack of large-scale public datasets and benchmarks in Chinese hampers the research and downstream applications of Chinese video-text Pre-training. Towards this end, we release and benchmark CNVid-3.5M, a large-scale public cross-modal dataset containing over 3.5M Chinese video-text pairs. We summarize our contributions by three verbs, i.e., “Build”, “Filter”, and “Pretrain”: 1) To build a public Chinese video-text dataset, we collect over 4.5M videos from the Chinese websites. 2) To improve the data quality, we propose a novel method to filter out 1M weakly-paired videos, resulting in the CNVid-3.5M dataset. And 3) we benchmark CNVid-3.5M with three mainstream pixel-level Pre-training architectures. At last, we propose the Hard Sample Curriculum Learning strategy to promote the Pre-training performance. To the best of our knowledge, CNVid-3.5M is the largest public video-text dataset in Chinese, and we provide the first pixel-level benchmarks for Chinese video-text Pre-training. The dataset, codebase, and pre-trained models are available at https://github.com/CNVid/CNVid-3.5M.
Tian Gan 0002, Qing Wang 0068, Xingning Dong, Xiangyuan Ren, Liqiang Nie, Qingpei Guo
CVPR3
2023 DBiased-P: Dual-Biased Predicate Predictor for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) is to abstract the objects and their semantic relationships within a given image. Current SGG performance is mainly limited by the biased predicate prediction caused by the long-tailed data distribution. Though many unbiased SGG methods have emerged to enhance the prediction of the tail predicates, their improvements on the tail predicates are often accompanied by the deterioration on the head ones, leading the prediction overly debiased. Toward this end, in this work, we propose a Dual-Biased Predicate Predictor (DBiased-P) to boost the unbiased SGG, which comprises a re-weighted primary classifier and an unweighted auxiliary classifier. The former classifier is tail-biased and used for the final predicate prediction, while the latter one is head-biased and designed to boost the head predicate prediction of the primary classifier by a head-oriented soft regularization. Experiments conducted on Visual Genome and Open Image datasets indicate the superiority of our DBiased-P in unbiased SGG, which significantly improves the recall@50 of the state-of-the-art unbiased SGG method DT2-ACBS from 23.3% to 55.5% as well as the mean recall@50 from 35.9% to 37.7%.
Xianjing Han, Xuemeng Song, Xingning Dong, Yinwei Wei, Meng Liu 0006, Liqiang Nie
IEEE Trans. Multim.3
2022 Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation
abstract
Scene Graph Generation, which generally follows a regular encoder-decoder pipeline, aims to first encode the visual contents within the given image and then parse them into a compact summary graph. Existing SGG approaches generally not only neglect the insufficient modality fusion between vision and language, but also fail to provide informative predicates due to the biased relationship predictions, leading SGG far from practical. Towards this end, we first present a novel Stacked Hybrid-Attention network, which facilitates the intra-modal refinement as well as the intermodal interaction, to serve as the encoder. We then devise an innovative Group Collaborative Learning strategy to optimize the decoder. Particularly, based on the observation that the recognition capability of one classifier is limited towards an extremely unbalanced dataset, we first deploy a group of classifiers that are expert in distinguishing different subsets of classes, and then cooperatively optimize them from two aspects to promote the unbiased SGG. Experiments conducted on VG and GQA datasets demonstrate that, we not only establish a new state-of-the-art in the unbiased metric, but also nearly double the performance compared with two baselines. Our code is available at https://github.com/dongxingning/SHA-GCL-for-SGG.
Xingning Dong, Tian Gan 0002, Xuemeng Song, Jianlong Wu, Liqiang Nie
CVPR1
2022 Divide-and-Conquer Predictor for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to detect the objects and their pairwise predicates in an image. Existing SGG methods mainly fulfil the challenging predicate prediction task that involves severe long-tailed data distribution with a single classifier. However, we argue that this may be enough to differentiate predicates that present obvious differences (e.g.,$on$and$near$), but not sufficient to distinguish similar predicates that only have subtle differences (e.g.,$on$and$standing~on$). Towards this end, we divide the predicate prediction into a few sub-tasks with a Divide-and-Conquer Predictor (DC-Predictor). Specifically, we first develop an offline pattern-predicate correlation mining algorithm to discover the similar predicates that share the same object interaction pattern. Based on that, we devise a general pattern classifier and a set of specific predicate classifiers for DC-Predictor. The former works on recognizing the pattern of a given object pair and routing it to the corresponding specific predicate classifier, while the latter aims to differentiate similar predicates in each specific pattern. In addition, we introduce the Bayesian Personalized Ranking loss in each specific predicate classifier to enhance the pairwise differentiation between head predicates and their similar ones. Experiments on VG150 and GQA datasets show the superiority of our model over state-of-the-art methods.
Xianjing Han, Xingning Dong, Xuemeng Song, Tian Gan 0002, Yibing Zhan, Yan Yan 0002, Liqiang Nie
IEEE Trans. Circuits Syst. Video Technol.2