VLDB 2026 Research / reviewers in the wild / expert
Gongfu Li
dblp:245/3489
· DBLP profile ↗
14ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0003-0655-5317ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scenario Shared Instance Modeling for Click-through Rate PredictionabstractMulti-scenario recommendation (MSR) is a popular training paradigm in industrial platforms for uniformly integrating information from multiple scenarios and serving them simultaneously. A key challenge in MSR research is accurately identifying the commonalities and distinctive information between scenarios. Currently, most existing MSR methods focus on implicitly extracting this information from the architectural level. However, this continues to increase the complexity and training overhead of MSR. Furthermore, the custom components responsible for extracting implicit information in each MSR method are too dependent on the specific MSR architecture and are not easily reused in other methods. Given these challenges, we first show in a motivating experiment that it may be beneficial to explicitly select a reasonable set of shared instances that can affect parameter optimization in all scenarios during the training of MSR, i.e., to explicitly obtain the critical information required for MSR from the data level. Then, this paper proposes SSIM with an adaptive selection network. Specifically, SSIM can be integrated with existing MSR methods in a lightweight way to adaptively select an informative and shareable subset of instances from each scenario to improve recommendations. In particular, the selected multi-scenario shared subset has extraordinary reusability and can be easily saved to benefit model training of various future MSR models. Finally, we evaluate SSIM and demonstrate its effectiveness through experiments on two public multi-scenario benchmarks and an online A/B test. Dugang Liu, Chaohua Yang 0002, Yuwen Fu, Xing Tang 0007, Gongfu Li, Fuyuan Lyu, Xiuqiang He 0001, Zhong Ming 0001 |
KDD (1) | 5 |
| 2024 | Contrastive topic-enhanced network for video captioning
Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Hong Man, Xiangmin Xu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Hugs Bring Double Benefits: Unsupervised Cross-Modal Hashing with Multi-granularity Aligned Transformers
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
Int. J. Comput. Vis. | 6 |
| 2024 | Confidence-Aware Sentiment Quantification via Sentiment Perturbation ModelingabstractSentiment Quantification aims to detect the overall sentiment polarity of users from a set of reviews corresponding to a target. Existing methods equally treat and aggregate individual reviews' sentiment to judge the overall sentiment polarity. However, the confidence of each review is not equal in sentiment quantification where sentiment perturbation arising from high- and low-confidence reviews may degrade the accuracy of Sentiment Quantification. Specifically, fake reviews with deceptive sentiments are low confidence, which perturbs the overall sentiment prediction. Whereas, some reviews generated by responsible users are high confidence. They contain authoritative suggestions so they should be emphasized in Sentiment Quantification. In this paper, we design and build COSE, a confidence-aware sentiment quantification framework, which can measure the confidence of individual reviews to eliminate sentiment perturbation and facilitate sentiment quantification. We design a Review Graph that achieves review confidence modeling in an unsupervised manner and obtains review confidence representations. Moreover, we develop a dynamic fusion attention mechanism, which produces sentiment “de-perturbation” vectors to eliminate the sentiment perturbation based on the confidence representations. Extensive experiments on large-scale review datasets validate the significant superiority of COSE over the state-of-the-art. Xiangyun Tang, Dongliang Liao, Meng Shen 0001, Liehuang Zhu, Shen Huang, Gongfu Li, Hong Man, Jin Xu 0014 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Complementarity-Aware Space Learning for Video-Text RetrievalabstractIn general, videos are powerful at recording physical patterns (e.g., spatial layout) while texts are great at describing abstract symbols (e.g., emotion). When video and text are used in multi-modal tasks, they are claimed to be complementary and their distinct information is crucial. However, when it comes to cross-modal tasks (e.g., retrieval), existing works usually use their common part in the form of common space learning while their distinct information is abandoned. In this paper, we argue that distinct information is also beneficial for cross-modal retrieval. To address this problem, we propose a divide-and-conquer learning approach, namely Complementarity-aware Space Learning (CSL), by recasting this challenge into learning of two spaces (i.e., latent and symbolic spaces) to simultaneously explore their common and distinct information by considering multi-modal complementary character. Specifically, we first propose to learn a symbolic space from video with a memory-based video encoder and a symbolic generator. In contrast, we also introduce learning a latent space from text with a text encoder and a memory-based latent feature selector. Finally, we propose a complementarity-aware loss by integrating two spaces to facilitate video-text retrieval tasks. Extensive experiments show that our approach outperforms existing state-of-the-art methods by 5.1%, 2.1% and 0.9% of R@10 for text-to-video retrieval on three benchmarks, respectively. Ablation study also verifies that the distinct information from video and text improves the retrieval performance. Trained models and source code have been released athttps://github.com/NovaMind-Z/CSL. Jinkuan Zhu, Pengpeng Zeng, Lianli Gao, Gongfu Li, Dongliang Liao, Jingkuan Song |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | End-to-End Pre-Training With Hierarchical Matching and Momentum Contrast for Text-Video RetrievalabstractLately, video-language pre-training and text-video retrieval have attracted significant attention with the explosion of multimedia data on the Internet. However, existing approaches for video-language pre-training typically limit the exploitation of the hierarchical semantic information in videos, such as frame semantic information and global video semantic information. In this work, we present an end-to-end pre-training network with Hierarchical Matching and Momentum Contrast named HMMC. The key idea is to explore the hierarchical semantic information in videos via multilevel semantic matching between videos and texts. This design is motivated by the observation that if a video semantically matches a text (can be a title, tag or caption), the frames in this video usually have semantic connections with the text and show higher similarity than frames in other videos. Hierarchical matching is mainly realized by two proxy tasks: Video-Text Matching (VTM) and Frame-Text Matching (FTM). Another proxy task: Frame Adjacency Matching (FAM) is proposed to enhance the single visual modality representations while training from scratch. Furthermore, momentum contrast framework was introduced into HMMC to form a multimodal momentum contrast framework, enabling HMMC to incorporate more negative samples for contrastive learning which contributes to the generalization of representations. We also collected a large-scale Chinese video-language dataset (over 763k unique videos) named CHVTT to explore the multilevel semantic connections between videos and texts. Experimental results on two major Text-video retrieval benchmark datasets demonstrate the advantages of our methods. We release our code at https://github.com/cheetah003/HMMC. Wenxue Shen, Jingkuan Song, Xiaosu Zhu, Gongfu Li, Heng Tao Shen |
IEEE Trans. Image Process. | 4 |
| 2023 | Keyword-Based Diverse Image Retrieval With Variational Multiple Instance GraphabstractThe task of cross-modal image retrieval has recently attracted considerable research attention. In real-world scenarios, keyword-based queries issued by users are usually short and have broad semantics. Therefore, semantic diversity is as important as retrieval accuracy in such user-oriented services, which improves user experience. However, most typical cross-modal image retrieval methods based on single point query embedding inevitably result in low semantic diversity, while existing diverse retrieval approaches frequently lead to low accuracy due to a lack of cross-modal understanding. To address this challenge, we introduce an end-to-end solution termed variational multiple instance graph (VMIG), in which a continuous semantic space is learned to capture diverse query semantics, and the retrieval task is formulated as a multiple instance learning problems to connect diverse features across modalities. Specifically, a query-guided variational autoencoder is employed to model the continuous semantic space instead of learning a single-point embedding. Afterward, multiple instances of the image and query are obtained by sampling in the continuous semantic space and applying multihead attention, respectively. Thereafter, an instance graph is constructed to remove noisy instances and align cross-modal semantics. Finally, heterogeneous modalities are robustly fused under multiple losses. Extensive experiments on two real-world datasets have well verified the effectiveness of our proposed solution in both retrieval accuracy and semantic diversity. Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Da Cao, Hong Man |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Hugs Are Better Than Handshakes: Unsupervised Cross-Modal Transformer Hashing with Multi-granularity Alignment
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
BMVC | 6 |
| 2022 | A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalabstractCross-modal retrieval aims to build correspondence between multiple modalities by learning a common representation space. Typically, an image can match multiple texts semantically and vice versa, which significantly increases the difficulty of this task. To address this problem, probabilistic embedding is proposed to quantify these many-to-many relationships. However, existing datasets (e.g., MS-COCO) and metrics (e.g., Recall@K) cannot fully represent these diversity correspondences due to non-exhaustive annotations. Based on this observation, we utilize semantic correlation computed by CIDEr to find the potential correspondences. Then we present an effective metric, named Average Semantic Precision (ASP), which can measure the ranking precision of semantic correlation for retrieval sets. Additionally, we introduce a novel and concise objective, coined Differentiable ASP Approximation (DAA). Concretely, DAA can optimize ASP directly by making the ranking function of ASP differentiable through a sigmoid function. To verify the effectiveness of our approach, extensive experiments are conducted on MS-COCO, CUB Captions, and Flickr30K, which are commonly used in cross-modal retrieval. The results show that our approach obtains superior performance over the state-of-the-art approaches on all metrics. The code and trained models are released at https://github.com/leolee99/2022-NeurIPS-DAA. Jingkuan Song, Lianli Gao, Pengpeng Zeng, Haonan Zhang 0003, Gongfu Li |
NeurIPS | 6 |
| 2022 | Hybrid Contrastive Quantization for Efficient Cross-View Video RetrievalabstractWith the recent boom of video-based social platforms (e.g., YouTube and TikTok), video retrieval using sentence queries has become an important demand and attracts increasing research attention. Despite the decent performance, existing text-video retrieval models in vision and language communities are impractical for large-scale Web search because they adopt brute-force search based on high-dimensional embeddings. To improve efficiency, Web search engines widely apply vector compression libraries (e.g., FAISS [26]) to post-process the learned embeddings. Unfortunately, separate compression from feature encoding degrades the robustness of representations and incurs performance decay. To pursue a better balance between performance and efficiency, we propose the first quantized representation learning method for cross-view video retrieval, namely Hybrid Contrastive Quantization (HCQ). Specifically, HCQ learns both coarse-grained and fine-grained quantizations with transformers, which provide complementary understandings for texts and videos and preserve comprehensive semantic information. By performing Asymmetric-Quantized Contrastive Learning (AQ-CL) across views, HCQ aligns texts and videos at coarse-grained and multiple fine-grained levels. This hybrid-grained learning strategy serves as strong supervision on the cross-view video quantization model, where contrastive learning at different levels can be mutually promoted. Extensive experiments on three Web video benchmark datasets demonstrate that HCQ achieves competitive performance with state-of-the-art non-compressed retrieval methods while showing high efficiency in storage and computation. Code and configurations are available at https://github.com/gimpong/WWW22-HCQ. Jinpeng Wang 0002, Bin Chen 0011, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shutao Xia, Jin Xu 0014 |
WWW | 5 |
| 2021 | Hierarchical Coherence Modeling for Document Quality AssessmentabstractText coherence plays a key role in document quality assessment. Most existing text coherence methods only focus on similarity of adjacent sentences. However, local coherence exists in sentences with broader contexts and diverse rhetoric relations, rather than just adjacent sentences similarity. Besides, the highlevel text coherence is also an important aspect of document quality. To this end, we propose a hierarchical coherence model for document quality assessment. In our model, we implement a local attention mechanism to capture the location semantics, bilinear tensor layer for measure coherence and max-coherence pooling for acquiring high-level coherence. We evaluate the proposed method on two realistic tasks: news quality judgement and automated essay scoring. Experimental results demonstrate the validity and superiority of our work. Dongliang Liao, Gongfu Li |
AAAI | 3 |
| 2021 | Adaptive Feature Weight Learning For Robust Clustering Problem with Sparse ConstraintabstractClustering task has been greatly developed in recent years like partition-based and graph-based methods. However, in terms of improving robustness, most existing algorithms only focus on noise and outliers between data, while ignoring the noise in feature space. To deal with this situation, we propose a novel weight learning mechanism to adaptively reweight each feature in the data. Combining with the clustering task, we further propose a robust fuzzy K-Means model based on the auto-weighted feature learning, which can effectively reduce the proportion of noisy features. Besides, a regularization term is introduced into our model to make the sample-to-clusters memberships of each sample have suitable sparsity. Specifically, we design an effective strategy to determine the value of the regularization parameter. The experimental results on both synthetic and real-world datasets demonstrate that our model has better performance than other classical algorithms. Feiping Nie 0001, Wei Chang 0002, Xuelong Li 0001, Jin Xu 0014, Gongfu Li |
ICASSP | 5 |
| 2020 | Cognitive Representation Learning of Self-Media Online Article QualityabstractThe automatic quality assessment of self-media online articles is an urgent and new issue, which is of great value to the online recommendation and search. Different from traditional and well-formed articles, self-media online articles are mainly created by users, which have the appearance characteristics of different text levels and multi-modal hybrid editing, along with the potential characteristics of diverse content, different styles, large semantic spans and good interactive experience requirements. To solve these challenges, we establish a joint model CoQAN in combination with the layout organization, writing characteristics and text semantics, designing different representation learning subnetworks, especially for the feature learning process and interactive reading habits on mobile terminals. It is more consistent with the cognitive style of expressing an expert's evaluation of articles. We have also constructed a large scale real-world assessment dataset. Extensive experimental results show that the proposed framework significantly outperforms state-of-the-art methods, and effectively learns and integrates different factors of the online article quality assessment. Shen Huang, Gongfu Li, Qiang Deng, Dongliang Liao, Pengda Si, Yujiu Yang 0001 |
ACM Multimedia | 3 |
| 2019 | Popularity Prediction on Online Articles with Deep Fusion of Temporal Process and Content FeaturesabstractPredicting the popularity of online article sheds light to many applications such as recommendation, advertising and information retrieval. However, there are several technical challenges to be addressed for developing the best of predictive capability. (1) The popularity fluctuates under impacts of external factors, which are unpredictable and hard to capture. (2) Content and meta-data features, largely determining the online content popularity, are usually multi-modal and nontrivial to model. (3) Besides, it also needs to figure out how to integrate temporal process and content features modeling for popularity prediction in different lifecycle stages of online articles. In this paper, we propose a Deep Fusion of Temporal process and Content features (DFTC) method to tackle them. For modeling the temporal popularity process, we adopt the recurrent neural network and convolutional neural network. For multi-modal content features, we exploit the hierarchical attention network and embedding technique. Finally, a temporal attention fusion is employed for dynamically integrating all these parts. Using datasets collected from WeChat, we show that the proposed model significantly outperforms state-of-the-art approaches on popularity prediction. Dongliang Liao, Gongfu Li, Weiqing Liu |
AAAI | 3 |