VLDB 2026 Research / reviewers in the wild / expert
Yiqing Cai
dblp:47/10243
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BeLink: Behavior graph network for unsupervised user identity linkage
Xingkong Ma, Mengmeng Guo, Houjie Qiu, Yiqing Cai |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | QCRD: Quality-guided Contrastive Rationale Distillation for Large Language ModelsabstractThe deployment of large language models (LLMs) faces considerable challenges concerning resource constraints and inference efficiency. Recent research has increasingly focused on smaller, task-specific models enhanced by distilling knowledge from LLMs. However, prior studies have often overlooked the diversity and quality of knowledge, especially the untapped potential of negative knowledge. Constructing effective negative knowledge remains severely understudied. In this paper, we introduce a novel framework called quality-guided contrastive rationale distillation aimed at enhancing reasoning capabilities through contrastive knowledge learning. For positive knowledge, we enrich its diversity through temperature sampling and employ self-consistency for further denoising and refinement. For negative knowledge, we propose an innovative self-adversarial approach that generates low-quality rationales by sampling previous iterations of smaller language models, embracing the idea that one can learn from one’s own weaknesses. A contrastive loss is developed to distill both positive and negative knowledge into smaller language models, where an online-updating discriminator is integrated to assess qualities of rationales and assign them appropriate weights, optimizing the training process. Through extensive experiments across multiple reasoning tasks, we demonstrate that our method consistently outperforms existing distillation techniques, yielding higher-quality rationales. Yiqing Cai, Zhida Huang |
EMNLP | 4 |
| 2025 | Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal ModelsabstractWei Wang, Zhaowei Li, Qi Xu, Linfeng Li, YiQing Cai, Botian Jiang, Hang Song, Xingcan Hu, Pengyu Wang, Li Xiao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wei Wang 0378, Yiqing Cai, Botian Jiang, Xingcan Hu, Pengyu Wang 0006, Li Xiao 0002 |
EMNLP | 5 |
| 2025 | PersAD: adversarial protection against text-based personality inference
Houjie Qiu, Xingkong Ma, Bo Liu 0014, Yiqing Cai, Baoyun Peng |
Data Min. Knowl. Discov. | 4 |
| 2025 | Psycholinguistic knowledge-guided graph network for personality detection of silent users
Houjie Qiu, Xingkong Ma, Bo Liu 0014, Yiqing Cai, Zhaoyun Ding |
Inf. Process. Manag. | 4 |
| 2024 | Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationabstractIn the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset. Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He |
AAAI | 3 |
| 2024 | GroundingGPT: Language Enhanced Multi-modal Grounding ModelabstractZhaowei Li, Qi Xu, Dong Zhang, Hang Song, YiQing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang, Tao Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yiqing Cai, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang |
ACL (1) | 5 |
| 2024 | SSBM: A spatially separated boxes-based multi-tab website fingerprinting model
Xueshu Hong, Xingkong Ma, Yiqing Cai, Bo Liu 0014 |
J. Netw. Comput. Appl. | 4 |
| 2023 | Explicit Invariant Feature Induced Cross-Domain Crowd CountingabstractCross-domain crowd counting has shown progressively improved performance. However, most methods fail to explicitly consider the transferability of different features between source and target domains. In this paper, we propose an innovative explicit Invariant Feature induced Cross-domain Knowledge Transformation framework to address the inconsistent domain-invariant features of different domains. The main idea is to explicitly extract domain-invariant features from both source and target domains, which builds a bridge to transfer more rich knowledge between two domains. The framework consists of three parts, global feature decoupling (GFD), relation exploration and alignment (REA), and graph-guided knowledge enhancement (GKE). In the GFD module, domain-invariant features are efficiently decoupled from domain-specific ones in two domains, which allows the model to distinguish crowds features from backgrounds in the complex scenes. In the REA module both inter-domain relation graph (Inter-RG) and intra-domain relation graph (Intra-RG) are built. Specifically, Inter-RG aggregates multi-scale domain-invariant features between two domains and further aligns local-level invariant features. Intra-RG preserves taskrelated specific information to assist the domain alignment. Furthermore, GKE strategy models the confidence of pseudolabels to further enhance the adaptability of the target domain. Various experiments show our method achieves state-of-theart performance on the standard benchmarks. Code is available at https://github.com/caiyiqing/IF-CKT. Yiqing Cai, Lianggangxu Chen, Haoyue Guan, Shaohui Lin, Changhong Lu, Changbo Wang, Gaoqi He |
AAAI | 1 |
| 2023 | Video-based spatio-temporal scene graph generation with efficient self-supervision tasks
Lianggangxu Chen, Yiqing Cai, Changhong Lu, Changbo Wang, Gaoqi He |
Multim. Tools Appl. | 2 |
| 2023 | Global Representation Guided Adaptive Fusion Network for Stable Video Crowd CountingabstractModern crowd counting methods in natural scenes, even when video datasets are available, are mostly based on images. Because of background interference or occlusion in the scene, these methods can easily lead to mutations and instability in density prediction. There has been minimal research on how to exploit the inherent consistency among adjacent frames to achieve high estimation accuracy of video sequences. In this study, we explore the long-term global temporal consistency in the video sequence and propose a novel Global Representation Guided Adaptive Fusion Network (GRGAF) for video crowd counting. The primary aim is to establish a long-term temporal representation among consecutive frames to guide the density estimation of local frames, which can alleviate the prediction instability caused by background noise and occlusions in crowd scenes. Moreover, in order to further enforce the temporal consistency, we apply the generative adversarial learning scheme and design a global-local joint loss, which can make the estimated density maps more temporally coherent. Extensive experiments on four challenging video-based crowd counting datasets (FDST, DroneCrowd, MALL and UCSD) demonstrate that our method makes effective use of spatio-temporal information of video and outperforms the other state-of-the-art approach. Yiqing Cai, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
IEEE Trans. Multim. | 1 |
| 2022 | DH-GCN: Saliency-Aware Complex Scene Graph Generation Using Dual-Hierarchy Graph Convolutional NetworkabstractIn reality, complex scene plagues numerous scene graph generation models because realistic scene contains myriad of objects and complicated relationships. Most current methods suffer poor performance when encountering complex scenes. We find that there are two principal reasons for this phenomenon. First, the construction of graph loses sight of the hierarchy of objects. Second, there exists redundant information in feature optimization. To facilitate this issue, this paper proposes an innovative dual-hierarchy graph convolutional network (DH-GCN), which is a conceptually elegant and efficient top-down approach. In specific, DH-GCN leverages salient object detector to hierarchize objects and give gist nodes more accurate representation. Moreover, the dual-hierarchy message propagation is designed to refine the representation hierarchically and eliminate redundant information. Systematic experiments on Visual Genome dataset show the superiority of our method over strong baseline methods. Jiale Lu, Lianggangxu Chen, Yiqing Cai, Haoyue Guan, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 3 |
| 2022 | Exploring Contextual Relationships in 3D Cloud Points by Semantic Knowledge MiningabstractAbstract 3D scene graph generation (SGG) aims to predict the class of objects and predicates simultaneously in one 3D point cloud scene with instance segmentation. Since the underlying semantic of 3D point clouds is spatial information, recent ideas of the 3D SGG task usually face difficulties in understanding global contextual semantic relationships and neglect the intrinsic 3D visual structures. To build the global scope of semantic relationships, we first propose two types of Semantic Clue (SC) from entity level and path level, respectively. SC can be extracted from the training set and modeled as the co‐occurrence probability between entities. Then a novel Semantic Clue aware Graph Convolution Network (SC‐GCN) is designed to explicitly model each SC of which the message is passed in their specific neighbor pattern. For constructing the interactions between the 3D visual and semantic modalities, a visual‐language transformer (VLT) module is proposed to jointly learn the correlation between 3D visual features and class label embeddings. Systematic experiments on the 3D semantic scene graph (3DSSG) dataset show that our full method achieves state‐of‐the‐art performance. Lianggangxu Chen, Jiale Lu, Yiqing Cai, Changbo Wang, Gaoqi He |
Comput. Graph. Forum | 3 |
| 2021 | Leveraging Intra-Domain Knowledge to Strengthen Cross-Domain Crowd CountingabstractUnsupervised cross-domain counting research using synthetic datasets becomes imminent when considering the laborious labeling for supervised methods. However, the existing methods only focus on learning domain shared knowledge to narrow the gap between the source domain and target domain (inter-domain gap). Nevertheless, these methods do not consider the enormous distribution gap among the target domain data itself (intra-domain gap). In this paper, we propose a two-step domain adaptation method with multi-level feature response branches, which further uses the intra-domain knowledge to strengthen the target domain’s adaptability. Specifically, we first use different feature response branches to learn inter-domain knowledge more robustly, reducing the prediction inconsistency of different scenarios. Subsequently, the trained model is used to generate pseudo-labels for the target domain. The entire model was retrained by using pseudo-labels. Various experiments on synthetic dataset GCC and three real public datasets validate our proposed method’s availability with higher accuracy. Yiqing Cai, Lianggangxu Chen, Zhenwei Ma, Changhong Lu, Changbo Wang, Gaoqi He |
ICME | 1 |
| 2014 | Cyclic network automata and cohomological waves
Yiqing Cai, Robert Ghrist |
IPSN | 1 |
| 2012 | Cyclic network automata for indoor sensor networkabstractFollowing Baryshnikov-Coffman-Kwak [1], we use network cyclic cellular automata to generate a decentralized protocol, with only a small fraction of sensors awake. The work here shows that waves of awake-state nodes turn corners and automatically solve pusuit/evasion-type problems without centralized coordination. Yiqing Cai, Robert Ghrist |
IPSN | 1 |