VLDB 2026 Research / reviewers in the wild / expert
Jinhua Gao
dblp:119/6596
· DBLP profile ↗
22ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 8 since 2021Databases, data management, data science and information retrieval · 11 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsabstractShiyao Cui, QingLin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shiyao Cui, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Han Qiu 0001, Minlie Huang |
ACL (1) | 6 |
| 2025 | Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language ModelsabstractRetrieval-Augmented Language Models boost task performance, owing to the retriever that provides external knowledge.Although crucial, the retriever primarily focuses on semantics relevance, which may not always be effective for generation.Thus, utility-based retrieval has emerged as a promising topic, prioritizing passages that provide valid benefits for downstream tasks.However, due to insufficient understanding, capturing passage utility accurately remains unexplored.This work proposes SCARLet, a framework for training utility-based retrievers in RALMs, which incorporates two key factors, multi-task generalization and inter-passage interaction.First, SCAR-Let constructs shared context on which training data for various tasks is synthesized.This mitigates semantic bias from context differences, allowing retrievers to focus on learning task-specific utility and generalize across tasks.Next, SCARLet uses a perturbation-based attribution method to estimate passage-level utility for shared context, which reflects interactions between passages and provides more accurate feedback.We evaluate our approach on ten datasets across various tasks, both indomain and out-of-domain, showing that retrievers trained by SCARLet consistently improve the overall performance of RALMs. Yilong Xu, Jinhua Gao, Xiaoming Yu, Yuanhai Xue, Baolong Bi, Huawei Shen, Xueqi Cheng 0001 |
EMNLP | 2 |
| 2025 | ALiiCE: Evaluating Positional Fine-grained Citation GenerationabstractYilong Xu, Jinhua Gao, Xiaoming Yu, Baolong Bi, Huawei Shen, Xueqi Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yilong Xu, Jinhua Gao, Xiaoming Yu, Baolong Bi, Huawei Shen, Xueqi Cheng 0001 |
NAACL (Long Papers) | 2 |
| 2025 | A Study on Simulating Directional Land Surface Emissivity Based on Kernel-Driven Models and Its Application to the Generalized Split-Window AlgorithmabstractIn radiometric measurements, the emissivity of natural objects exhibits a dependence on the viewing angle. Ignoring the angular effect of surface emissivity can increase the uncertainty of land surface temperature (LST) retrievals. To mitigate this issue, we evaluated the simulation performance of 11 parametric kernel-driven models (KDMs) and developed directional emissivity models using MYD21 and MYD03 products. Afterward, the directional and classification-based emissivities were input into the refined GSW algorithm to retrieve LSTs with and without considering angular effects (LST_GSW_DE and LST_GSW_CE, respectively). Coupled with the MYD21 LST product (LST_TES), three LSTs were evaluated via SURFRAD in situ data and ERA5-Land products. The main findings were as follows: (1) The RMSEs of directional emissivity simulated by different KDMs ranged from ˜0.0003 to ˜0.001, and their performance differences were generally slight, indicating that parameterized KDMs demonstrate reliable simulation performance in satellite-based directional emissivity modeling. (2) The directional emissivity simulation performances of different KDMs were ranked as follows: dual-kernel model (with both hotspot and base shape kernels) ≥ multikernel model > single-kernel model. The USEA and GUTA-sparse models exhibited advantages over the other KDMs when simulating impervious surfaces during the daytime. (3) We evaluated the three types of retrieved LSTs via SURFRAD in situ data. The rankings of the RMSE and MBE values were consistent: LST_TES was optimal, followed by LST_GSW_DE and LST_GSW_CE, with average RMSEs of 2.47 K, 2.62 K, and 2.80 K, respectively. Furthermore, we evaluated the three types of retrieved LSTs against the ERA5-Land data, and the rankings of the RMSE and MBE values were also consistent: LST_TES was comparable to (slightly better than) LST_GSW_DE in some seasons and consistently better than LST_GSW_CE. The average RMSEs were 2.45 K, 2.52 K, and 2.60 K. In addition, the RMSE and MBE values at different VZAs for the three LSTs increased with increasing VZA, especially when the VZA was greater than 40°. The results demonstrated that it is feasible to use KDMs to simulate directional emissivity from satellite data, offering theoretical interpretability and addressing the issues of discrete and missing emissivity data. Future studies could be devoted to establishing new KDMs or kernels that conform to different land surface and solar illumination conditions to improve the LST retrieval accuracy. Hao Sun 0003, Dandan Wang 0003, Zhiwei He 0004, Bo-Hui Tang, Zhenheng Xu, Jinhua Gao, Tian Zhang 0025, Huanyu Xu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Towards Continual Knowledge Graph Embedding via Incremental DistillationabstractTraditional knowledge graph embedding (KGE) methods typically require preserving the entire knowledge graph (KG) with significant training costs when new knowledge emerges. To address this issue, the continual knowledge graph embedding (CKGE) task has been proposed to train the KGE model by learning emerging knowledge efficiently while simultaneously preserving decent old knowledge. However, the explicit graph structure in KGs, which is critical for the above goal, has been heavily ignored by existing CKGE methods. On the one hand, existing methods usually learn new triples in a random order, destroying the inner structure of new KGs. On the other hand, old triples are preserved with equal priority, failing to alleviate catastrophic forgetting effectively. In this paper, we propose a competitive method for CKGE based on incremental distillation (IncDE), which considers the full use of the explicit graph structure in KGs. First, to optimize the learning order, we introduce a hierarchical strategy, ranking new triples for layer-by-layer learning. By employing the inter- and intra-hierarchical orders together, new triples are grouped into layers based on the graph structure features. Secondly, to preserve the old knowledge effectively, we devise a novel incremental distillation mechanism, which facilitates the seamless transfer of entity representations from the previous layer to the next one, promoting old knowledge preservation. Finally, we adopt a two-stage training paradigm to avoid the over-corruption of old knowledge influenced by under-trained new knowledge. Experimental results demonstrate the superiority of IncDE over state-of-the-art baselines. Notably, the incremental distillation mechanism contributes to improvements of 0.2%-6.5% in the mean reciprocal rank (MRR) score. More exploratory experiments validate the effectiveness of IncDE in proficiently learning new knowledge while preserving old knowledge across all time steps. Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Ziyu Shang, Jinhua Gao, Ke Ji, Yanhe Liu |
AAAI | 5 |
| 2024 | Fast and Continual Knowledge Graph Embedding via Incremental LoRA
Jiajun Liu 0005, Wenjun Ke 0002, Peng Wang 0004, Jinhua Gao, Ziyu Shang, Zijie Xu 0003, Ke Ji |
IJCAI | 5 |
| 2024 | CPMF: An Integrated Technology for Generating 30-m, All-Weather Land Surface Temperature by Coupling Physical Model, Machine Learning, and Spatiotemporal Fusion Model
Jinhua Gao, Hao Sun 0003, Zhenheng Xu, Tian Zhang 0025, Huanyu Xu, Xiang Zhao 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Towards Incremental NER Data Augmentation via Syntactic-aware Insertion TransformerabstractNamed entity recognition (NER) aims to locate and classify named entities in natural language texts. Most existing high-performance NER models employ a supervised paradigm, which requires a large quantity of high-quality annotated data during training. In order to help NER models perform well in few-shot scenarios, data augmentation approaches attempt to build extra data by means of random editing or by using end-to-end generation with PLMs. However, these methods focus on only the fluency of generated sentences, ignoring the syntactic correlation between the new and raw sentences. Such uncorrelation also brings low diversity and inconsistent labeling of synthetic samples. To fill this gap, we present SAINT (Syntactic-Aware InsertioN Transformer), a hard-constraint controlled text generation model that incorporates syntactic information. The proposed method operates by inserting new tokens between existing entities in a parallel manner. During insertion procedure, new tokens will be added taking both semantic and syntactic factors into account. Hence the resulting sentence can retain the syntactic correctness with respect to the raw data. Experimental results on two benchmark datasets, i.e., Ontonotes and Wikiann, demonstrate the comparable performance of SAINT over the state-of-the-art baselines. Wenjun Ke 0002, Zongkai Tian, Qi Liu 0056, Peng Wang 0004, Jinhua Gao |
IJCAI | 5 |
| 2023 | Zero-shot stance detection via multi-perspective contrastive learning with unlabeled data
Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
Inf. Process. Manag. | 2 |
| 2022 | Few-Shot Stance Detection via Target-Aware Prompt DistillationabstractStance detection aims to identify whether the author of a text is in favor of, against, or neutral to a given target. The main challenge of this task comes two-fold: few-shot learning resulting from the varying targets and the lack of contextual information of the targets. Existing works mainly focus on solving the second issue by designing attention-based models or introducing noisy external knowledge, while the first issue remains under-explored. In this paper, inspired by the potential capability of pre-trained language models (PLMs) serving as knowledge bases and few-shot learners, we propose to introduce prompt-based fine-tuning for stance detection. PLMs can provide essential contextual information for the targets and enable few-shot learning via prompts. Considering the crucial role of the target in stance detection task, we design target-aware prompts and propose a novel verbalizer. Instead of mapping each label to a concrete word, our verbalizer maps each label to a vector and picks the label that best captures the correlation between the stance and the target. Moreover, to alleviate the possible defect of dealing with varying targets with a single hand-crafted prompt, we propose to distill the information learned from multiple prompts. Experimental results show the superior performance of our proposed model in both full-data and few-shot scenarios. Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
SIGIR | 2 |
| 2022 | ConsistSum: Unsupervised Opinion Summarization with the Consistency of Aspect, Sentiment and SemanticabstractUnsupervised opinion summarization techniques are designed to condense the review data and summarize informative and salient opinions in the absence of golden references. Existing dominant methods generally follow a two-stage framework: first creating the synthetic "review-summary" paired datasets and then feeding them into the generative summary model for supervised training. However, these methods mainly focus on semantic similarity in synthetic dataset creation, ignoring the consistency of aspects and sentiments in synthetic pairs. Such inconsistency also brings a gap to the training and inference of the summarization model. Wenjun Ke 0002, Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
WSDM | 2 |
| 2021 | Semantic-Syntax Cascade Injection Model for Aspect Sentiment Triple Extraction
Wenjun Ke 0002, Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
PAKDD (2) | 2 |
| 2021 | Capturing SQL Query Overlapping via Subtree Copy for Cross-Domain Context-Dependent SQL Generation
Ruizhuo Zhao, Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
PAKDD (2) | 2 |
| 2021 | Incorporating explicit syntactic dependency for aspect level sentiment classification
Wenjun Ke 0002, Jinhua Gao, Huawei Shen, Xueqi Cheng 0001 |
Neurocomputing | 2 |
| 2021 | Learning diffusion model-free and efficient influence function for influence maximization from information cascades
Qi Cao 0005, Huawei Shen, Jinhua Gao, Xueqi Cheng 0001 |
Knowl. Inf. Syst. | 3 |
| 2020 | Label-Consistency based Graph Neural Networks for Semi-supervised Node ClassificationabstractGraph neural networks (GNNs) achieve remarkable success in graph-based semi-supervised node classification, leveraging the information from neighboring nodes to improve the representation learning of target node. The success of GNNs at node classification depends on the assumption that connected nodes tend to have the same label. However, such an assumption does not always work, limiting the performance of GNNs at node classification. In this paper, we propose label-consistency based graph neural network (LC-GNN), leveraging node pairs unconnected but with the same labels to enlarge the receptive field of nodes in GNNs. Experiments on benchmark datasets demonstrate the proposed LC-GNN outperforms traditional GNNs in graph-based semi-supervised node classification. We further show the superiority of LC-GNN in sparse scenarios with only a handful of labeled nodes. Bingbing Xu 0001, Huawei Shen, Jinhua Gao, Xueqi Cheng 0001 |
SIGIR | 5 |
| 2020 | Popularity Prediction on Social Platforms with Coupled Graph Neural NetworksabstractPredicting the popularity of online content on social platforms is an important task for both researchers and practitioners. Previous methods mainly leverage demographics, temporal and structural patterns of early adopters for popularity prediction. However, most existing methods are less effective to precisely capture the cascading effect in information diffusion, in which early adopters try to activate potential users along the underlying network. In this paper, we consider the problem of network-aware popularity prediction, leveraging both early adopters and social networks for popularity prediction. We propose to capture the cascading effect explicitly, modeling the activation state of a target user given the activation state and influence of his/her neighbors. To achieve this goal, we propose a novel method, namely CoupledGNN, which uses two coupled graph neural networks to capture the interplay between node activation states and the spread of influence. By stacking graph neural network layers, our proposed method naturally captures the cascading effect along the network in a successive manner. Experiments conducted on both synthetic and real-world Sina Weibo datasets demonstrate that our method significantly outperforms the state-of-the-art methods for popularity prediction. Qi Cao 0005, Huawei Shen, Jinhua Gao, Bingzheng Wei, Xueqi Cheng 0001 |
WSDM | 3 |
| 2019 | Learning Binary Hash Codes for Fast Anchor Link Retrieval across NetworksabstractUsers are usually involved in multiple social networks, without explicit anchor links that reveal the correspondence among different accounts of the same user across networks. Anchor link prediction aims to identify the hidden anchor links, which is a fundamental problem for user profiling, information cascading, and cross-domain recommendation. Although existing methods perform well in the accuracy of anchor link prediction, the pairwise search manners on inferring anchor links suffer from big challenge when being deployed in practical systems. To combat the challenges, in this paper we propose a novel embedding and matching architecture to directly learn binary hash code for each node. Hash codes offer us an efficient index to filter out the candidate node pairs for anchor link prediction. Extensive experiments on synthetic and real world large-scale datasets demonstrate that our proposed method has high time efficiency without loss of competitive prediction accuracy in anchor link prediction. Yongqing Wang 0005, Huawei Shen, Jinhua Gao, Xueqi Cheng 0001 |
WWW | 3 |
| 2018 | Towards Efficient Detection of Overlapping Communities in Massive NetworksabstractCommunity detection is essential to analyzing and exploring natural networks such as social networks, biological networks, and citation networks. However, few methods could be used as off-the-shelf tools to detect communities in real world networks for two reasons. On the one hand, most existing methods for community detection cannot handle massive networks that contain millions or even hundreds of millions of nodes. On the other hand, communities in real world networks are generally highly overlapped, requiring that community detection method could capture the mixed community membership. In this paper, we aim to offer an off-the-shelf method to detect overlapping communities in massive real world networks. For this purpose, we take the widely-used Poisson model for overlapping community detection as starting point and design two speedup strategies to achieve high efficiency. Extensive tests on synthetic and large scale real networks demonstrate that the proposed strategies speedup the community detection method based on Poisson model by 1 to 2 orders of magnitudes, while achieving comparable accuracy at community detection. Bing-Jie Sun, Huawei Shen, Jinhua Gao, Wentao Ouyang, Xueqi Cheng 0001 |
AAAI | 3 |
| 2017 | A Non-negative Symmetric Encoder-Decoder Approach for Community DetectionabstractCommunity detection or graph clustering is crucial to understanding the structure of complex networks and extracting relevant knowledge from networked data. Latent factor model, e.g., non-negative matrix factorization and mixed membership block model, is one of the most successful methods for community detection. Latent factor models for community detection aim to find a distributed and generally low-dimensional representation, or coding, that captures the structural regularity of network and reflects the community membership of nodes. Existing latent factor models are mainly based on reconstructing a network from the representation of its nodes, namely network decoder, while constraining the representation to have certain desirable properties. These methods, however, lack an encoder that transforms nodes into their representation. Consequently, they fail to give a clear explanation about the meaning of a community and suffer from undesired computational problems. In this paper, we propose a non-negative symmetric encoder-decoder approach for community detection. By explicitly integrating a decoder and an encoder into a unified loss function, the proposed approach achieves better performance over state-of-the-art latent factor models for community detection task. Moreover, different from existing methods that explicitly impose the sparsity constraint on the representation of nodes, the proposed approach implicitly achieves the sparsity of node representation through its symmetric and non-negative properties, making the optimization much easier than competing methods based on sparse matrix factorization. Bing-Jie Sun, Huawei Shen, Jinhua Gao, Wentao Ouyang, Xueqi Cheng 0001 |
CIKM | 3 |
| 2017 | Cascade Dynamics Modeling with Attention-based Recurrent Neural NetworkabstractAn ability of modeling and predicting the cascades of resharing is crucial to understanding information propagation and to launching campaign of viral marketing. Conventional methods for cascade prediction heavily depend on the hypothesis of diffusion models, e.g., independent cascade model and linear threshold model. Recently, researchers attempt to circumvent the problem of cascade prediction using sequential models (e.g., recurrent neural network, namely RNN) that do not require knowing the underlying diffusion model. Existing sequential models employ a chain structure to capture the memory effect. However, for cascade prediction, each cascade generally corresponds to a diffusion tree, causing cross-dependence in cascade---one sharing behavior could be triggered by its non-immediate predecessor in the memory chain. In this paper, we propose to an attention-based RNN to capture the cross-dependence in cascade. Furthermore, we introduce a \emph{coverage} strategy to combat the misallocation of attention caused by the memoryless of traditional attention mechanism. Extensive experiments on both synthetic and real world datasets demonstrate the proposed models outperform state-of-the-art models at both cascade prediction and inferring diffusion tree. Yongqing Wang 0005, Huawei Shen, Shenghua Liu, Jinhua Gao, Xueqi Cheng 0001 |
IJCAI | 4 |
| 2017 | Marked Temporal Dynamics Modeling Based on Recurrent Neural Network
Yongqing Wang 0005, Shenghua Liu, Huawei Shen, Jinhua Gao, Xueqi Cheng 0001 |
PAKDD (1) | 4 |