VLDB 2026 Research / reviewers in the wild / expert
Yueheng Sun
dblp:83/6984
· DBLP profile ↗
38ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0003-4569-8193ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 22 since 2021Databases, data management, data science and information retrieval · 13 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynSpectral: A Multi-channel Temporal Spectral GNN with Frequency Decomposition for Dynamic Graphs
Runguo Tao, Tianpeng Li, Minglai Shao 0001, Wenjun Wang 0002, Xuan Guo 0005, Yueheng Sun |
DASFAA (2) | 6 |
| 2026 | Beyond semantics: Exploiting propagation structures with dual-adapter LLMs for fake news detection
Aojie Si, Shizhan Chen, Xiaobao Wang, Yueheng Sun, Jiaai Guo |
Neural Networks | 4 |
| 2025 | Combining Loss-aware Curriculum Learning with Incomplete Graph Neural NetworksabstractGraph neural networks (GNNs) have achieved great success in node classification tasks. However, most graph neural networks are incomplete. For example, the reference of each article is subjectively introduced by the author in the citation network, which leads to an incomplete citation network, especially the imbalance of categories. Directly training a GNN classifier with raw data tends to under-represent samples from minority classes, leading to suboptimal performance. This paper presents a novel framework, named CL2IGNN, in which an embedding space is constructed to encode the similarity among the nodes. New samples are synthesized in this space to ensure their genuineness. Additionally, an edge generator is trained simultaneously to model the relational information and provide it to these new samples. In order to improve the accuracy of node classification, we assess the quality of each data node and progressively incorporate the training dataset into the model, increasing the difficulty step by step. Our approach can effectively reduce bias and variance, mitigate the impact of noisy data, and improve overall accuracy. The code is public at https://anonymous.4open.science/r/CL2IGNN-1832. Keao Xi, Gaoke Zhang, Yueheng Sun, Wenjun Wang 0002 |
ICASSP | 4 |
| 2024 | LLM-Driven Multimodal Opinion Expression Identification
Bonian Jia, Huiyao Chen, Yueheng Sun, Meishan Zhang, Min Zhang 0005 |
INTERSPEECH | 3 |
| 2024 | Diffusion Review-Based Recommendation
Xiangfu He, Qiyao Peng 0001, Minglai Shao 0001, Yueheng Sun |
KSEM (5) | 4 |
| 2024 | Automatic Noise Generation and Reduction for Text ClassificationabstractLabel noise is an important issue in machine learning, which might lead to negative influences on various tasks. Given that real benchmarks for evaluation of noise reduction methods are limited, plenty of studies construct pseudo noisy data to verify their proposed methods. However, very few works have realized the rationality of the noise generation strategies. If the generated pseudo datasets are biased, their final conclusions might also be problematic. In this work, we focus on text classification of natural language processing (NLP) to investigate various pseudo noise generation methods, which is the first work of this line for NLP. In particular, we compare the noise generated with crowdsourcing noise, a kind of real noise as gold-standard, to evaluate these noise generation methods. After then, we measure and compare the performance of representative noise reduction methods respectively based on the data of crowdsourcing and our top-ranked pseudo noisy generation strategies. We conduct experiments on five text classification datasets, offering detailed comparison results as well as discussions. Huiyao Chen, Yueheng Sun, Meishan Zhang, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Group-Based Personalized News Recommendation with Long- and Short-Term Fine-Grained MatchingabstractPersonalized news recommendation aims to help users find news content they prefer, which has attracted increasing attention recently. There are two core issues in news recommendation: learning news representation and matching candidate news with user interests. In this context, “candidate” indicates potential for interest. Due to the superior ability to understand natural language demonstrated by Pretrained Language Models (PLMs), recent works utilize PLMs (e.g., BERT) to strengthen news modeling, obtaining more accurate user interest matching and achieving notable improvement in news recommendation. However, the existing PLM-based methods are usually incapable of fully exploring the fine-grained (i.e., word-level) relatedness between user behaviors and candidate news due to the heavy computational cost brought by PLMs. In this article, we propose a group-based personalized news recommendation method with long- and short-term matching mechanisms between users and candidate news based on PLMs to learn fine-grained matching efficiently and effectively. In our approach, we design to group user historical clicked news into chunks with quite shorter news sequences according to their clicked timestamps, which could alleviate the computation issues of PLMs. PLMs are applied in each group jointly with the candidate news to capture their word-level interaction, and global group-level matching is learned across different groups. In addition, the group-based mechanism could be naturally adapted for long- and short-term user representation learning, in which we build users’ long preferences from the representations of all groups and treat the last group as short interests, respectively. Finally, we employ a gate network to dynamically unify the group-level, long- and short-term representations, yielding comprehensive user-news matching effectively. Extensive experiments are conducted on two real-world datasets. The results show that our proposed method achieves superior performance in news recommendations. Hongyan Xu 0001, Qiyao Peng 0001, Hongtao Liu 0008, Yueheng Sun, Wenjun Wang 0002 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | A Dual-branch Learning Model with Gradient-balanced Loss for Long-tailed Multi-label Text ClassificationabstractMulti-label text classification has a wide range of applications in the real world. However, the data distribution in the real world is often imbalanced, which leads to serious long-tailed problems. For multi-label classification, due to the vast scale of datasets and existence of label co-occurrence, how to effectively improve the prediction accuracy of tail labels without degrading the overall precision becomes an important challenge. To address this issue, we propose A Dual-Branch Learning Model with Gradient-Balanced Loss (DBGB) based on the paradigm of existing pre-trained multi-label classification SOTA models. Our model consists of two main long-tailed module improvements. First, with the shared text representation, the dual-classifier is leveraged to process two kinds of label distributions; one is the original data distribution and the other is the under-sampling distribution for head labels to strengthen the prediction for tail labels. Second, the proposed gradient-balanced loss can adaptively suppress the negative gradient accumulation problem related to labels, especially tail labels. We perform extensive experiments on three multi-label text classification datasets. The results show that the proposed method achieves competitive performance on overall prediction results compared to the state-of-the-art methods in solving the multi-label classification, with significant improvement on tail-label accuracy. Yitong Yao, Peng Zhang 0002, Yueheng Sun |
ACM Trans. Inf. Syst. | 4 |
| 2023 | Multi-Order Relations Hyperbolic Fusion for Heterogeneous GraphsabstractHeterogeneous graphs with multiple node and edge types are prevalent in real-world scenarios. However, most methods use meta-paths on the original graph structure to learn information in heterogeneous graphs, and these methods only consider pairwise relations and rely on meta-paths. In this paper, we use simplicial complexes to extract higher-order relations containing multiple nodes from heterogeneous graphs. We also discover power-law structures in both the heterogeneous graph and the extracted simplicial complex. Thus, we propose the Simplicial Hyperbolic Attention Network (SHAN), a graph neural network for heterogeneous graphs. SHAN extracts simplicial complexes and the original graph structure from the heterogeneous graph to represent multi-order relations between nodes. Next, SHAN uses hyperbolic multi-perspective attention to learn the importance of different neighbors and relations in hyperbolic space. Finally, SHAN integrates multi-order relations to obtain a more comprehensive node representation. We conducted extensive experiments to verify the effectiveness of SHAN and the results of node classification experiments on three publicly available heterogeneous graph datasets demonstrate that SHAN outperforms representative baseline models. Yueheng Sun, Minglai Shao 0001 |
CIKM | 2 |
| 2023 | Robust Few-Shot Graph Anomaly Detection via Graph Coarsening
Yueheng Sun, Tianpeng Li, Minglai Shao 0001 |
KSEM (1) | 2 |
| 2023 | Joint Community and Structural Hole Spanner Detection via Graph Contrastive Learning
Wenjun Wang 0002, Tianpeng Li, Minglai Shao 0001, Jiye Liu, Yueheng Sun |
KSEM (4) | 6 |
| 2023 | Learning graph deep autoencoder for anomaly detection in multi-attributed networks
Minglai Shao 0001, Qiyao Peng 0001, Jun Zhao 0017, Zhan Pei, Yueheng Sun |
Knowl. Based Syst. | 6 |
| 2023 | Heterogeneous network representation learning based on role feature extraction
Yueheng Sun, Mengyu Jia, Minglai Shao 0001 |
Pattern Recognit. | 1 |
| 2023 | Fine-Grained Domain Adaptation for Chinese Syntactic ProcessingabstractSyntactic processing is fundamental to natural language processing. It provides rich and comprehensive syntax information in sentences that could be potentially beneficial for downstream tasks. Recently, pretrained language models have shown great success in Chinese syntactic processing, which typically involves word segmentation, POS tagging, and dependency parsing. However, the on-going research never ends since performance would be degraded drastically when tested on a highly-discrepant domain. This problem is widely accepted as domain adaptation, where the test domain differs from the training domain in supervised learning. Self-training is one promising solution for it, and straightforward source-to-target adaptation has already shown remarkable effectiveness in previous work. While this strategy ignores the fact that sentences of the target domain sentences may have very different gaps from the source training domain. More specifically, sentences with large gaps might fail by direct self-training adaptation. To this end, we propose fine-grained domain adaptation for Chinese syntactic processing in this work, aiming to model the gaps between the source and the target domains accurately and progressively. The key idea is to divide the target domain into fine-grained subdomains by using a specified domain distance metric, and then perform gradual self-training on the subdomains. We further offer an intuitive theoretical illustration based on the theory of Kumar et al. (2020) approximately. In addition, a novel representation learning framework is proposed to encode fine-grained subdomains effectively, aiming to utilize the above idea fully. Experimental results on benchmark datasets show that our method can achieve significant improvements over a variety of baselines. Meishan Zhang, Peiming Guo, Peijie Jiang, Dingkun Long, Yueheng Sun, Pengjun Xie, Min Zhang 0005 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Enhancing RDF Verbalization with Descriptive and Relational KnowledgeabstractRDF verbalization has received increasing interest, which aims to generate a natural language description of the knowledge base. Sequence-to-sequence models based on Transformer are able to obtain strong performance equipped with pre-trained language models such as BART and T5. However, in spite of the general performance gain introduced by the pre-trained models, the performance of the task is still limited by the small scale of the training dataset. To address the problem, we propose two orthogonal strategies to enhance the representation learning of RDF triples. Concretely, two types of knowledge are introduced, i.e., descriptive knowledge and relational knowledge, respectively. The descriptive knowledge indicates the semantic information of self definition, and the relational knowledge indicates the semantic information learned from the structural context. We further combine the descriptive and relational knowledge together to enhance the representation learning. Experimental results on the WebNLG and SemEval-2010 datasets show that the two types of knowledge can both enhance the model performance, and their combination is able to obtain further improvements in most cases, providing new state-of-the-art results. Meishan Zhang, Shuang Liu 0007, Yueheng Sun, Nan Duan 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency TransferabstractUnsupervised cross-lingual transfer has been shown great potentials for dependency parsing of the low-resource languages when there is no annotated treebank available. Recently, the self-training method has received increasing interests because of its state-of-the-art performance in this scenario. In this work, we advance the method further by coupling it with curriculum learning, which guides the self-training in an easy-to-hard manner. Concretely, we present a novel metric to measure the instance difficulty of a dependency parser which is trained mainly on a Treebank from a resource-rich source language. By using the metric, we divide a low-resource target language into several fine-grained sub-languages by their difficulties, and then apply iterative-self-training progressively on these sub-languages. To fully explore the auto-parsed training corpus from sub-languages, we exploit an improved parameter generation network to model the sub-languages for better representation learning. Experimental results show that our final curriculum-style self-training can outperform a range of strong baselines, leading to new state-of-the-art results on unsupervised cross-lingual dependency parsing. We also conduct detailed experimental analyses to examine the proposed approach in depth for comprehensive understandings. Peiming Guo, Shen Huang, Peijie Jiang, Yueheng Sun, Meishan Zhang, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing AnnotationsabstractRecent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy.Crowdsourcing is one practical solution for this problem, aiming to create a large-scale but quality-unguaranteed corpus.In this work, we investigate Chinese OEI with extremelynoisy crowdsourcing annotations, constructing a dataset at a very low cost.Following Zhang et al. (2021), we train the annotator-adapter model by regarding all annotations as goldstandard in terms of crowd annotators, and test the model by using a synthetic expert, which is a mixture of all annotators.As this annotatormixture for testing is never modeled explicitly in the training phase, we propose to generate synthetic training samples by a pertinent mixup strategy to make the training and testing highly consistent.The simulation experiments on our constructed dataset show that crowdsourcing is highly promising for OEI, and our proposed annotator-mixup can further enhance the crowdsourcing modeling. Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Xiaobin Wang, Min Zhang 0005 |
ACL (1) | 3 |
| 2022 | Role-Oriented Dynamic Network EmbeddingabstractExploring the differences and important patterns of nodes from the perspective of roles has gradually developed into an interesting and important topic in network analysis. However, existing role-oriented network embedding methods focus more on identifying underlying roles for static network, which leads to complex temporal behaviors being overlooked and degraded performance facing dynamic network. The few role analytics methods for dynamic networks either cannot learn general node representations or fail to discovery role transitions of nodes. In this work, we propose a unified framework RDNE (Role-oriented Dynamic Network Embedding) to tackle such challenges, which aim to learn multiple embeddings for individual nodes based on time-varying structural behaviors. Based on regular equivalence, RDNE propagates the structural features over the graph to derive the initial role-oriented representations. Then, it applies capsule network to further model the mapping between nodes and roles, which is the first time capsule network is used for role discovery. For the varying and temporal dependence within dynamic network, we utilize the Gated Recurrent Unit to compute historical information and use historical information to influence the generation of representations at the next snapshot. Comprehensive experiments on both synthetic and real-world networks validate the superiority of the proposed RDNE. Wenjun Wang 0002, Minglai Shao 0001, Yueheng Sun, Pengfei Jiao |
IEEE Big Data | 4 |
| 2022 | Domain-Specific NER via Retrieving Correlated SamplesabstractSuccessful Machine Learning based Named Entity Recognition models could fail on texts from some special domains, for instance, Chinese addresses and e-commerce titles, where requires adequate background knowledge. Such texts are also difficult for human annotators. In fact, we can obtain some potentially helpful information from correlated texts, which have some common entities, to help the text understanding. Then, one can easily reason out the correct answer by referencing correlated samples. In this paper, we suggest enhancing NER models with correlated samples. We draw correlated samples by the sparse BM25 retriever from large-scale in-domain unlabeled data. To explicitly simulate the human reasoning process, we perform a training-free entity type calibrating by majority voting. To capture correlation features in the training stage, we suggest to model correlated samples by the transformer-based multi-instance cross-encoder. Empirical results on datasets of the above two domains show the efficacy of our methods. Xin Zhang 0097, Yong Jiang 0005, Xiaobin Wang, Xuming Hu, Yueheng Sun, Pengjun Xie, Meishan Zhang |
COLING | 5 |
| 2022 | Towards Personalized Review Generation with Gated Multi-source Fusion Network
Hongtao Liu 0008, Wenjun Wang 0002, Hongyan Xu 0001, Qiyao Peng 0001, Pengfei Jiao, Yueheng Sun |
DASFAA (3) | 6 |
| 2022 | Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text GenerationabstractImage-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades.Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new perspective for image-to-text toward spatial semantics.Given an image and two objects inside it, VSD aims to produce one description focusing on the spatial perspective between the two objects.Accordingly, we manually annotate a dataset to facilitate the investigation of the newly-introduced task and build several benchmark encoder-decoder models by using VL-BART and VL-T5 as backbones.In addition, we investigate pipeline and joint end-to-end architectures for incorporating visual spatial relationship classification (VSRC) information into our model.Finally, we conduct experiments on our benchmark dataset to evaluate all our models.Results show that our models are impressive, providing accurate and human-like spatial-oriented text descriptions.Meanwhile, VSRC has great potential for VSD, and the joint end-to-end architecture is the better choice for their integration.We make the dataset and codes public for research purposes. Yu Zhao 0043, Jianguo Wei, Zhichao Lin, Yueheng Sun, Meishan Zhang, Min Zhang 0005 |
EMNLP | 4 |
| 2022 | Geometry interaction network alignment
Yinghui Wang 0005, Wenjun Wang 0002, Zixu Zhen, Qiyao Peng 0001, Pengfei Jiao, Minglai Shao 0001, Yueheng Sun |
Neurocomputing | 8 |
| 2022 | Multi-users interaction anomalous subgraph detection for event mining
Yang Yu 0030, Wenjun Wang 0002, Minglai Shao 0001, Ying Sun 0005, Yueheng Sun, Qiang Tian |
Neurocomputing | 6 |
| 2022 | Fake news detection via knowledgeable prompt learning
Gongyao Jiang, Shuang Liu 0007, Yu Zhao 0043, Yueheng Sun, Meishan Zhang |
Inf. Process. Manag. | 4 |
| 2022 | STHGCN: A spatiotemporal prediction framework based on higher-order graph convolution networks
Jun Wang 0193, Wenjun Wang 0002, Wei Yu 0016, Keyong Jia, Xiaoming Li 0006, Yueheng Sun, Yuqing Xu |
Knowl. Based Syst. | 8 |
| 2021 | Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionabstractXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Pengjun Xie. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Pengjun Xie |
ACL/IJCNLP (1) | 3 |
| 2021 | A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS TaggingabstractDomain adaption for word segmentation and POS tagging is a challenging problem for Chinese lexical processing.Self-training is one promising solution for it, which struggles to construct a set of high-quality pseudo training instances for the target domain.Previous work usually assumes a universal sourceto-target adaption to collect such pseudo corpus, ignoring the different gaps from the target sentences to the source domain.In this work, we start from joint word segmentation and POS tagging, presenting a fine-grained domain adaption method to model the gaps accurately.We measure the gaps by one simple and intuitive metric, and adopt it to develop a pseudo target domain corpus based on finegrained subdomains incrementally.A novel domain-mixed representation learning model is proposed accordingly to encode the multiple subdomains effectively.The whole process is performed progressively for both corpus construction and model training.Experimental results on a benchmark dataset show that our method can gain significant improvements over a vary of baselines.Extensive analyses are performed to show the advantages of our final domain adaption model as well. Peijie Jiang, Dingkun Long, Yueheng Sun, Meishan Zhang, Pengjun Xie |
EMNLP (1) | 3 |
| 2021 | A Graph-Based Neural Model for End-to-End Frame Semantic ParsingabstractFrame semantic parsing is a semantic analysis task based on FrameNet which has received great attention recently.The task usually involves three subtasks sequentially: (1) target identification, (2) frame classification and (3) semantic role labeling.The three subtasks are closely related while previous studies model them individually, which ignores their intern connections and meanwhile induces error propagation problem.In this work, we propose an end-to-end neural model to tackle the task jointly.Concretely, we exploit a graphbased method, regarding frame semantic parsing as a graph construction problem.All predicates and roles are treated as graph nodes, and their relations are taken as graph edges.Experiment results on two benchmark datasets of frame semantic parsing show that our method is highly competitive, resulting in better performance than pipeline models. Zhichao Lin, Yueheng Sun, Meishan Zhang |
EMNLP (1) | 2 |
| 2021 | The mass, fake news, and cognition security
Bin Guo 0001, Yasan Ding, Yueheng Sun, Shuai Ma 0001, Ke Li 0045, Zhiwen Yu 0001 |
Frontiers Comput. Sci. | 3 |
| 2020 | CPNSA: Cascade Prediction with Network Structure Attention
Chaochao Liu, Wenjun Wang 0002, Pengfei Jiao, Yueheng Sun, Xiaoming Li 0006, Xue Chen 0005 |
CollaborateCom (1) | 4 |
| 2020 | Cascade modeling with multihead self-attentionabstractModeling how information diffuses across social network platforms can be widely used. Recently, researchers have used deep learning methods to model information cascades and forecast their progression without dependence on the hypothesis of the underlying diffusion model. Most of these studies use sequential models (e.g., recurrent neural networks, RNNs) and model cascades of information spread without using the network structure information. However, the network structure information substantially affects information spread, cross-dependence should be considered in cascade modeling, and recurrent neural networks produce poor results on long sequence modeling. To solve these issues, in this paper, we present a new cascade modeling method with the multihead self-attention mechanism. We design an encoder that combines network structure information with multihead self-attention to learn the representations of cascades and consider diverse user dependencies on the network. Experiments are conducted on both synthetic and real-world datasets. The results show that the proposed method performs better on the long sequence cascade prediction than the state-of-the-art methods. Chaochao Liu, Wenjun Wang 0002, Pengfei Jiao, Xue Chen 0005, Yueheng Sun |
IJCNN | 5 |
| 2019 | Surrounding-Based Attention Networks for Aspect-Level Sentiment Classification
Yueheng Sun, Xianchen Wang, Hongtao Liu 0008, Wenjun Wang 0002, Pengfei Jiao |
ICANN (4) | 1 |
| 2019 | NRSA: Neural Recommendation with Summary-Aware Attention
Qiyao Peng 0001, Peiyi Wang, Wenjun Wang 0002, Hongtao Liu 0008, Yueheng Sun, Pengfei Jiao |
KSEM (1) | 5 |
| 2019 | REET: Joint Relation Extraction and Entity Typing via Multi-task Learning
Hongtao Liu 0008, Peiyi Wang, Fangzhao Wu, Pengfei Jiao, Wenjun Wang 0002, Xing Xie 0001, Yueheng Sun |
NLPCC (1) | 7 |
| 2019 | Community structure enhanced cascade prediction
Chaochao Liu, Wenjun Wang 0002, Yueheng Sun |
Neurocomputing | 3 |
| 2018 | NE-FLGC: Network Embedding Based on Fusing Local (First-Order) and Global (Second-Order) Network Structure with Node Content
Hongyan Xu 0001, Hongtao Liu 0008, Wenjun Wang 0002, Yueheng Sun, Pengfei Jiao |
PAKDD (2) | 4 |
| 2018 | Exploring temporal community structure and constant evolutionary pattern hiding in dynamic networks
Pengfei Jiao, Wei Yu 0016, Wenjun Wang 0002, Xiaoming Li 0006, Yueheng Sun |
Neurocomputing | 5 |
| 2008 | Finding question-answer pairs from online forumsabstractOnline forums contain a huge amount of valuable user generated content. In this paper we address the problem of extracting question-answer pairs from forums. Question-answer pairs extracted from forums can be used to help Question Answering services (e.g. Yahoo! Answers) among other applications. We propose a sequential patterns based classification method to detect questions in a forum thread, and a graph based propagation method to detect answers for questions in the same thread. Experimental results show that our techniques are very promising. Gao Cong, Chin-Yew Lin, Young-In Song, Yueheng Sun |
SIGIR | 5 |