Yongjun Li 0006

dblp:35/4145-6 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-3184-6805ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 DSAD: A dual-stream aligned detection framework for LLM-generated exam answers
Yongjun Li 0006, Tao You, Linxin Liu
Expert Syst. Appl.2
2026 R2IS: Resilient and robust neighborhood rough feature selection using combination mutual information
Gengsen Li, Binbin Sang, Shaoguo Cui, Yongjun Li 0006
Neurocomputing5
2026 A Multiview Trajectory Representation Learning Method for User Identification
abstract
With the development of location-based social networks, more and more users post their check-ins on social sites. Matching accounts on different social sites based on users’ spatiotemporal trajectories has become a hot issue in current research. Existing methods do not take into account the temporal characteristics, spatial characteristics, and sequential characteristics of the user’s trajectory at the same time, resulting in the inability to obtain accurate trajectory features. To solve this problem, we propose a new user identification approach based on the spatiotemporal characteristics of user trajectories, which consists of three parts: 1) convert trajectories into graph structures that preserve both spatial connectivity and access order, then use a graph attention network (GAT) to extract spatial representations—enabling effective modeling of users’ mobility structural patterns; 2) design a time-slice enhanced long short-term memory (LSTM) model to extract temporal periodicity and sequential features, which is robust to sparse trajectories with uneven time intervals; and 3) adopt a contrastive learning framework with a modified SimCLR to mine hard negative samples, solving the problem of positive–negative sample imbalance and enhancing the discriminability of trajectory representations. Extensive experiments on real datasets demonstrate that the proposed approach is superior to baseline methods in terms of performance and efficiency.
Jiazheng Fu, Yongjun Li 0006, Dongxu Li 0003
IEEE Trans. Comput. Soc. Syst.3
2026 AlLNet: Alternative Learning is Better for Multimodal Fake News Detection
abstract
The significance of multimodal feature fusion in multimodal-based fake news detection is crucial. However, each modality’s contribution to the final detection outcome varies, and due to substantial differences in data distribution, a simple fusion approach is often inadequate. While existing works attempt to perform multimodal fusion using attention mechanisms, they still face the challenge of modal laziness—where a modality may be underutilized during training due to its low weight, resulting in insufficient learning. In this article, we propose the model alternative learning network (AlLNet). AlLNet addresses the issue of modal laziness by learning unimodal features separately during the training stage and employing an alternating independent training process. Additionally, it incorporates contrast learning to enhance the discriminative ability of image features. During the inference stage, AlLNet evaluates the uncertainty of different modalities using entropy and combines their predictions dynamically with appropriate weights to produce the final result. Moreover, AlLNet is designed to be lightweight and effectively handles modalities missing, making it well-suited for real-world applications. Our experiments conducted on three real-world datasets in English and Chinese demonstrate the advantages and effectiveness of our approach.
Yiyuan Zhu, Yongjun Li 0006
IEEE Trans. Comput. Soc. Syst.2
2025 TACSNet: Two-stage alignment cross-modal differential sensing network for concept prediction
Yuxin Ji, Yongjun Li 0006, Zeyuan Qu, Jiali Wei, Wenhui Sun
Knowl. Based Syst.2
2024 Multi-interest sequential recommendation with contrastive learning and temporal analysis
Yongjun Li 0006
Knowl. Based Syst.3
2024 Learning Frequency-Aware Cross-Modal Interaction for Multimodal Fake News Detection
abstract
Recently, fake news detection (FND) is an essential task in the field of social network analysis, and multimodal detection methods that combine text and image have been significantly explored in the last five years. However, the physical features of images that can be clearly shown in the frequency level are often ignored, and thus cross-modal feature extraction and interaction still remain a great challenge when the frequency domain is introduced for multimodal FND. To address this issue, we propose a frequency-aware cross-modal interaction network (FCINet) for multimodal FND in this article. First, a triple-branch encoder with robust feature extraction capacity is proposed to explore the representation of frequency, spatial, and text domains, separately. Then, we design a parallel cross-modal interaction strategy to fully exploit the interdependencies among them to facilitate multimodal FND. Finally, a combined loss function including deep auxiliary supervision and event classification is introduced to improve the generalization ability for multitask training. Extensive experiments and visual analysis on two public real-world multimodal fake news datasets show that the presented FCINet obtains excellent performance and exceeds numerous state-of-the-art methods.
Yanfeng Liu, Yongjun Li 0006
IEEE Trans. Comput. Soc. Syst.3
2024 Fake Social Media News Detection Based on Forwarding User Representation
abstract
Recently, the proliferation of fake news has imposed a dramatic impact on the public, so the fake news detection has been attracting increasing attention. Most of existing methods employ traditional classification models or neural networks that consider news content, comments, and context to detect fake news. However, the characteristics of news vary by type, topic, and context, which are difficult to uniformly represent for fake news detection. In this article, we propose a novel fake news detection model based on retweeting user embedding using only the stable and easily accessible information of retweeting users. Our model includes three components. First, we use all retweeting users to construct the retweeting graph and then learn the user representation. Second, for each tweet, we aggregate the representations of its retweeting users to learn the tweet representation. Finally, the tweet representations are fed into a feedforward neural network to train the news classifier. To demonstrate the effectiveness of our method, we conduct experiments on two public datasets. The results show that our classifier performs well and achieves better results than other state-of-the-art methods.
Zhaojie Yan, Yongjun Li 0006, Lirong Huang, Wenli Ji
IEEE Trans. Comput. Soc. Syst.2
2024 A Trajectory-Oriented Locality-Sensitive Hashing Method for User Identification
abstract
User identification across social sites, which benefits many applications, has recently been attracting considerable attention. Most existing methods focused more on the effectiveness of user identification, rather than on efficiency. Matching as many crosssite user accounts as possible, which causes very high computation overhead posed by the full-scale pairwise comparisons, remains unsolved, especially when the number of users reaches tens of millions or more. To address this issue, we present a novel locality sensitive hashing method for user identification (UI-LSH), which consists of four components. (1) It involves embedding stay points into vectors, (2) and constructing locality-sensitive hashing families suitable for stay points. (3) It presents a method for projecting stay points into hash buckets that ensures the close stay points are placed in the same bucket with high probability. (4) It constructs the candidate user pairs based on the projection results. The experiments on three ground-truth datasets show that our method reduces the number of user pairs to be compared by as much as 81.87%, 67.68%, and 63.15%, respectively. Overall, UI-LSH holds great promise for significantly improving the efficiency of user identification.
Yongjun Li 0006, Wenli Ji
IEEE Trans. Knowl. Data Eng.1
2023 A Binary-Search-Based Locality-Sensitive Hashing Method for Cross-Site User Identification
abstract
In recent years, we have witnessed increasing attention paid to the problem of cross-site user identification (CSUI) with the bloom of social media. Despite noticeable progress in this field, the problem of enormous computation posed by the full-scale pairwise comparison still remains unsolved, which constrains the real application of the state-of-the-art approaches, especially when the user number reaches tens of millions. To address this issue, we propose a novel trajectory-oriented method, which aims at incorporating the locality-sensitive hashing (LSH) technology into the strategy of approximate nearest neighbors searching. Specifically, we design an LSH function based on binary search (BLSH) for processing user trajectory easily, and by searching the BLSH buckets, we cluster similar users to avoid the onerous full-scale pairwise comparison. Moreover, a hierarchy of discrete attention is devised to further improve the effectiveness, which develops our method into the mature one, hierarchically attentioned binary-search-based LSH (HA-BLSH). Extensive experiments on ground-truth datasets demonstrate that our HA-BLSH outperforms the cutting-edge approaches on efficiency and effectiveness. The further discussion indicates that our method is also quite flexible and able to accommodate more scenarios.
Wenqiang He, Yongjun Li 0006, Yinyin Zhang
IEEE Trans. Comput. Soc. Syst.2
2022 A Location Recall Strategy for Improving Efficiency of User-Generated Short Text Geolocalization
abstract
Currently, users’ daily lives are closely related to social sites. Geolocalizing user-generated short texts (UGSTs) can provide services for the wider fields of business and society. The existing works focus on improving the accuracy of geolocalization algorithms and reducing the prediction error. However, little work has been done on improving the computing efficiency of geolocalization algorithms. When dealing with large-scale data, the existing algorithm efficiency is greatly reduced. To solve this problem, we summarize three modes of user access to point of interests (PoIs), further propose a variety of location recall methods, and then introduce these geolocation recall methods into the existing algorithm. Geolocation recall generates a candidate location set for each user according to the user’s historical behavior records to reduce the search space of the geolocalization algorithm and improve its efficiency. We carried out experiments on the ground-truth datasets, compared the performance of various location recall methods, and verified the excellent location recall performance of three geolocalization algorithms. The experimental results show that location recall shortens the calculation time of the three algorithms by approximately 70% and further improves the accuracy of these algorithms.
Congjie Gao, Yongjun Li 0006, Jiaqi Yang 0003, Yinyin Zhang
IEEE Trans. Comput. Soc. Syst.2
2021 Measuring the short text similarity based on semantic and syntactic information
Jiaqi Yang 0003, Yongjun Li 0006, Congjie Gao, Yinyin Zhang
Future Gener. Comput. Syst.2
2021 Matching user accounts with spatio-temporal awareness across social networks
Yongjun Li 0006, Wenli Ji, Wei Dong 0006, Dongxu Li 0003
Inf. Sci.1
2020 Entity disambiguation with context awareness in user-generated short texts
Jiaqi Yang 0003, Yongjun Li 0006, Congjie Gao, Wei Dong 0006
Expert Syst. Appl.2
2020 Exploiting similarities of user friendship networks across social networks for user identification
Yongjun Li 0006, Zhaoting Su, Jiaqi Yang 0003, Congjie Gao
Inf. Sci.1
2019 Fine-Grained Geolocalization of User-Generated Short Text based on Weight Probability Model
abstract
Recently, the fine-grained geolocalization of User-Generated Short Text (UGST) has become increasingly important. Existing methods can not make full use of the location information in the UGSTs. Besides, existing works only consider the importance of terms for all locations, but do not distinguish the importance of the same term in different locations. To solve these problems, we propose a fine-grained geolocalization method based on a weight probability model (FGST-WP). The method mainly includes three parts: 1) Using the reverse maximum match algorithm to filter out UGSTs that do not contain any location indicative information. 2) Building coupling of terms and locations and adopting a mixed weight strategy to assign weights to terms. 3) Calculating the probability of non-geotagged UGST posted from each location and selecting k locations according to the top-k probabilities. Experiments on ground-truth datasets prove the superior performance of FGST-WP.
Congjie Gao, Yongjun Li 0006, Jiaqi Yang 0003
CIKM2
2019 Online User Representation Learning Across Heterogeneous Social Networks
abstract
Accurate user representation learning has been proven fundamental for many social media applications, including community detection, recommendation, etc. A major challenge lies in that, the available data in a single social network are usually very limited and sparse. In real life, many people are members of several social networks in the same time. Constrained by the features and design of each, any single social platform offers only a partial view of a user from a particular perspective. In this paper, we propose MV-URL, a multi-view user representation learning model to enhance user modeling by integrating the knowledge from various networks. Different from the traditional network embedding frameworks where either the whole framework is single-network based or each network involved is a homogeneous network, we focus on multiple social networks and each network in our task is a heterogeneous network. It's very challenging to effectively fuse knowledge in this setting as the fusion depends upon not only the varying relatedness of information sources, but also the target application tasks. MV-URL focuses on two tasks: user account linkage (i.e., to predict the missing true user account linkage across social media) and user attribute prediction. Extensive evaluations have been conducted on two real-world collections of linked social networks, and the experimental results show the superiority of MV-URL compared with existing state-of-art embedding methods. It can be learned online, and is trivially parallelizable. These qualities make it suitable for real world applications.
Weiqing Wang 0001, Hongzhi Yin, Xingzhong Du, Wen Hua, Yongjun Li 0006, Nguyen Quoc Viet Hung
SIGIR5
2019 Matching user accounts across social networks based on username and display name
Yongjun Li 0006, Hongzhi Yin, Quanqing Xu
World Wide Web1
2018 User Identification with Spatio-Temporal Awareness across Social Networks
abstract
User Identification with Spatio-Temporal awareness has been attracting much attention from academia. The existing methods not only handle temporal and spatial data separately but also do not consider conflictive check-in records. To tackle these problems, we propose a novel approach that consists of three parts. 1) Measure the similarity of users with a kernel density estimation based method, which handles the spatial and temporal data together; 2) Assign weights to check-in records following the idea of TFIDF, which highlights the discriminative check-in records; 3) Penalize the similarity of users based on the number of conflictive check-in records. The pair-wise users with similarity higher than a predefined threshold are considered to belong to the same individual. Experiments demonstrate the superiority of the proposed approach.
Wenli Ji, Yongjun Li 0006, Wei Dong 0006
CIKM3
2018 Building an Ethereum and IPFS-Based Decentralized Social Network System
abstract
Evolvement of blockchain technology has greatly changed the network and it makes many applications to be distributed, decentralized without loss of security. Ethereum is an open-source blockchain platform that provides a runtime environment for running smart contracts, which is called Ethereum Virtual Machine (EVM). Ethereum-based applications are usually referred to as Decentralized Applications (DApps), since they are based on the decentralized EVM, and its smart contracts. Meanwhile, distributed data store also evolves fast with the blockchain technology. Distributed storage develops to reduce the cost of the server side hardware and increase data availability. InterPlanetary File System (IPFS) is a protocol for distributed storage. IPFS stores immutable data, remove duplication, and obtain address information for storage nodes to search for files in the network. Many DApps have been created with the use of these technologies and one example is to use this design for a decentralized Twitter-like system that is resistant to censorship and single point of failure. This paper involves researching the blockchain technology and implementing a decentralised social network application on the Ethereum private blockchain with the use of smart contract and IPFS. In addition, it examines how blockchain and distributed storage can enable more functionalities of traditional social network systems.
Quanqing Xu, Zhiwen Song, Rick Siow Mong Goh, Yongjun Li 0006
ICPADS4
2018 Matching user accounts based on user generated content across social networks
Yongjun Li 0006, Hongzhi Yin, Quanqing Xu
Future Gener. Comput. Syst.1
2018 Expediting Binary Fuzzing with Symbolic Analysis
abstract
Fuzzing is an important method for binary vulnerability mining. It can analyze binary programs without their source codes, which is not easy to do by other technologies. But due to the blindness of input generation, binary fuzzing often falls into traps for a long time when the new mutated inputs cannot generate unexplored paths. In this paper, we propose an efficient and flexible fuzzing framework named Tinker. It defines the growth rate of path coverage to measure the current state of fuzzing. If the fuzzing falls into low-speed or blocked states, a symbolic analysis procedure is invoked to generate a new input which can help the fuzzing jump out of the trap. In the symbolic analysis procedure, we employ dynamic execution to track the traversed nodes. The untraversed branches are then identified according to the recorded data of American Fuzzy Lop (AFL) [M. Zalewski, American Fuzzy Lop (2014), http://lcamtuf.coredump.cx/afl/ ]. At last, we employ control flow graph (CFG) to construct complete paths to these branches and a new input is generated using symbolic execution. Moreover, to expedite the detection of vulnerabilities, we generate inputs which trigger more high-risk system calls first, such that the possibility of finding vulnerabilities can be improved. Tinker has been implemented and the experiments on DARPA CGC benchmark show that Tinker is more efficient in vulnerability mining than state-of-the-art binary vulnerability mining tools.
Luhang Xu, Liangze Yin, Wei Dong 0006, Weixi Jia, Yongjun Li 0006
Int. J. Softw. Eng. Knowl. Eng.5
2018 A deep dive into user display names across social networks
Yongjun Li 0006, Quanqing Xu, Hongzhi Yin
Inf. Sci.1
2018 A Comment on "Cross-Platform Identification of Anonymous Identical Users in Multiple Social Media Networks"
abstract
The Friend Relationship-Based User Identification (FRUI) algorithm is considered to be the ideal method. However, if the seed User Matched Pairs (UMPs) were not suitable, FRUI would stop early due to Controversial UMPs. We highlight this gap and propose a minor change to make FRUI a more general identification algorithm.
Yongjun Li 0006, Zhaoting Su
IEEE Trans. Knowl. Data Eng.1
2017 A Solution to Tweet-Based User Identification Across Online Social Networks
Yongjun Li 0006
ADMA1