Conghui Zheng

dblp:238/0928 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 CDCGAN: Class Distribution-aware Conditional GAN-based minority augmentation for imbalanced node classification
Bojia Liu, Conghui Zheng, Fuhui Sun, Li Pan 0002
Neural Networks2
2024 Global Prosperity or Local Monopoly? Understanding the Geography of App Popularity
abstract
App stores allow developers to globally distribute their apps to gain more users and attention. In the highly competitive market of app stores, developers need to cater to a large number of users spanning multiple countries. We posit that the characteristics of diverse geographical, linguistic, cultural, societal, and economic environments may impact the adoption of apps. In this paper, we take the first step to characterize popular apps across over 150 countries worldwide, and explore the potential correlations to a number of underlying factors including geography, language as well as cultural, societal, and economic dimensions. Our study is based on a longitudinal (one-year) dataset of daily app popularity from the iOS app stores, covering 154 regions around the world. We reveal that app popularity shows great diversity across the world, while similarities exist among countries that share geographical proximity and linguistic convergence. The differences in app popularity across regions can be further correlated with the cultural model and socioeconomic indices we adopt. On top of the dataset and findings, we implement a prediction task that contributes to app distribution, helping developers choose the right market to distribute and promote their apps. To the best of our knowledge, we are the first to attempt to provide a global understanding of the characteristics of app popularity across the mobile app ecosystem. Our observations can benefit stakeholders in the ecosystem, striving to improve app uptake.
Liu Wang 0002, Conghui Zheng, Haoyu Wang 0001, Xiapu Luo, Gareth Tyson, Yi Wang 0004, Shangguang Wang
MSR2
2024 JORA: Weakly Supervised User Identity Linkage via Jointly Learning to Represent and Align
abstract
The user identity linkage that establishes correspondence between users across networks is a fundamental issue in various social network applications. Efforts have recently been devoted to introducing network embedding techniques that map the different network users into the common representation space, thereby inferring user correspondence based on the similarities of their representations. However, existing studies that separately train the network embedding and space alignment in two stages may lead to conflict between the objectives of the two stages. Besides, the similarities between unlabeled cross-network user pairs are difficult to define and largely impact the result. Moreover, many previous methods still need plenty of labeled aligned user pairs to ensure performance, which may not be available. To address the above problems, we propose to solve the weakly-supervised user identity linkage problem via JOintly learning to Represent and Align, i.e., the JORA model. The architecture of JORA adopts the inductive graph convolutional network (GCN) that learns representations for each network. The model is jointly optimized by the representation learning component and alignment learning component. The former one aims to preserve the similarities between intranetwork users. The latter one aligns the different spaces by a projection function and aims to preserve the similarities between cross-network users. A specific attention mechanism is proposed to learn self-adaptive similarities for unlabeled user pairs during alignment learning and it reduces the error propagation caused by predefined similarities. The joint optimization helps perceive network characteristics during alignment and reduces the number of labeled users required. Experiments conducted on real social networks show that the proposed model achieves significantly better performance than the state-of-the-art methods.
Conghui Zheng, Li Pan 0002, Peng Wu 0013
IEEE Trans. Neural Networks Learn. Syst.1
2023 Homogeneous Entity Context Enhanced Representation Network for Temporal Knowledge Graph Reasoning
abstract
Aiming to extrapolate missing facts in the future timestamp, temporal knowledge graph (TKG) reasoning has significant practical value across various applications. Most existing methods mainly rely on direct historical interactions to extract the temporal features of entities and have achieved promising performance in some specific scenes. However, these methods are constrained when it comes to predicting emerging events that lack such direct historical interaction information. To address this limitation, we propose a novel approach called the Homogeneous Entity Context Enhanced Representation Network (HECERN). Homogeneous entities are defined as entities with similar behavior patterns. HECERN takes advantage of the complex interaction connections among the relations of homogeneous entities, which contain rich and relevant context information for accurate prediction of emerging events. Specifically, we generate corresponding entity context subgraph for each history subgraph and extract dependencies for homogeneous entities and neighboring entities based on these two subgraphs, respectively. In addition, we introduce an attention-guided fusion mechanism that dynamically integrates information from neighbor-level dependencies and homogeneous-level dependencies to effectively generate the final entity representation. The effectiveness of HECERN is demonstrated through comprehensive experiments on three publicly available event-based TKG datasets. The results clearly indicate that our proposed approach significantly outperforms state-of-the-art methods in accurately predicting future events. (The source code is available at https://anonymous.4open.science/r/HECERN-2C6B.)
Yujia Yang, Conghui Zheng, Li Pan 0002
ICDM2
2023 A Unified Generative Adversarial Learning Framework for Improvement of Skip-Gram Network Representation Learning Methods
abstract
Network Representation Learning (NRL), which aims to embed nodes into a latent, low-dimensional vector space while preserving some network properties, facilitates the network analysis tasks. The goal of most NRL methods is to make similar nodes represented similarly in the embedding space. Many methods adopt the skip-gram model to achieve such goal by maximizing the predictive probability among the context nodes for each center node. The context nodes are usually determined based on the concept of \textit{proximity} which is defined based on some explicit network features. However, these proximities may result in a loss of training samples and have limited discriminative power. We propose a general and unified generative adversarial learning framework to address the problems. The proposed framework can handle almost all kinds of networks in a unified way, including homogeneous plain networks, attribute augmented networks and heterogeneous networks. It can improve the performances of the most of the state-of-the-art skip-gram based NRL methods. Moreover, another unified and general NRL method is extended from the framework. It can learn the network representation independently. Extensive experiments on proximity preserving evaluation and two network analysis tasks, i.e., link prediction and node classifications, demonstrate the superiority and versatility of our framework.
Peng Wu 0013, Conghui Zheng, Li Pan 0002
IEEE Trans. Knowl. Data Eng.2
2023 Attribute Augmented Network Embedding Based on Generative Adversarial Nets
abstract
Network embedding is to learn low-dimensional representations of nodes while preserving necessary information for network analysis tasks. Though representations preserving both structure and attribute features have achieved in many real-world applications, learning these representations for networks with attribute information is difficult due to the heterogeneity between structure and attribute information. Many existing methods have been proposed to preserve explicit proximities between nodes, with optimization limited to node pairs with large structure and attribute proximities, which may lead to overfitting. To address the above problems, we adopt an attribute augmented network to represent attribute and structure information in a unified framework. Specifically, we study the problem of attribute augmented network embedding that exploits the strength of generative adversarial nets (ANGANs) in capturing the latent distribution of data to learn robust and informative representations of nodes. The ANGAN method obtains the low-dimensional representations of nodes through adversarial learning between the generative and discriminative models. The generative model approximates the underlying connectivity and attributes distributions of nodes by using the distributions generated from the learned representations. It is implemented by utilizing the properties of the attribute augmented network to improve the traditional Skip-gram model. The discriminative model is designed as a binary classifier to distinguish the truly connected node pairs from the generated ones. The pre-training algorithm and the teacher forcing approach are adopted to improve training efficiency and stability. Empirical results show that ANGAN generally outperforms state-of-the-art methods in various real-world applications, which demonstrates the effectiveness and generality of our method.
Conghui Zheng, Li Pan 0002, Peng Wu 0013
IEEE Trans. Neural Networks Learn. Syst.1
2022 CAMU: Cycle-Consistent Adversarial Mapping Model for User Alignment Across Social Networks
abstract
The user alignment problem that establishes a correspondence between users across networks is a fundamental issue in various social network analyses and applications. Since symbolic representations of users suffer from sparsity and noise when computing their cross-network similarities, the state-of-the-art methods embed users into the low-dimensional representation space, where their features are preserved and establish user correspondence based on the similarities of their low-dimensional embeddings. Many embedding-based methods try to align latent spaces of two networks by learning a mapping function before computing similarities. However, most of them learn the mapping function largely based on the limited labeled aligned user pairs and ignore the distribution discrepancy of user representations from different networks, which may lead to the overfitting problem and affect the performance. To address the above problems, we propose a cycle-consistent adversarial mapping model to establish user correspondence across social networks. The model learns mapping functions across the latent representation spaces, and the representation distribution discrepancy is addressed through the adversarial training between the mapping functions and the discriminators as well as the cycle-consistency training. Besides, the proposed model utilizes both labeled and unlabeled users in the training process, which may alleviate the overfitting problem and reduce the number of labeled users required. Results of extensive experiments demonstrate the effectiveness of the proposed model on user alignment on real social networks.
Conghui Zheng, Li Pan 0002, Peng Wu 0013
IEEE Trans. Cybern.1
2021 Mining Set of Interested Communities with Limited Exemplar Nodes for Network Based Services
abstract
Community detection provides invaluable help for various network based services, such as marketing and product recommendation. A specific service usually requires a set of interested communities rather than all communities in the network. In this paper, we address the cases where some exemplar nodes are provided in advance and the set of interested communities is mined for some specific services. Providing sufficient and suitable priori exemplars is not an easy task in most cases. With inadequate priori knowledge, most of recent community detection methods may fail to capture the requirements of a service. We describe the service requirements' essence by a so-called interested attribute subspace with large importance weights on some focus attributes, and study the problem of detecting the set of interested communities based on the guidance of the most limited exemplar information, i.e., two exemplar nodes from any potential interested community. An Interested Subspace and Community Mining (ISCM) method is proposed. In ISCM, a priori knowledge extension technique is designed at first by utilizing the neighborhood of the two exemplar nodes to get more exemplar nodes. Then the interested subspace is inferred from the extension. Finally the set of interested communities are located and mined by the guidance of the interested subspace. Experiments on synthetic datasets demonstrate the effectiveness and efficiency of our method and applications on real-world datasets show its application values for network based services.
Peng Wu 0013, Li Pan 0002, Conghui Zheng
IEEE Trans. Serv. Comput.3
2020 Multimodal Deep Network Embedding With Integrated Structure and Attribute Information
abstract
Network embedding is the process of learning low-dimensional representations for nodes in a network while preserving node features. Existing studies only leverage network structure information and emphasize the preservation of structural features. However, nodes in real-world networks often have a rich set of attributes providing extra semantic information. It has been demonstrated that both structural and attribute features are important for network analysis tasks. To preserve both features, we investigate the problem of integrating structure and attribute information to perform network embedding and propose a multimodal deep network embedding (MDNE) method. MDNE captures the non-linear network structures and the complex interactions among structures and attributes using a deep model consisting of multiple layers of non-linear functions. Since structures and attributes are two different types of information, a multimodal learning method is adopted to pre-process them and help the model to better capture the correlations between node structure and attribute information. We define the loss function employing structural and attribute proximities to preserve the respective features, and the representations are obtained by minimizing the loss function. Results of extensive experiments on four real-world data sets show that the proposed method performs significantly better than baselines on a variety of tasks, which demonstrates the effectiveness and generality of our method.
Conghui Zheng, Li Pan 0002, Peng Wu 0013
IEEE Trans. Neural Networks Learn. Syst.1