EDBT 2026 Demo / reviewers in the wild / expert
Tiantian He 0001
dblp:151/4420-1
· DBLP profile ↗
40ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0003-4839-681XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 11 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Structurally Stabilized Representations for Lossless DNA StorageabstractThis paper presents Reed-Solomon coded single-stranded representation learning (RSRL), a novel end-to-end model for learning representations for lossless DNA data storage. In contrast to existing learning-based methods, RSRL is inspired by both error-correction codec and structural biology. Specifically, RSRL first learns the representations for the subsequent storage from the binary data transformed by the Reed-Solomon codec (RS code). Then, the representations are masked by an RS-code-informed mask to focus on correcting the burst errors occurring in the learning process. The synergy of RS masks and graph attention enables active error localization, breaking through the limitations of traditional passive error correction. With the decoded representations with error corrections, a novel biologically stabilized loss is formulated to regularize the data representations to possess stable single-stranded structures. By incorporating these novel strategies, RSRL can learn highly durable, dense, and lossless representations for subsequent storage tasks in DNA sequences. The proposed RSRL has been compared with a number of baselines in real-world tasks of multi-type data storage. The experimental results obtained demonstrate that RSRL can store diverse types of data with much higher information density and durability, but much lower error rates. Ben Cao, Xue Li 0019, Tiantian He 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008 |
AAAI | 3 |
| 2026 | A Structural Knowledge Enhanced Re-ranking method with large language models for temporal knowledge graph prediction
Bo Li 0001, Bin Song 0001, Tiantian He 0001, Yew-Soon Ong |
Pattern Recognit. | 3 |
| 2026 | Adaptive PID-Incorporated Nonnegative Latent Factor AnalysisabstractHigh-dimensional and incomplete (HDI) data are ubiquitous for representing the intricate interactions among a vast number of nodes arising from diverse real application scenarios. Nonnegative latent factor (NLF) models have demonstrated their effectiveness in extracting critical latent features from HDI data by leveraging a single latent factor (LF)-dependent, nonnegative, and multiplicative update (SLF-NMU) algorithm. However, the SLF-NMU algorithm always leads NLF models to converge sluggishly as it updates an LF based only on the current updated information. To address this critical issue, we present APNLF, which innovatively adopts an adaptive proportional–integral–derivative (PID) controller to enable the learning process of SLF-NMU to be more efficient. The proposed APNLF encompasses the following twofold ideas: 1) establishing a PID-increment-based SLF-NMU (PSN) algorithm, which updates an LF by comprehensively modeling the current, past, and future update increment information guided by the principle of a PID controller; and 2) designing an effective fuzzy reasoning rule to implement all hyperparameters adaptation, which boosts model’s applicability in practical applications. Moreover, APNLF’s convergence analysis is provided in theory. Experiments on five HDI datasets demonstrate that the APNLF model outperforms the state-of-the-art models in efficiency and accuracy on an HDI matrix. The source code is available athttps://github.com/Aaaaapplege/APNLF-Development Jinli Li, Ye Yuan 0014, Tiantian He 0001, Xin Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Voronoi-grid-based Pareto Front Learning and Its Application to Collaborative Federated LearningabstractMulti-objective optimization (MOO) exists extensively in machine learning, and aims to find a set of Pareto-optimal solutions, called the Pareto front, e.g., it is fundamental for multiple avenues of research in federated learning (FL). Pareto-Front Learning (PFL) is a powerful method implemented using Hypernetworks (PHNs) to approximate the Pareto front. This method enables the acquisition of a mapping function from a given preference vector to the solutions on the Pareto front. However, most existing PFL approaches still face two challenges: (a) sampling rays in high-dimensional spaces; (b) failing to cover the entire Pareto Front which has a convex shape. Here, we introduce a novel PFL framework, called as PHN-HVVS, which decomposes the design space into Voronoi grids and deploys a genetic algorithm (GA) for Voronoi grid partitioning within high-dimensional space. We put forward a new loss function, which effectively contributes to more extensive coverage of the resultant Pareto front and maximizes the HV Indicator. Experimental results on multiple MOO machine learning tasks demonstrate that PHN-HVVS outperforms the baselines significantly in generating Pareto front. Also, we illustrate that PHN-HVVS advances the methodologies of several recent problems in the FL field. The code is available at https://github.com/buptcmm/phnhvvs. Qiqi Liu, Tiantian He 0001, Yew-Soon Ong, Yaochu Jin, Qicheng Lao, Han Yu 0001 |
ICML | 4 |
| 2025 | COSTA: Contrastive Spatial and Temporal Debiasing framework for next POI recommendation
Zhu Sun 0001, Tiantian He 0001, Shanshan Feng 0001, Guanfeng Liu 0001 |
Neural Networks | 4 |
| 2025 | Natural Gas Pipeline Leak Detection Based on Dual Feature Drift in Acoustic SignalsabstractDetecting leaks in natural gas pipelines using acoustic signals typically requires extensive prior knowledge and complex parameter designs, making it challenging to handle background noise and data distribution disparities simultaneously. This article proposes a dual-feature drift framework utilizing a nonparametric design approach for acoustic signal-based leak detection. This framework consists of two core technologies: first, feature backward normalization. Low-dimensional drift factors are designed based on transformed acoustic signals to exponentially normalize the time-periodic features of the signal feature matrix, thereby eliminating strong background noise. Second, constructing the feature drift layer within a one-dimensional convolutional neural network. Weighted parameters constrain a high-dimensional feature matrix, developing drift factors that perform exponential drift on each feature, thus enhancing gradient constraints and eliminating data distribution differences during model training. This framework achieves a fault identification accuracy of 95.46% for natural gas pipeline leaks, outperforming competing methods and representing a novel approach to intelligent pipeline leak detection. Lizhong Yao, Yu Zhang 0289, Ling Wang 0001, Rui Li 0087, Tiantian He 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Graph Linear Convolution Pooling for Learning in Incomplete High-Dimensional DataabstractHigh-dimensional and incomplete (HDI) data are frequently encountered in diverse real-world applications involving complex interactions among numerous nodes. Approaches based on latent feature analysis (LFA) have proven effective in performing representation learning in HDI data. Nevertheless, they cannot handle the high-order connectivity among nodes in HDI data well, resulting in severe accuracy loss. To address the previously mentioned issue, we present a novel model in this paper, namely Graph Linear Convolution Pooling Network (GLCPN). The proposed GLCPN adopts the three-fold ideas. First, it leverages simplified graph convolutions to efficiently capture high-order connectivity among nodes for learning representations of matrix factorization. Second, a simple yet effective priori convolution operator is adopted by each graph neural layer to capture node-node collaboration for aggregation. Third, a locality-enhanced pooling scheme is designed to holistically utilize multi-layer representations of the neighborhood. Therefore, GLCPN can effectively acquire the hidden information in HDI data with high efficiency. In addition, we have conducted a theoretical analysis demonstrating that the proposed GLCPN is more expressive compared with existing graph neural networks for HDI data. Extensive experiments have been further conducted on ten well-established HDI datasets from various applications. The experimental results demonstrate that the proposed GLCPN significantly outperforms state-of-the-art models for learning representations in HDI data evaluated by accuracy and efficiency metrics. Fanghui Bi, Tiantian He 0001, Yew-Soon Ong, Xin Luo 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Discovering Spatiotemporal-Individual Coupled Features From Nonstandard Tensors - A Novel Dynamic Graph Mixer ApproachabstractIn this article, we present the dynamic graph mixer (DGM), a novel model for learning spatiotemporal-individual coupled features from high-dimensional and incomplete (HDI) tensors, which frequently represent dynamic interactions among real-world data samples. In contrast to existing methods, the proposed DGM possesses the following three advantages when learning representations from HDI tensors. First, it performs light graph message passing based on the conjoint attentions learned by jointly modeling latent features and implicit structures to extract the high-order connectivity. Second, a multilayer nonlinear tensor neural network (TNN) is adopted to learn the intricate attribute features of node-node-time from different views. Third, it follows the Tucker decomposition paradigm in a data density-oriented modeling mechanism to integrate node representations, preserving the overall multidimensional interaction patterns. In addition, we provide theoretical evidence that the key components in DGM can significantly improve expressiveness. Extensive experiments conducted on eight testing datasets of HDI tensors demonstrate that DGM outperforms state-of-the-art methods in both learning accuracy and efficiency. Fanghui Bi, Tiantian He 0001, Yew-Soon Ong, Xin Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Proximal-ADMM-Incorporated Nonnegative Latent-Factorization-of-Tensors Model for Representing Dynamic Cryptocurrency Transaction NetworkabstractCryptocurrency services, as one of the most successful applications of blockchain technology, have recently garnered significant attention from the graph learning community. Its large-scale dynamic transaction records contain a variety of behavioral patterns and rich knowledge involving accounts, making the dynamic cryptocurrency transaction network embedding (DCTNE) a hot, yet thorny research topic. As the trading accounts increase and time accumulates, considerable transaction services are dispersed into various time slots, leading to very sparse transaction data within a time slot, that is, the transaction service data is high-dimensional and incomplete (HDI). To efficiently mine high-value knowledge from HDI data, this article proposes a proximal-ADMM-incorporated nonnegative latent-factorization-of-tensors (PNL) model for DCTNE that adopts threefold ideas: 1) incorporating the proximal terms into the alternating-direction-method-of-multipliers (ADMMs)-based learning scheme to reduce the oscillations for high estimation accuracy and fast convergence; 2) implementing a parallel training process with hyperparameter self-adaptation for high computational efficiency; and 3) proving that the proximal-incorporated learning scheme guarantees the convergence to a Karush–Kuhn–Tucker (KKT) stationary point. Experimental results on eight real-world DCTNs show that the PNL significantly outperforms several state-of-the-art (SOTA) models, demonstrating not only high efficiency and accuracy in performing DCTNE, but also strong potential to enhance the operational reliability and stability of cryptocurrency transaction systems. Xin Liao 0003, Hao Wu 0061, Tiantian He 0001, Xin Luo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | FedCompetitors: Harmonious Collaboration in Federated Learning with Competing ParticipantsabstractFederated learning (FL) provides a privacy-preserving approach for collaborative training of machine learning models. Given the potential data heterogeneity, it is crucial to select appropriate collaborators for each FL participant (FL-PT) based on data complementarity. Recent studies have addressed this challenge. Similarly, it is imperative to consider the inter-individual relationships among FL-PTs where some FL-PTs engage in competition. Although FL literature has acknowledged the significance of this scenario, practical methods for establishing FL ecosystems remain largely unexplored. In this paper, we extend a principle from the balance theory, namely “the friend of my enemy is my enemy”, to ensure the absence of conflicting interests within an FL ecosystem. The extended principle and the resulting problem are formulated via graph theory and integer linear programming. A polynomial-time algorithm is proposed to determine the collaborators of each FL-PT. The solution guarantees high scalability, allowing even competing FL-PTs to smoothly join the ecosystem without conflict of interest. The proposed framework jointly considers competition and data heterogeneity. Extensive experiments on real-world and synthetic data demonstrate its efficacy compared to five alternative approaches, and its ability to establish efficient collaboration networks among FL-PTs. Shanli Tan, Hao Cheng 0014, Han Yu 0001, Tiantian He 0001, Yew-Soon Ong, Chong-Jun Wang, Xiaofeng Tao 0001 |
AAAI | 5 |
| 2024 | Free-Rider and Conflict Aware Collaboration Formation for Cross-Silo Federated LearningabstractFederated learning (FL) is a machine learning paradigm that allows multiple FL participants (FL-PTs) to collaborate on training models without sharing private data. Due to data heterogeneity, negative transfer may occur in the FL training process. This necessitates FL-PT selection based on their data complementarity. In cross-silo FL, organizations that engage in business activities are key sources of FL-PTs. The resulting FL ecosystem has two features: (i) self-interest, and (ii) competition among FL-PTs. This requires the desirable FL-PT selection strategy to simultaneously mitigate the problems of free riders and conflicts of interest among competitors. To this end, we propose an optimal FL collaboration formation strategy -FedEgoists- which ensures that: (1) a FL-PT can benefit from FL if and only if it benefits the FL ecosystem, and (2) a FL-PT will not contribute to its competitors or their supporters. It provides an efficient clustering solution to group FL-PTs into coalitions, ensuring that within each coalition, FL-PTs share the same interest. We theoretically prove that the FL-PT coalitions formed are optimal since no coalitions can collaborate together to improve the utility of any of their members. Extensive experiments on widely adopted benchmark datasets demonstrate the effectiveness of FedEgoists compared to nine state-of-the-art baseline methods, and its ability to establish efficient collaborative networks in cross-silos FL with FL-PTs that engage in business activities. Xiaoli Tang 0001, Tiantian He 0001, Yew-Soon Ong, Qiqi Liu, Qicheng Lao, Han Yu 0001 |
NeurIPS | 4 |
| 2024 | Road Network Representation Learning with the Third Law of GeographyabstractRoad network representation learning aims to learn compressed and effective vectorized representations for road segments that are applicable to numerous tasks. In this paper, we identify the limitations of existing methods, particularly their overemphasis on the distance effect as outlined in the First Law of Geography. In response, we propose to endow road network representation with the principles of the recent Third Law of Geography. To this end, we propose a novel graph contrastive learning framework that employs geographic configuration-aware graph augmentation and spectral negative sampling, ensuring that road segments with similar geographic configurations yield similar representations, and vice versa, aligning with the principles stated in the Third Law. The framework further fuses the Third Law with the First Law through a dual contrastive learning objective to effectively balance the implications of both laws. We evaluate our framework on two real-world datasets across three downstream tasks. The results show that the integration of the Third Law significantly improves the performance of road segment representations in downstream tasks. Haicang Zhou, Weiming Huang 0001, Yile Chen 0001, Tiantian He 0001, Gao Cong, Yew-Soon Ong |
NeurIPS | 4 |
| 2024 | SCG: A Novel Spatiotemporal Coupling Graph Convolutional Network-Incorporated Approach for Dynamic QoS EstimationabstractDynamic Quality-of-Service (QoS) data capturing temporal variations in user-service interactions are essential source for service selection and user behavior understanding. Approaches based on Latent Feature Analysis (LFA) have shown to be beneficial for discovering effective temporal patterns in QoS data. However, existing methods cannot well model the spatiality and temporality implied in dynamic interactions in a unified form, causing abundant accuracy loss for missing QoS estimation. To address the problem, this paper presents a novel Graph Convolutional Network (GCN)-based dynamic QoS estimator namely Spatiotemporal Coupling GCN (SCG) model with the three-fold ideas as below. First, SCG builds its dynamic graph convolutional rules by incorporating generalized tensor product framework, for unified modeling of spatial and temporal patterns. Second, SCG combines the heterogeneous GCN layer with tensor factorization, for effective representation learning on time-varying bipartite user-service graphs. Third, it further simplifies the dynamic GCN structure to lower the training difficulties. Extensive experiments have been conducted on two large-scale widely-adopted QoS datasets describing throughput and response time. The results demonstrate that SCG realizes higher QoS estimation accuracy compared with the state-of-the-arts, illustrating it can learn powerful representations to users and cloud services. Fanghui Bi, Tiantian He 0001 |
SMC | 2 |
| 2024 | Discrete Multi-View Feature Propagation Preserving Graph ClusteringabstractGraph clustering is a fundamental and challenging learning task, which is conventionally approached by grouping similar vertices based on edge structure and feature similarity. In contrast to previous methods, in this paper, we investigate how multi-view feature propagation can influence cluster discovery in graph data. To this end, we present Discrete Multi-View Feature Propagation Preserving Graph Clustering (DMVFPPGC), a novel method that leverages multi-view feature propagation to enhance cluster identification in graph data. DMVFPPGC employs a unified objective function that utilizes graph topology and multi-view vertex features to determine vertex cluster membership, regularized by a module that supports key latent feature propagation. We derive an iterative algorithm to optimize this function, prove model convergence within a finite number of iterations, and analyze its computational complexity. Our experiments on various real-world graphs demonstrate the superior clustering performance of DMVFPPGC compared to well-established methods, manifesting its effectiveness across different scenarios. Zhixuan Duan, Fanghui Bi, Tiantian He 0001 |
SMC | 4 |
| 2024 | Polarized message-passing in graph neural networksabstractIn this paper, we present Polarized message-passing (PMP), a novel paradigm to revolutionize the design of message-passing graph neural networks (GNNs). In contrast to existing methods, PMP captures the power of node-node similarity and dissimilarity to acquire dual sources of messages from neighbors. The messages are then coalesced to enable GNNs to learn expressive representations from sparse but strongly correlated neighbors. Three novel GNNs based on the PMP paradigm, namely PMP graph convolutional network (PMP-GCN), PMP graph attention network (PMP-GAT), and PMP graph PageRank network (PMP-GPN) are proposed to perform various downstream tasks. Theoretical analysis is also conducted to verify the high expressiveness of the proposed PMP-based GNNs. In addition, an empirical study of five learning tasks based on 12 real-world datasets is conducted to validate the performances of PMP-GCN, PMP-GAT, and PMP-GPN. The proposed PMP-GCN, PMP-GAT, and PMP-GPN outperform numerous strong message-passing GNNs across all five learning tasks, demonstrating the effectiveness of the proposed PMP paradigm. Tiantian He 0001, Yang Liu 0007, Yew-Soon Ong, Xin Luo 0001 |
Artif. Intell. | 1 |
| 2024 | A prediction method of diabetes comorbidity based on non-negative latent features
Leming Zhou, Kechen Liu, Hanshu Qin, Tiantian He 0001 |
Neurocomputing | 5 |
| 2024 | Differentiable Clustering for Graph AttentionabstractGraph clusters (or communities) represent important graph structural information. In this paper, we presentDifferentiableClustering for graphATtention (DCAT). To the best of our knowledge, DCAT is the first solution that incorporates graph clustering into graph attention networks (GAT) to learn cluster-aware attention scores for semi-supervised learning tasks. In DCAT, we propose a novel approach to formunderlineating graph clustering as an auxiliary differentiable objective based on modunderlinearity maximization, which can be optimized together with the learning objective of GAT for a semi-supervised task. Specifically, we propose a solution to relaxing modunderlinearity maximization from a discrete optimization problem to a differentiable objective with theoretical guarantee so that we can learn cluster-aware attention scores by jointly learning from graph clustering and a semi-supervised learning task. To address the computational challenge, we further propose to reformunderlineate the constraint introduced by the clustering objective into a new form. Our analysis shows that DCAT allocates higher attention scores to nodes within the same cluster, allowing them to have a higher influence in node representation learning, and thus DCAT will generate better node representations for downstream applications. The experimental resunderlinets on commonly used datasets show that DCAT outperforms popunderlinear and state-of-the-art graph neural networks. Haicang Zhou, Tiantian He 0001, Yew-Soon Ong, Gao Cong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | A Fast Nonnegative Autoencoder-Based Approach to Latent Feature Analysis on High-Dimensional and Incomplete DataabstractHigh-Dimensional and Incomplete (HDI) data are frequently encountered in various Big Data-related applications. Despite its incompleteness, an HDI data repository contains rich knowledge and patterns concerning the complex interactions among numerous nodes. Recently, a Neural Network (NN)-based approach to Latent Feature Analysis (LFA) model becomes popular owing to its strong representation learning ability to HDI data. Nevertheless, existing NN-based LFA models neglect the inherent nonnegativity in most HDI data, resulting in representation accuracy loss. Motivated by this discovery, this study innovatively proposes a Fast Nonnegative AutoEncoder (FNAE)-based approach to LFA on HDI data, whose ideas are three-fold: a) constructing a multilayered autoencoder subject to nonnegativity constraints for high representation learning ability; b) incorporating the data density-oriented modeling mechanism into FNAE's input and output layers for high computational and storage efficiency; and c) implementing an Adam-based single latent factor-dependent, nonnegative and multiplicative update algorithm for efficient model training as well as fulfilling the nonnegativity constraints. Experimental results on eight commonly-adopted HDI matrices from industrial applications demonstrate that the proposed FNAE significantly outperforms several state-of-the-art NN-based LFA models in both estimation accuracy for missing links of an HDI matrix and computational efficiency. Fanghui Bi, Tiantian He 0001, Xin Luo 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Piggybacking on past problem for faster optimization in aluminum electrolysis process design
Lizhong Yao, Tiantian He 0001, Haijun Luo |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Semisupervised Graph Neural Networks for Graph ClassificationabstractGraph classification aims to predict the label associated with a graph and is an important graph analytic task with widespread applications. Recently, graph neural networks (GNNs) have achieved state-of-the-art results on purely supervised graph classification by virtue of the powerful representation ability of neural networks. However, almost all of them ignore the fact that graph classification usually lacks reasonably sufficient labeled data in practical scenarios due to the inherent labeling difficulty caused by the high complexity of graph data. The existing semisupervised GNNs typically focus on the task of node classification and are incapable to deal with graph classification. To tackle the challenging but practically useful scenario, we propose a novel and general semisupervised GNN framework for graph classification, which takes full advantage of a slight amount of labeled graphs and abundant unlabeled graph data. In our framework, we train two GNNs as complementary views for collaboratively learning high-quality classifiers using both labeled and unlabeled graphs. To further exploit the view itself, we constantly select pseudo-labeled graph examples with high confidence from its own view for enlarging the labeled graph dataset and enhancing predictions on graphs. Furthermore, the proposed framework is investigated on two specific implementation regimes with a few labeled graphs and the extremely few labeled graphs, respectively. Extensive experimental results demonstrate the effectiveness of our proposed semisupervised GNN framework for graph classification on several benchmark datasets. Yu Xie 0009, Yanfeng Liang, Maoguo Gong, A. K. Qin 0001, Yew-Soon Ong, Tiantian He 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | Two-Stream Graph Convolutional Network-Incorporated Latent Feature AnalysisabstractHistorical Quality-of-Service (QoS) data describing existing user-service invocations are vital to understanding user behaviors and cloud service conditions. Collaborative Filtering (CF) models based on Matrix Factorization (MF) have proven to be highly efficient in performing representation learning in QoS data. However, its performance is hindered by its linear inherence and implicit encoding of collaborative QoS signal. To address this critical issue, we present a novel approach in this article, dubbed asTwo-streamGraph convolutional network-incorporatedLatentFeatureAnalysis (TGLFA). The proposed TGLFA is significantly different from previous approaches to representation learning in QoS data in the following three aspects. First, it constructs a multilayer fully-connected network to capture the attribute characteristics representing the nonlinear latent features of service. Second, TGLFA constructs a biparty graph to represent the user-service interactions, where the light graph convolutional network is adopted to acquire the high-order connectivity in QoS data. Last, Aiming to improve computational efficiency, the proposed approach considers the mechanism for data density-oriented modeling when building the input and output layers. Detailed experimental results on eight large-scale cases constructed on two real QoS datasets demonstrate that the proposed TGLFA significantly outperforms its state-of-the-art peers in both estimation accuracy for missing QoS data and computational efficiency. The notable results show that TGLFA is a novel and effective approach to QoS data representation learning. Fanghui Bi, Tiantian He 0001, Yuetong Xie, Xin Luo 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | A Two-Stream Light Graph Convolution Network-based Latent Factor Model for Accurate Cloud Service QoS EstimationabstractHistorical Quality-of-Service (QoS) data regarding past user-service invocations are vital to understand the user behaviors and cloud service conditions. A Matrix Factorization (MF)-based Collaborative Filtering (CF) model has proven to be highly effective in performing representation learning to such QoS data. However, its performance is hindered by its linear interaction and implicit encoding of collaborative QoS signal. To address this critical issue, this paper presents a Two-stream Light Graph Convolution Network-based latent factor (TLGCN) model with the three-fold ideas: 1) constructing a multilayered and fully-connected network to represent services’ nonlinear latent features; 2) integrating the user-service interactions, i.e., the bipartite graph structure into the representation learning process with a light graph convolution network for illustrating the high-order connectivity information in QoS data; and 3) incorporating the data density-oriented modeling mechanism into the input and output of TLGCN for high computational efficiency. Experimental results on two real QoS datasets demonstrate that the proposed TLGCN model significantly outperforms its state-of-the-art peers in both estimation accuracy for missing QoS data and computational efficiency. Fanghui Bi, Tiantian He 0001, Xin Luo 0001 |
ICDM | 2 |
| 2022 | A multiobjective prediction model with incremental learning ability by developing a multi-source filter neural network for the electrolytic aluminium processabstractAbstract Improving current efficiency and reducing energy consumption are two important technical goals of the electrolytic aluminum process (EAP). However, because the process involves complex noise characteristics (i.e., unknown types, redundant distributions and variable forms), it is very difficult to accurately develop a multiobjective prediction model. To overcome this problem, in this paper, a novel framework of multiobjective incremental learning based on a multi-source filter neural network (MSFNN) is presented. The proposed framework first presents a “multi-source filter” (MSF) technique that utilizes the mean and variance in the unscented Kalman filter (UKF) to guide the importance function of the particle filter (PF) based on a density kernel estimation method. Then, the MSF is embedded in the mutated neural network to adjust weights in real time. Third, weights are calculated and normalized by a modified importance function, which is the basis for further optimizing a secondary sampling based on sampling importance resampling (SIR). Finally, the incremental learning model with two objectives (i.e., process power consumption and current efficiency) based on the MSFNN in the EAP is established. The presented framework has been verified by the real-world EAP and some closely related methods. All test results indicate that the MSFNN’s relative prediction errors of the above two objectives are controlled within 0.51% and 0.38%, respectively and prove that MSFNN has significant competitive advantages over other recent filtering network models. Successfully establishment of the proposed framework provides a model foundation for multiobjective optimization problems in the EAP. Lizhong Yao, Tiantian He 0001, Shouxin Liu, Ling Nie |
Appl. Intell. | 3 |
| 2022 | Co-Learning Bayesian OptimizationabstractBayesian optimization (BO) is well known to be sample efficient for solving black-box problems. However, BO algorithms may get stuck in suboptimal solutions even with plenty of samples. Intrinsically, such a suboptimal problem of BO can attribute to the poor surrogate accuracy of the trained Gaussian process (GP), particularly that in the regions where the optimal solutions locate. Hence, we propose to build multiple GP models instead of a single GP surrogate to complement each other, thus resolving the suboptimal problem of BO. Nevertheless, according to the bias-variance tradeoff equation, the individual prediction errors can increase when increasing the diversity of models, which may lead even worse overall surrogate accuracy. On the other hand, based on the theory of the Rademacher complexity, it has been proven that exploiting the agreement of models on unlabeled information can reduce the complexity of hypothesis space, therefore achieving the required surrogate accuracy with fewer samples. Such value of model agreement has been extensively demonstrated for co-training style algorithms to boost model accuracy with a small portion of samples. Inspired by the above, we propose a novel BO algorithm labeled as co-learning BO (CLBO), which exploits both model diversity and agreement on unlabeled information to improve the overall surrogate accuracy with limited samples, therefore achieving more efficient global optimization. Through tests on five numerical toy problems and three engineering benchmarks, the effectiveness of the proposed CLBO has been well demonstrated. Zhendong Guo, Yew-Soon Ong, Tiantian He 0001, Haitao Liu 0002 |
IEEE Trans. Cybern. | 3 |
| 2022 | Vicinal Vertex Allocation for Matrix Factorization in NetworksabstractIn this article, we present a novel matrix-factorization-based model, labeled here as Vicinal vertex allocated matrix factorization (VVAMo), for uncovering clusters in network data. Different from the past related efforts of network clustering, which consider the edge structure, vertex features, or both in their design, the proposed model includes the additional detail on vertex inclinations with respect to topology and features into the learning. In particular, by taking the latent preferences between vicinal vertices into consideration, VVAMo is then able to uncover network clusters composed of proximal vertices that share analogous inclinations, and correspondingly high structural and feature correlations. To ensure such clusters are effectively uncovered, we propose a unified likelihood function for VVAMo and derive an alternating algorithm for optimizing the proposed function. Subsequently, we provide the theoretical analysis of VVAMo, including the convergence proof and computational complexity analysis. To investigate the effectiveness of the proposed model, a comprehensive empirical study of VVAMo is conducted using extensive commonly used realistic network datasets. The results obtained show that VVAMo attained superior performances over existing classical and state-of-the-art approaches. Tiantian He 0001, Lu Bai 0005, Yew-Soon Ong |
IEEE Trans. Cybern. | 1 |
| 2021 | Learning Conjoint Attentions for Graph Neural NetsabstractIn this paper, we present Conjoint Attentions (CAs), a class of novel learning-to-attend strategies for graph neural networks (GNNs). Besides considering the layer-wise node features propagated within the GNN, CAs can additionally incorporate various structural interventions, such as node cluster embedding, and higher-order structural correlations that can be learned outside of GNN, when computing attention scores. The node features that are regarded as significant by the conjoint criteria are therefore more likely to be propagated in the GNN. Given the novel Conjoint Attention strategies, we then propose Graph conjoint attention networks (CATs) that can learn representations embedded with significant latent features deemed by the Conjoint Attentions. Besides, we theoretically validate the discriminative capacity of CATs. CATs utilizing the proposed Conjoint Attention strategies have been extensively tested in well-established benchmarking datasets and comprehensively compared with state-of-the-art baselines. The obtained notable performance demonstrates the effectiveness of the proposed Conjoint Attentions. Tiantian He 0001, Yew-Soon Ong, Lu Bai 0005 |
NeurIPS | 1 |
| 2021 | Multi-source propagation aware network clustering☆abstractNetwork cluster analysis is of great importance as it is closely related to diverse applications, such as social community detection, biological module identification, and document segmentation. Aiming to effectively uncover clusters in the network data, a number of computational approaches , which utilize network topology , single vector of vertex features, or both the aforementioned, have been proposed. However, most prevalent approaches are incapable of dealing with those contemporary network data whose vertices are characterized by features collected from multiple sources. To address this challenge, in this paper, we propose a novel framework, dubbed Multi-Source Propagation Aware Network Clustering (MSPANC) for uncovering clusters in network data possessing multiple sources of vertex features. Different from most previous approaches, MSPANC is able to infer the cluster preference for each vertex utilizing both network topology and multi-source vertex features. To improve the practical significance of the discovered clusters, the learning of cluster membership is also involved into the modeling of the maximization of intra-cluster propagation regarding multi-source features. We propose a unified objective function for MSPANC to perform the clustering task and derive an alternative manner of learning algorithm for model optimization. Besides, we theoretically prove the convergence of the algorithm for optimizing MSPANC. The proposed model has been tested on five real-world datasets, including social, biological and document networks, and has been compared with several competitive baselines. The remarkable experimental results validate the effectiveness of MSPANC. Tiantian He 0001, Yew-Soon Ong, Pengwei Hu 0001 |
Neurocomputing | 1 |
| 2020 | Precise object detection using adversarially augmented local/global feature fusionabstractObject detection, which aims at recognizing or locating the objects of interest in remote sensing imagery with high spatial resolutions (HSR), plays a significant role in many real-world scenarios, e.g., environment monitoring, urban planning, civil infrastructure construction, disaster rescuing, and geographic image retrieval . As a long-lasting challenging problem in both machine learning and geoinformatics communities, many approaches have been proposed to tackle it. However, previous methods always overlook the abundant information embedded in the HSR remote sensing images. The effectiveness of these methods, e.g., accuracy of detection, is therefore limited to some extent. To overcome the mentioned challenge, in this paper, we propose a novel two-phase deep framework, dubbed GLGOD-Net, to effectively detect meaningful objects in HSR images . GLGOD-Net firstly attempts to learn the enhanced deep representations from super-resolution image data . Fully utilizing the augmented image representations, GLGOD-Net then learns the fused representations into which both local and global latent features are implanted. Such fused representations learned by GLGOD-Net can be used to precisely detect different objects in remote sensing images . The proposed framework has been extensively tested on a real-world HSR image dataset for object detection and has been compared with several strong baselines. The remarkable experimental results validate the effectiveness of GLGOD-Net. The success of GLGOD-Net not only advances the cutting-edge of image data analytics , but also promotes the corresponding applicability of deep learning in remote sensing imagery . Xiaobing Han, Tiantian He 0001, Yew-Soon Ong, Yanfei Zhong |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Contextual Correlation Preserving Multiview Featured Graph ClusteringabstractGraph clustering, which aims at discovering sets of related vertices in graph-structured data, plays a crucial role in various applications, such as social community detection and biological module discovery. With the huge increase in the volume of data in recent years, graph clustering is used in an increasing number of real-life scenarios. However, the classical and state-of-the-art methods, which consider only single-view features or a single vector concatenating features from different views and neglect the contextual correlation between pairwise features, are insufficient for the task, as features that characterize vertices in a graph are usually from multiple views and the contextual correlation between pairwise features may influence the cluster preference for vertices. To address this challenging problem, we introduce in this paper, a novel graph clustering model, dubbed contextual correlation preserving multiview featured graph clustering (CCPMVFGC) for discovering clusters in graphs with multiview vertex features. Unlike most of the aforementioned approaches, CCPMVFGC is capable of learning a shared latent space from multiview features as the cluster preference for each vertex and making use of this latent space to model the inter-relationship between pairwise vertices. CCPMVFGC uses an effective method to compute the degree of contextual correlation between pairwise vertex features and utilizes view-wise latent space representing the feature-cluster preference to model the computed correlation. Thus, the cluster preference learned by CCPMVFGC is jointly inferred by multiview features, view-wise correlations of pairwise features, and the graph topology. Accordingly, we propose a unified objective function for CCPMVFGC and develop an iterative strategy to solve the formulated optimization problem. We also provide the theoretical analysis of the proposed model, including convergence proof and computational complexity analysis. In our experiments, we extensively compare the proposed CCPMVFGC with both classical and state-of-the-art graph clustering methods on eight standard graph datasets (six multiview and two single-view datasets). The results show that CCPMVFGC achieves competitive performance on all eight datasets, which validates the effectiveness of the proposed model. Tiantian He 0001, Yang Liu 0007, Tobey H. Ko, Keith C. C. Chan, Yew-Soon Ong |
IEEE Trans. Cybern. | 1 |
| 2019 | Manifold Regularized Stochastic Block ModelabstractStochastic block models (SBMs) play essential roles in network analysis, especially in those related to unsupervised learning (clustering). Many SBM-based approaches have been proposed to uncover network clusters, by means of maximizing the block-wise posterior probability that generates edges bridging vertices. However, none of them is capable of inferring the cluster preference for each vertex through simultaneously modeling block-wise edge structure, vertex features, and similarities between pairwise vertices. To fill this void, we propose a novel SBM dubbed manifold regularized stochastic model (MrSBM) to perform the task of unsupervised learning in network data in this paper. Besides modeling edges that are within or connecting blocks, MrSBM also considers modeling vertex features utilizing the probabilities of vertex-cluster preference and feature-cluster contribution. In addition, MrSBM attempts to generate manifold similarity of pairwise vertices utilizing the inferred vertex-cluster preference. As a result, the inference of cluster preference may well capture the comparability in the manifold. We design a novel process for network data generation, based on which, we specify the model structure and formulate the network clustering problem using a novel likelihood function. To guarantee MrSBM learns the optimal cluster preference for each vertex, we derive an effective Expectation-Maximization based algorithm for model fitting. MrSBM has been tested on five sets of real-world network data and has been compared with both classical and state-of-the-art approaches to network clustering. The competitive experimental results validate the effectiveness of MrSBM. Tiantian He 0001, Lu Bai 0005, Yew-Soon Ong |
ICTAI | 1 |
| 2019 | Measuring Boundedness for Protein Complex Identification in PPI NetworksabstractThe problem of identifying protein complexes in Protein-Protein Interaction (PPI) networks is usually formulated as the problem of identifying dense regions in such networks. In this paper, we present a novel approach, called TBPCI, to identify protein complexes based instead on the concept of a measure of boundedness. Such a measure is defined as an objective function of a Jaccard Index-based connectedness measure which takes into consideration how much two proteins within a network are connected to each other, and an association measure which takes into consideration how much two connecting proteins are associated based on their attributes found in the Gene Ontology database. Based on the above two measures, the objective function is derived to capture how strong the proteins can be considered as bounded together and the objective value is therefore referred as the aggregated degree of boundedness. To identify protein complexes, TBPCI computes the degree of boundedness between all possible pairwise proteins. Then, TBPCI uses a Breadth-First-Search method to determine whether a protein-pair should be incorporated into the same complex. TBPCI has been tested with several real data sets and the experimental results show it is an effective approach for identifying protein complexes in PPI networks. Tiantian He 0001, Keith C. C. Chan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Learning Latent Patterns in Molecular Data for Explainable Drug Side Effects Prediction
Pengwei Hu 0001, Zhu-Hong You, Tiantian He 0001, Shaochun Li, Shuhang Gu, Keith C. C. Chan |
BIBM | 3 |
| 2018 | Learning Latent Factors in Linked Multi-modality Data
Tiantian He 0001, Keith C. C. Chan |
ISMIS | 1 |
| 2018 | Clustering in Networks with Multi-Modality AttributesabstractNetwork clustering is one of the most significant tasks of network analytics. To discover network clusters, there have been many approaches proposed, utilizing network topology, or node attributes. However, there are no effective approaches that are able to discover clusters in the network with multiple modalities of attributes. In this paper, we propose a novel clustering model, called CNMMA, to discover network clusters using edge structure, and multi-modality attributes associated with vertices. Assuming edge structure, and node attributes are generated by corresponding low dimensional latent spaces (matrices), CNMMA can learn an optimal latent matrix representing the cluster membership for each vertex in the network. Besides, CNMMA makes use of an effective method to regulate the latent spaces w.r.t. edge structure and node attributes so that those vertices sharing similar edges and modality-wise attributes are more possible to be assigned with the same cluster labels. CNMMA has been tested with several real-world networks, which contain multiple modalities of node attributes, and has been compared with state-of-the-art approaches to network clustering. The experimental results show that CNMMA outperforms most approaches in most datasets. The clusters discovered by CNMMA are better matched with the ground truth. Tiantian He 0001, Keith C. C. Chan, Libin Yang |
WI | 1 |
| 2018 | An eigenvector based center selection for fast training scheme of RBFNN
Yan-Xing Hu, Jane You, James Nga-Kwok Liu, Tiantian He 0001 |
Inf. Sci. | 4 |
| 2018 | Evolutionary Graph Clustering for Protein Complex IdentificationabstractThis paper presents a graph clustering algorithm, called EGCPI, to discover protein complexes in protein-protein interaction (PPI) networks. In performing its task, EGCPI takes into consideration both network topologies and attributes of interacting proteins, both of which have been shown to be important for protein complex discovery. EGCPI formulates the problem as an optimization problem and tackles it with evolutionary clustering. Given a PPI network, EGCPI first annotates each protein with corresponding attributes that are provided in Gene Ontology database. It then adopts a similarity measure to evaluate how similar the connected proteins are taking into consideration the network topology. Given this measure, EGCPI then discovers a number of graph clusters within which proteins are densely connected, based on an evolutionary strategy. At last, EGCPI identifies protein complexes in each discovered cluster based on the homogeneity of attributes performed by pairwise proteins. EGCPI has been tested with several real data sets and the experimental results show EGCPI is very effective on protein complex discovery, and the evolutionary clustering is helpful to identify protein complexes in PPI networks. The software of EGCPI can be downloaded via: https://github.com/hetiantian1985/EGCPI. Tiantian He 0001, Keith C. C. Chan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | MISAGA: An Algorithm for Mining Interesting Subgraphs in Attributed GraphsabstractAn attributed graph contains vertices that are associated with a set of attribute values. Mining clusters or communities, which are interesting subgraphs in the attributed graph is one of the most important tasks of graph analytics. Many problems can be defined as the mining of interesting subgraphs in attributed graphs. Algorithms that discover subgraphs based on predefined topologies cannot be used to tackle these problems. To discover interesting subgraphs in the attributed graph, we propose an algorithm called mining interesting subgraphs in attributed graph algorithm (MISAGA). MISAGA performs its tasks by first using a probabilistic measure to determine whether the strength of association between a pair of attribute values is strong enough to be interesting. Given the interesting pairs of attribute values, then the degree of association is computed for each pair of vertices using an information theoretic measure. Based on the edge structure and degree of association between each pair of vertices, MISAGA identifies interesting subgraphs by formulating it as a constrained optimization problem and solves it by identifying the optimal affiliation of subgraphs for the vertices in the attributed graph. MISAGA has been tested with several large-sized real graphs and is found to be potentially very useful for various applications. Tiantian He 0001, Keith C. C. Chan |
IEEE Trans. Cybern. | 1 |
| 2018 | Discovering Fuzzy Structural Patterns for Graph AnalyticsabstractMany real-world datasets can be represented as attributed graphs that contain vertices, each of which is associated with a set of attribute values. Discovering clusters, or communities, which are structural patterns in these graphs, are one of the most important tasks in graph analysis. To perform the task, a number of algorithms have been proposed. Some of them detect clusters of particular topological properties, whereas some others discover them mainly based on attribute information. Also, most of the algorithms discover disjoint clusters only. As a result, they may not be able to detect more meaningful clusters hidden in the attributed graph. To do so more effectively, we propose an algorithm, called FSPGA, to discover fuzzy structural patterns for graph analytics. FSPGA performs the task of cluster discovery as a fuzzy-constrained optimization problem, which takes into consideration both the graph topology and attribute values. FSPGA has been tested with both synthetic and real-world graph datasets and is found to be efficient and effective at detecting clusters in attributed graphs. FSPGA is a promising fuzzy algorithm for structural pattern detection in attributed graphs. Tiantian He 0001, Keith C. C. Chan |
IEEE Trans. Fuzzy Syst. | 1 |
| 2017 | Deep Fusion of Multiple Networks for Learning Latent Social CommunitiesabstractThe rapid development of techniques results in a growing diversity of social network data which require for analysis. Therefore, the deeper understanding of latent knowledge representing the social network data needs learning by combining the insights obtained from multiple, diverse networks carrying heterogeneous information featuring the interrelationship between vertices. In this manuscript, we propose a novel deepmodel- based approach to learn latent structural representation from multi-domain social network data. The algorithm, which we call Deep Multiple Networks Fusion (DMNF), is able to discover an aggregated deep representation, by taking into consideration multiple networks, which represent heterogeneous information carried by the social network data. To perform the task, DMNF first constructs a network representing the total degree of interrelationship between pairwise vertices by utilizing a fusion method to compute such degree taking into consideration heterogeneous information embedded in the network data, e.g., node connection, and attribute relativity. Given the fused network data, DMNF attempts to learn the latent network representation making use of a deep neural network model. Such learned representation is able to reveal the latent structure, e.g., social communities, and clusters in the social network. DMNF has been tested with two sets of real social network data and compared with several prevalent approaches to network community detection. The experimental results show that the latent representation found by DMNF may match well with the ground-truth communities and DMNF is able to outperform the state-of-the-art approaches to detecting social network communities. Pengwei Hu 0001, Tiantian He 0001, Keith C. C. Chan, Henry Leung 0001 |
ICTAI | 2 |
| 2014 | Evolutionary community detection in social networksabstractAs people that share common characteristics and interests tend to communicate with each other more frequently, they form communities within social networks. Several methods have been developed to discover such communities based on topological metrics. These methods have been used to successfully discover communities that are relatively large, but for communities characterized by members interacting more frequently with each other rather than interacting with many others, we propose here an effective method which is based on the use of an evolutionary algorithm (EA) called ECDA. Given a social network represented as a graph, unlike existing approaches, ECDA considers both topological metrics of the graph and the attributes of the vertices and edges when detecting for communities in the network. It performs its task by formulating the community detection problem as an optimization problem. By computing a measure of statistical significance for each attribute of the vertices, ECDA looks for communities in a network that have maximal connection significance within a community and minimal significance between any two communities. With such a strategy, ECDA partitions a network into different communities consisting of members with similar attributes within and different attributes without. Unlike other EAs, ECDA adopts a reproduction process consisting of special crossover and mutation operators, called Self-Evolution, to speed up the evolutionary process. ECDA has been tested with several real datasets and its performance is found to be very promising. Tiantian He 0001, Keith C. C. Chan |
IEEE Congress on Evolutionary Computation | 1 |