Yike Guo

dblp:g/YikeGuo · also Yi-Ke Guo · DBLP profile ↗
← Back
24ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0002-3075-2161ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 11 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 5Big Data, Cloud & Distributed Data Systems · 4Information Retrieval & Web Search · 2Database Systems & Data Management · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Spatiotemporal Graph Learning with Direct Volumetric Information Passing and Feature Enhancement
abstract
Data-driven learning of physical systems has attracted significant attention, where many neural models have been developed. In particular, mesh-based graph neural networks (GNNs) have demonstrated considerable potential in modeling spatiotemporal dynamics across arbitrary geometric domains. However, the existing node-edge message-passing and aggregation mechanism in GNNs limits the representation learning capability. In this paper, we propose a dual-module framework, Cell-embedded and Feature-enhanced Graph Neural Network (CeFeGNN), for learning spatiotemporal dynamics. Specifically, we embed learnable cell attributions to the common node-edge message passing process, thereby better capturing the spatial dependency of regional features. Such a strategy essentially upgrades the local aggregation scheme from first order (e.g., from edge to node) to a higher order (e.g., from volume and edge to node), which takes advantage of volumetric information in message passing. Meanwhile, a novel feature-enhanced block is designed to further improve the model's performance and alleviate the over-smoothing problem. Extensive experiments on various PDE systems and a real-world dataset demonstrate that CeFeGNN achieves superior performance compared with other baselines.
Yuan Mi, Qi Wang 0123, Xueqin Hu, Yike Guo, Ji-Rong Wen, Yang Liu 0130, Hao Sun 0002
KDD (1)4
2025 Conservation-informed Graph Learning for Spatiotemporal Dynamics Prediction
abstract
Data-centric methods have shown great potential in understanding and predicting spatiotemporal dynamics, enabling better design and control of the object system. However, deep learning models often lack interpretability, fail to obey intrinsic physics, and struggle to cope with the various domains. While geometry-based methods, e.g., graph neural networks (GNNs), have been proposed to further tackle these challenges, they still need to find the implicit physical laws from large datasets and rely excessively on rich labeled data. In this paper, we herein introduce the conservation-informed GNN (CiGNN), an end-to-end explainable learning framework, to learn spatiotemporal dynamics based on limited training data. The network is designed to conform to the general conservation law via symmetry, where conservative and non-conservative information passes over a multiscale space enhanced by a latent temporal marching strategy. The efficacy of our model has been verified in various spatiotemporal systems based on synthetic and real-world datasets, showing superiority over baseline models. Results demonstrate that CiGNN exhibits remarkable accuracy and generalizability, and is readily applicable to learning for prediction of various spatiotemporal dynamics in a spatial domain with complex geometry.
Yuan Mi, Pu Ren, Hongteng Xu, Hongsheng Liu 0002, Zidong Wang 0010, Yike Guo, Ji-Rong Wen, Hao Sun 0002, Yang Liu 0005
KDD (1)6
2024 Hierarchical Linear Symbolized Tree-Structured Neural Processes
abstract
Traditional Neural Processes (NPs) and their variants aim to learn relationships between context sample points but do not consider multi-level information, resulting in a limited ability to learn complex distributions.This paper draws inspiration from features such as the hierarchical nature and interpretability of tree-like structures.This paper proposes a Hierarchical Linear Symbolized Treestructured Neural Processes (HLNPs) architecture.This framework utilizes variables to build a top-down hierarchical linear symbolized tree-structured network architecture, enhancing positional representation information in a hierarchical manner along the deterministic path.In the latent distribution, the hierarchical linear symbolized tree-structured network approximates functions discretely through a layered approach.By decomposing the latent complex distribution into several simpler sub-problems using sum and product symbols, the upper bound of optimization is thereby increased.The tree structure discretizes variables to capture model uncertainty in the form of entropy.This approach also imparts a causal effect to the HLNPs model.Finally, we demonstrate the effectiveness of the HLNPs models for 1D data, Bayesian optimization, and 2D data.
Jinyang Tai, Yike Guo
KDD2
2024 GNN-MgrPool: Enhanced graph neural networks with multi-granularity pooling for graph classification
Haichao Sun, Guoyin Wang 0001, Qun Liu 0005, Yike Guo
Inf. Sci.4
2022 FTAP: Feature transferring autonomous machine learning pipeline
Xing Wu 0001, Cheng Chen 0075, Mingyu Zhong, Jianjia Wang, Quan Qian, Junfeng Yao, Yike Guo
Inf. Sci.9
2021 Label Dependent Attention Model for Disease Risk Prediction Using Multimodal Electronic Health Records
abstract
Disease risk prediction has attracted increasing attention in the field of modern healthcare, especially with the latest advances in artificial intelligence (AI). Electronic health records (EHRs), which contain heterogeneous patient information, are widely used in disease risk prediction tasks. One challenge of applying AI models for risk prediction lies in generating interpretable evidence to support the prediction results while retaining the prediction ability. In order to address this problem, we propose the method of jointly embedding words and labels whereby attention modules learn the weights of words from medical notes according to their relevance to the names of risk prediction labels. This approach boosts interpretability by employing an attention mechanism and including the names of prediction tasks in the model. However, its application is only limited to the handling of textual inputs such as medical notes. In this paper, we propose a label dependent attention model LDAM to 1) improve the interpretability by exploiting Clinical-BERT (a biomedical language model pre-trained on a large clinical corpus) to encode biomedically meaningful features and labels jointly; 2) extend the idea of joint embedding to the processing of timeseries data, and develop a multi-modal learning framework for integrating heterogeneous information from medical notes and time-series health status indicators. To demonstrate our method, we apply LDAM to the MIMIC-III dataset to predict different disease risks. We evaluate our method both quantitatively and qualitatively. Specifically, the predictive power of LDAM will be shown, and case studies will be carried out to illustrate its interpretability.
Qing Yin, Yunya Song, Yike Guo, Xian Yang 0001
ICDM4
2020 The assessment of small bowel motility with attentive deformable neural network
Xing Wu 0001, Mingyu Zhong, Yike Guo, Hamido Fujita
Inf. Sci.3
2019 An Information-Theoretical Framework for Cluster Ensemble
abstract
Cluster ensemble is a very important tool that aggregates several base clusterings to generate a single output clustering with improved robustness and stability. However, the quality of the final clustering is often affected by uncertainties on the generation and integration of base clusterings. In this paper, we develop an information-theoretical framework which makes an effort to obtain a final clustering with high consensus on both the original data set and the base clustering set by minimizing the two uncertainties of cluster ensemble. In this framework, we provide a weighted consensus measure based on information entropy to evaluate the quality of a clustering, the similarity between clusters and the similarity between objects. Based on the measure, we propose three weighted cluster ensemble algorithms with different ensemble strategies in the framework, including the weighted feature consensus algorithm, the weighted relabeling consensus algorithm and the weighted pairwise-similarity consensus algorithm. In the experimental analysis, we compare the proposed algorithms with other existing clustering ensemble algorithms on several data sets. The comparison results illustrate the proposed algorithms are very effective and robust.
Liang Bai 0001, Jiye Liang, Hangyuan Du, Yike Guo
IEEE Trans. Knowl. Data Eng.4
2018 Deep Sequence Learning with Auxiliary Information for Traffic Prediction
abstract
Predicting traffic conditions from online route queries is a challenging task as there are many complicated interactions over the roads and crowds involved. In this paper, we intend to improve traffic prediction by appropriate integration of three kinds of implicit but essential factors encoded in auxiliary information. We do this within an encoder-decoder sequence learning framework that integrates the following data: 1) offline geographical and social attributes. For example, the geographical structure of roads or public social events such as national celebrations; 2) road intersection information. In general, traffic congestion occurs at major junctions; 3) online crowd queries. For example, when many online queries issued for the same destination due to a public performance, the traffic around the destination will potentially become heavier at this location after a while. Qualitative and quantitative experiments on a real-world dataset from Baidu have demonstrated the effectiveness of our framework.
Binbing Liao, Jingqing Zhang, Chao Wu 0001, Douglas McIlwraith, Tong Chen 0006, Shengwen Yang, Yike Guo, Fei Wu 0001
KDD7
2018 Crowdsourcing with online quantitative design analysis
abstract
Design is a balancing act between people’s competing concerns, design options and design performance. Recently collecting data on such concerns such as sustainability or aesthetics has become possible through online crowdsourcing, particularly in 3d. However, such systems rarely present more than a single design alternative or allow users to change the design and seldom provide quantitative design analysis to gauge design performance. This precludes a more participatory approach including a wider audience and their insight in the design process. To improve the design process we propose a system to assist the design team in exploring the balance of concerns, design options and their performance. We augment a 3d visualisation crowdsourcing environment with quantitative on-demand assessment of design variants run in the cloud. This enables crowdsourced exploration of the design space and its performance. Automated participant tracking and explicit submitted feedback on design options are collated and presented to aid the design team in balancing the demands of urban master planning. We report application of this system to an urban masterplan with Arup.
David Birch, Alvise Simondetti, Yike Guo
Adv. Eng. Informatics3
2017 eTRIKS analytical environment: A modular high performance framework for medical data analysis
abstract
Translational research is quickly becoming a science driven by big data. Improving patient care, developing personalized therapies and new drugs depend increasingly on an organization's ability to rapidly and intelligently leverage complex molecular and clinical data from a variety of large-scale partner and public sources. As analysing these large-scale datasets becomes computationally increasingly expensive, traditional analytical engines are struggling to provide a timely answer to the questions that biomedical scientists are asking. Designing such a framework is developing for a moving target as the very nature of biomedical research based on big data requires an environment capable of adapting quickly and efficiently in response to evolving questions. The resulting framework consequently must be scalable in face of large amounts of data, flexible, efficient and resilient to failure. In this paper we design the eTRIKS Analytical Environment (eAE), a scalable and modular framework for the efficient management and analysis of large scale medical data, in particular the massive amounts of data produced by high-throughput technologies. We particularly discuss how we design the eAE as a modular and efficient framework enabling us to add new components or replace old ones easily. We further elaborate on its use for a set of challenging big data use cases in medicine and drug discovery.
Axel Oehmichen, Florian Guitton, Kai Sun 0005, Jean Grizet, Thomas Heinis, Yike Guo
IEEE BigData6
2017 Fast graph clustering with a new description model for community detection
Liang Bai 0001, Xueqi Cheng 0001, Jiye Liang, Yike Guo
Inf. Sci.4
2017 ACM TIST Special Issue on Data-Driven Intelligence for Wireless Networking
abstract
No abstract available.
Wenwu Zhu 0001, Jean C. Walrand, Yike Guo, Zhi Wang 0001
ACM Trans. Intell. Syst. Technol.3
2013 Elastic algorithms for guaranteeing quality monotonicity in big data mining
abstract
When mining large data volumes in big data applications users are typically willing to use algorithms that produce acceptable approximate results satisfying the given resource and time constraints. Two key challenges arise when designing such algorithms. The first relates to reasoning about tradeoffs between the quality of data mining output, e.g. prediction accuracy for classification tasks and available resource and time budgets. The second is organizing the computation of the algorithm to guarantee producing better quality of results as more budget is used. Little work has addressed these two challenges together in a generic way. In this paper, we propose a novel framework for developing elastic big data mining algorithms. Based on Shannon's entropy, an information-theoretic approach is introduced to reason about how result quality is affected by the allocated budget. This is then used to guide the development of algorithms that adapt to the available time budgets while guaranteeing producing better quality results as more budgets are used. We demonstrate the application of the framework by developing elastic k-Nearest Neighbour (kNN) classification and collaborative filtering (CF) recommendation algorithms as two examples. The core of both elastic algorithms is to use a naïve kNN classification or CF algorithm over R-tree data structures that successively approximate the entire datasets. Experimental evaluation was performed using prediction accuracy as quality metric on real datasets. The results show that elastic mining algorithms indeed produce results with consistent increase in observable qualities, i.e., prediction accuracy, in practice.
Rui Han 0001, Lei Nie 0008, Moustafa Ghanem, Yike Guo
IEEE BigData4
2013 Building a generic platform for big sensor data application
abstract
The drive toward smart cities alongside the rising adoption of personal sensors is leading to a torrent of sensor data. While systems exist for storing and managing sensor data, the real value of such data is the insight which can be generated from it. However there is currently no platform which enables sensor data to be taken from collection, through use in models to produce useful data products. The architecture of such a platform is a current research question in the field of Big Data and Smart Cities. In this paper we explore five key challenges in this field and provide a response through a sensor data platform “Concinnity” which can take sensor data from collection to final product via a data repository and workflow system. This will enable rapid development of applications built on sensor data using data fusion and the integration and composition of models to form novel workflows. We summarize the key features of our approach, exploring how it enables value to be derived from sensor data efficiently.
Chun-Hsiang Lee, David Birch, Chao Wu 0001, Dilshan Silva, Orestis Tsinalis, Yang Li 0003, Shulin Yan, Moustafa Ghanem, Yike Guo
IEEE BigData9
2013 Enhanced user data privacy with pay-by-data model
abstract
Personal data collection is becoming pervasive these days, these data has the risk of being abused by current application and application marketplace model, because only the price of application is explicitly indicated without clear agreement on usage of data, and the granularity of data access authentication is not enough to protect users privacy. In this short paper, we propose a new model of user data privacy. Data usage of the application is explicitly shown, and controlled by an authentication service, to protect users from the abuse of their data, especially in mobile application.
Chao Wu 0001, Yike Guo
IEEE BigData2
2005 Using dragpushing to refine centroid text classifiers
abstract
We present a novel algorithm, DragPushing, for automatic text classification. Using a training data set, the algorithm first calculates the prototype vectors, or centroids, for each of the available document classes. Using misclassified examples, it then iteratively refines these centroids; by dragging the centroid of a correct class towards a misclassified example and in the same time pushing the centroid of an incorrect class away from the misclassified example. The algorithm is simple to implement and is computationally very efficient. Evaluation experiments conducted on two benchmark collections show that its classification accuracy is comparable to that of more complex methods, such as support vector machines (SVM).
Songbo Tan, Xueqi Cheng 0001, Bin Wang 0004, Moustafa Ghanem, Yike Guo
SIGIR6
2003 InfoGrid: providing information integration for knowledge discovery
Nikolaos Giannadakis, Anthony Rowe 0002, Moustafa Ghanem, Yike Guo
Inf. Sci.4
2002 Discovery net: towards a grid of knowledge discovery
abstract
This paper provides a blueprint for constructing collaborative and distributed knowledge discovery systems within Grid-based computing environments. The need for such systems is driven by the quest for sharing knowledge, information and computing resources within the boundaries of single large distributed organisations or within complex Virtual Organisations (VO) created to tackle specific projects. The proposed architecture is built on top of a resource federation management layer and is composed of a set of different resources. We show how this architecture will behave during a typical KDD process design and deployment, how it enables the execution of complex and distributed data mining tasks with high performance and how it provides a community of e-scientists with means to collaborate, retrieve and reuse both KDD algorithms, discovery processes and knowledge in a visual analytical environment.
Vasa Curcin, Moustafa Ghanem, Yike Guo, Anthony Rowe 0002, Jameel Syed, Patrick Wendel
KDD3
2000 New paradigms in information visualization
abstract
We present three new visualization front-ends that aid navigation through the set of documents returned by a search engine (hit documents). We cluster the hit documents to visually group these documents and label the groups with related words. The different front-ends cater for different user needs, but all can browse cluster information as well as drilling up or down in one or more clusters and refining the search using one or more of the suggested related keywords.
Peter Au, Matthew Carey, Shalini Sewraz, Yike Guo, Stefan M. Rüger
SIGIR4
1999 Probing Knowledge in Distributed Data Mining
Yike Guo, Janjao Sutiwaraphun
PAKDD1
1999 Editorial
Yike Guo, Robert L. Grossman
Data Min. Knowl. Discov.1
1997 Parallel Induction Algorithms for Data Mining
John Darlington, Yike Guo, Janjao Sutiwaraphun, Hing Wing To
IDA2
1997 Large Scale Data Mining: Challenges and Responses
Jaturon Chattratichat, John Darlington, Moustafa Ghanem, Yike Guo, Harald Frank Hüning, Janjao Sutiwaraphun, Hing Wing To
KDD4