Stanley Kok

dblp:74/2264 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-8624-086XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Knowledge representation and reasoning · 25% Graph learning · 23% Probabilistic and Bayesian machine learning · 18%
Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 40% Knowledge graphs · 35% Data mining · 14%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
1.012026
Practical Global and Local Bounds in Gaussian Process Regression via Chaining · AAAI 2026
Machine learning › Trustworthy machine learning
uncertainty estimation
1.012026
Practical Global and Local Bounds in Gaussian Process Regression via Chaining · AAAI 2026
Machine learning › Graph learning
graph clustering
0.912025
Unsupervised Multiple Kernel Learning for Graphs via Ordinality Preservation · ICLR 2025
Machine learning › Graph learning
graph kernel
0.912025
Unsupervised Multiple Kernel Learning for Graphs via Ordinality Preservation · ICLR 2025
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
multiple kernel learning
0.912025
Unsupervised Multiple Kernel Learning for Graphs via Ordinality Preservation · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning
0.662016
Unifying Logical and Statistical AI · LICS 2016
Learning Markov Logic Networks Using Structural Motifs · ICML 2010
Learning Markov logic network structure via hypergraph lifting · ICML 2009
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
markov logic networks
0.652016
Unifying Logical and Statistical AI · LICS 2016
Learning Markov Logic Networks Using Structural Motifs · ICML 2010
Learning Markov logic network structure via hypergraph lifting · ICML 2009
Machine learning › Learning theory
generalization bounds
0.622026
Practical Global and Local Bounds in Gaussian Process Regression via Chaining · AAAI 2026
Unsupervised Multiple Kernel Learning for Graphs via Ordinality Preservation · ICLR 2025
Knowledge graphs
knowledge graph embedding
0.512021
BiQUE: Biquaternionic Embeddings of Knowledge Graphs · EMNLP (1) 2021
Query processing and optimization › query optimization
cost-based optimization
0.522016
CrowdOp: Query optimization for declarative crowdsourcing systems · ICDE 2016
CrowdOp: Query Optimization for Declarative Crowdsourcing Systems · IEEE Trans. Knowl. Data Eng. 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.332010
Learning Markov Logic Networks Using Structural Motifs · ICML 2010
Learning Markov logic network structure via hypergraph lifting · ICML 2009
Learning the structure of Markov logic networks · ICML 2005
Data integration and cleaning
entity resolution
0.112011
Collective graph identification · KDD 2011
Data mining › structured data mining
graph mining
0.112011
Collective graph identification · KDD 2011
Knowledge graphs
link prediction
0.112011
Collective graph identification · KDD 2011
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.112016
Unifying Logical and Statistical AI · LICS 2016
Data integration and cleaning
crowdsourced data processing
0.112015
CrowdOp: Query Optimization for Declarative Crowdsourcing Systems · IEEE Trans. Knowl. Data Eng. 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning
probabilistic logic
0.112006
Unifying Logical and Statistical AI · AAAI 2006
Logic in computer science
first-order logic
0.112006
Unifying Logical and Statistical AI · AAAI 2006
Computational complexity › complexity of reasoning
tractable fragments
0.112006
Unifying Logical and Statistical AI · AAAI 2006
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
inductive logic programming
0.112005
Learning the structure of Markov logic networks · ICML 2005
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation
0.012004
Web-scale information extraction in knowitall: (preliminary results) · WWW 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
fact extraction
0.012004
Web-scale information extraction in knowitall: (preliminary results) · WWW 2004
Natural language and speech › Information extraction and text analysis › knowledge discovery from text
knowledge base population
0.012004
Web-scale information extraction in knowitall: (preliminary results) · WWW 2004
Natural language and speech › Information extraction and text analysis › web information extraction
web-scale information extraction
0.012004
Web-scale information extraction in knowitall: (preliminary results) · WWW 2004
Information retrieval › query processing
web query processing
0.012004
Web-scale information extraction in knowitall: (preliminary results) · WWW 2004

Methods — techniques the papers use, named apart from their topics

kernel-specific refinement · 1.0chaining · 1.0biquaternion · 1.0probability simplex · 0.9ordinality preservation · 0.9multiple kernel learning · 0.9geometric transformations · 0.5geometric transformation · 0.5pseudo-likelihood · 0.4query plan generation · 0.2markov chain monte carlo · 0.2inductive logic programming · 0.2cost-latency trade-off · 0.2simulation · 0.2crowdsourcing experiments · 0.2cost-based optimization · 0.2probabilistic inference · 0.1collective classification · 0.1
YearPublicationVenuePosition
2026 Practical Global and Local Bounds in Gaussian Process Regression via Chaining
abstract
Gaussian process regression (GPR) is a popular nonparametric Bayesian method that provides predictive uncertainty estimates and is widely used in safety-critical applications. While prior research has introduced various uncertainty bounds, most existing approaches require access to specific input features, and rely on posterior mean and variance estimates or the tuning of hyperparameters. These limitations hinder robustness and fail to capture the model’s global behavior in expectation. To address these limitations, we propose a chaining-based framework for estimating upper and lower bounds on the expected extreme values over unseen data, without requiring access to specific input features. We provide kernel-specific refinements for commonly used kernels such as RBF and Matérn, in which our bounds are tighter than generic constructions. We further improve numerical tightness by avoiding analytical relaxations. In addition to global estimation, we also develop a novel method for local uncertainty quantification at specified inputs. This approach leverages chaining geometry through partition diameters, adapting to local structures without relying on posterior variance scaling. Our experimental results validate the theoretical findings and demonstrate that our method outperforms existing approaches on both synthetic and real-world datasets.
Junyi Liu 0003, Stanley Kok
AAAI2
2025 Unsupervised Multiple Kernel Learning for Graphs via Ordinality Preservation
abstract
Learning effective graph similarities is crucial for tasks like clustering, yet selecting the optimal kernel to evaluate such similarities in unsupervised settings remains a major challenge. Despite the development of various graph kernels, determining the most appropriate one for a specific task is particularly difficult in the absence of labeled data. Existing methods often struggle to handle the complex structure of graph data and rely on heuristic approaches that fail to adequately capture the global relationships between graphs. To overcome these limitations, we propose Unsupervised Multiple Kernel Learning for Graphs (UMKL-G), a model that combines multiple graph kernels without requiring labels or predefined local neighbors. Our approach preserves the topology of the data by maintaining ordinal relationships among graphs through a probability simplex, allowing for a unified and adaptive kernel learning process. We provide theoretical guarantees on the stability, robustness, and generalization of our method. Empirical results demonstrate that UMKL-G outperforms individual kernels and other state-of-the-art methods, offering a robust solution for unsupervised graph analysis.
Stanley Kok
ICLR2
2023 Towards Integration of Discriminability and Robustness for Document-Level Relation Extraction
abstract
Document-level relation extraction (DocRE)predicts relations for entity pairs that rely on long-range context-dependent reasoning in a document.As a typical multi-label classification problem, DocRE faces the challenge of effectively distinguishing a small set of positive relations from the majority of negative ones.This challenge becomes even more difficult to overcome when there exists a significant number of annotation errors in the dataset.In this work, we aim to achieve better integration of both the discriminability and robustness for the DocRE problem.Specifically, we first design an effective loss function to endow high discriminability to both probabilistic outputs and internal representations.We innovatively customize entropy minimization and supervised contrastive learning for the challenging multi-label and long-tailed learning problems.To ameliorate the impact of label errors, we equipped our method with a novel negative label sampling strategy to strengthen the model robustness.In addition, we introduce two new data regimes to mimic more realistic scenarios with annotation errors and evaluate our sampling strategy.Experimental results verify the effectiveness of each component and show that our method achieves new state-ofthe-art results on the DocRED dataset, its recently cleaned version, Re-DocRED, and the proposed data regimes. 1
Stanley Kok, Lidong Bing
EACL2
2021 BiQUE: Biquaternionic Embeddings of Knowledge Graphs
abstract
Knowledge graph embeddings (KGEs) compactly encode multi-relational knowledge graphs (KGs).Existing KGE models rely on geometric operations to model relational patterns.Euclidean (circular) rotation is useful for modeling patterns such as symmetry, but cannot represent hierarchical semantics.In contrast, hyperbolic models are effective at modeling hierarchical relations, but do not perform as well on patterns on which circular rotation excels.It is crucial for KGE models to unify multiple geometric transformations so as to fully cover the multifarious relations in KGs.To do so, we propose BiQUE, a novel model that employs biquaternions to integrate multiple geometric transformations, viz., scaling, translation, Euclidean rotation, and hyperbolic rotation.BiQUE makes the best tradeoffs among geometric operators during training, picking the best one (or their best combination) for each relation.Experiments on five datasets show BiQUE's effectiveness.
Stanley Kok
EMNLP (1)2
2021 Referent graph embedding model for name entity recognition of Chinese car reviews
Zhao Fang, Qiang Zhang 0010, Stanley Kok, Anning Wang, Shanlin Yang
Knowl. Based Syst.3
2016 CrowdOp: Query optimization for declarative crowdsourcing systems
abstract
We propose CROWDOP, a cost-based query optimization approach for declarative crowdsourcing systems. CROWDOP considers both cost and latency in the query optimization objectives and generates query plans that provide a good balance between the cost and latency. We develop efficient algorithms in CROWDOP for optimizing three types of queries: selection, join and complex selection-join queries. We validate our approach via extensive experiments by simulation as well as with the real crowd on Amazon Mechanical Turk.
Ju Fan, Meihui Zhang 0001, Stanley Kok, Meiyu Lu, Beng Chin Ooi
ICDE3
2016 Unifying Logical and Statistical AI
abstract
Intelligent agents must be able to handle the complexity and uncertainty of the real world. Logical AI has focused mainly on the former, and statistical AI on the latter. Markov logic combines the two by attaching weights to first-order formulas and viewing them as templates for features of Markov networks. Inference algorithms for Markov logic draw on ideas from satisfiability, Markov chain Monte Carlo and knowledge-based model construction. Learning algorithms are based on the voted perceptron, pseudo-likelihood and inductive logic programming. Markov logic has been successfully applied to a wide variety of problems in natural language understanding, vision, computational biology, social networks and others, and is the basis of the open-source Alchemy system.
Pedro M. Domingos, Daniel Lowd, Stanley Kok, Aniruddh Nath, Hoifung Poon, Matthew Richardson, Parag Singla
LICS3
2015 CrowdOp: Query Optimization for Declarative Crowdsourcing Systems
abstract
We study the query optimization problem in declarative crowdsourcing systems. Declarative crowdsourcing is designed to hide the complexities and relieve the user of the burden of dealing with the crowd. The user is only required to submit an SQL-like query and the system takes the responsibility of compiling the query, generating the execution plan and evaluating in the crowdsourcing marketplace. A given query can have many alternative execution plans and the difference in crowdsourcing cost between the best and the worst plans may be several orders of magnitude. Therefore, as in relational database systems, query optimization is important to crowdsourcing systems that provide declarative query interfaces. In this paper, we proposeCrowdOp, a cost-based query optimization approach for declarative crowdsourcing systems.CrowdOpconsiders both cost and latency in query optimization objectives and generates query plans that provide a good balance between the cost and latency. We develop efficient algorithms in theCrowdOpfor optimizing three types of queries: selection queries, join queries, and complex selection-join queries. We validate our approach via extensive experiments by simulation as well as with the real crowd on Amazon Mechanical Turk.
Ju Fan, Meihui Zhang 0001, Stanley Kok, Meiyu Lu, Beng Chin Ooi
IEEE Trans. Knowl. Data Eng.3
2014 Language modeling with sum-product networks
abstract
Sum product networks (SPNs) are a new class of deep probabilistic models. They can contain multiple hidden layers while keeping their inference and training times tractable. An SPN consists of interleaving layers of sum nodes and product nodes. A sum node can be interpreted as a hidden variable, and a product node can be viewed as a feature capturing rich interactions among an SPN’s inputs. We show that the ability of SPN to use hidden layers to model complex dependencies among words, and its tractable inference and learning times, make it a suitable framework for a language model. Even though SPNs have been applied to a variety of vision problems [1, 2], we are the first to use it for language modeling. Our empirical comparisons with six previous language models indicate that our SPN has superior performance.
Wei-Chen Cheng, Stanley Kok, Hoai Vu Pham, Hai Leong Chieu, Kian Ming A. Chai
INTERSPEECH2
2011 Collective graph identification
abstract
Data describing networks (communication networks, transaction networks, disease transmission networks, collaboration networks, etc.) is becoming increasingly ubiquitous. While this observational data is useful, it often only hints at the actual underlying social or technological structures which give rise to the interactions. For example, an email communication network provides useful insight but is not the same as the "real" social network among individuals. In this paper, we introduce the problem of graph identification, i.e., the discovery of the true graph structure underlying an observed network. We cast the problem as a probabilistic inference task, in which we must infer the nodes, edges, and node labels of a hidden graph, based on evidence provided by the observed network. This in turn corresponds to the problems of performing entity resolution, link prediction, and node labeling to infer the hidden graph. While each of these problems have been studied separately, they have never been considered together as a coherent task. We present a simple yet novel approach to address all three problems simultaneously. Our approach, called C3, consists of Coupled Collective Classifiers that are iteratively applied to propagate information among solutions to the problems. We empirically demonstrate that C3 is superior, in terms of both predictive accuracy and runtime, to state-of-the-art probabilistic approaches on three real-world problems.
Galileo Namata, Stanley Kok, Lise Getoor
KDD2
2010 Learning Markov Logic Networks Using Structural Motifs
Stanley Kok, Pedro M. Domingos
ICML1
2010 Hitting the Right Paraphrases in Good Time
Stanley Kok, Chris Brockett
HLT-NAACL1
2009 Learning Markov logic network structure via hypergraph lifting
abstract
Markov logic networks (MLNs) combine logic and probability by attaching weights to first-order clauses, and viewing these as templates for features of Markov networks. Learning MLN structure from a relational database involves learning the clauses and weights. The state-of-the-art MLN structure learners all involve some element of greedily generating candidate clauses, and are susceptible to local optima. To address this problem, we present an approach that directly utilizes the data in constructing candidates. A relational database can be viewed as a hypergraph with constants as nodes and relations as hyperedges. We find paths of true ground atoms in the hypergraph that are connected via their arguments. To make this tractable (there are exponentially many paths in the hypergraph), we lift the hypergraph by jointly clustering the constants to form higherlevel concepts, and find paths in it. We variabilize the ground atoms in each path, and use them to form clauses, which are evaluated using a pseudo-likelihood measure. In our experiments on three real-world datasets, we find that our algorithm outperforms the state-of-the-art approaches.
Stanley Kok, Pedro M. Domingos
ICML1
2008 Extracting Semantic Networks from Text Via Relational Clustering
Stanley Kok, Pedro M. Domingos
ECML/PKDD (1)1
2007 Statistical predicate invention
abstract
We propose statistical predicate invention as a key problem for statistical relational learning. SPI is the problem of discovering new concepts, properties and relations in structured data, and generalizes hidden variable discovery in statistical models and predicate invention in ILP. We propose an initial model for SPI based on second-order Markov logic, in which predicates as well as arguments can be variables, and the domain of discourse is not fully known in advance. Our approach iteratively refines clusters of symbols based on the clusters of symbols they appear in atoms with (e.g., it clusters relations by the clusters of the objects they relate). Since different clusterings are better for predicting different subsets of the atoms, we allow multiple cross-cutting clusterings. We show that this approach outperforms Markov logic structure learning and the recently introduced infinite relational model on a number of relational datasets.
Stanley Kok, Pedro M. Domingos
ICML1
2006 Unifying Logical and Statistical AI
Pedro M. Domingos, Stanley Kok, Hoifung Poon, Matthew Richardson, Parag Singla
AAAI2
2005 Learning the structure of Markov logic networks
abstract
Markov logic networks (MLNs) combine logic and probability by attaching weights to first-order clauses, and viewing these as templates for features of Markov networks. In this paper we develop an algorithm for learning the structure of MLNs from relational databases, combining ideas from inductive logic programming (ILP) and feature induction in Markov networks. The algorithm performs a beam or shortest-first search of the space of clauses, guided by a weighted pseudo-likelihood measure. This requires computing the optimal weights for each candidate structure, but we show how this can be done efficiently. The algorithm can be used to learn an MLN from scratch, or to refine an existing knowledge base. We have applied it in two real-world domains, and found that it outperforms using off-the-shelf ILP systems to learn the MLN structure, as well as pure ILP, purely probabilistic and purely knowledge-based approaches.
Stanley Kok, Pedro M. Domingos
ICML1
2004 Web-scale information extraction in knowitall: (preliminary results)
abstract
Manually querying search engines in order to accumulate a large bodyof factual information is a tedious, error-prone process of piecemealsearch. Search engines retrieve and rank potentially relevantdocuments for human perusal, but do not extract facts, assessconfidence, or fuse information from multiple documents. This paperintroduces KnowItAll, a system that aims to automate the tedious process ofextracting large collections of facts from the web in an autonomous,domain-independent, and scalable manner.The paper describes preliminary experiments in which an instance of KnowItAll, running for four days on a single machine, was able to automatically extract 54,753 facts. KnowItAll associates a probability with each fact enabling it to trade off precision and recall. The paper analyzes KnowItAll's architecture and reports on lessons learned for the design of large-scale information extraction systems.
Oren Etzioni, Michael J. Cafarella, Doug Downey, Stanley Kok, Ana-Maria Popescu, Tal Shaked, Stephen Soderland, Daniel S. Weld, Alexander Yates
WWW4