Wee Hyong Tok

dblp:t/WHTok · DBLP profile ↗
← Back
20ranked-venue papers in the field
9as first author
5since 2021 · last 2025
0000-0001-8346-5003ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 15 (9 first)Data Mining & Knowledge Discovery · 5
YearPublicationVenuePosition
2025 KDD 2025 - AI Reasoning Day
abstract
Generative AI and the use of large language models (LLMs) are changing the way we work, create, play, and live. As we have witnessed in the past few years, there is significant progress in training LLMs to have a deep understanding of the semantics of language so that such models begin to perform ''reasoning''. (Human) Reasoning is the process of applying logic to derive conclusions based on new or existing information with the goal of finding the truth. Reasoning is a form of high-level human intelligence. There are many types of reasoning: mathematical reasoning, common sense reasoning, temporal reasoning, among others. Multi-hope reasoning with LLM is an emerging capability for LLMs with tens of billions of parameters. Such ''reasoning models'', including Sonnet 3.7, Chat GPT O1, have powered important application areas such as AI4coding, agentic workflow, among others. The first KDD AI Reasoning Day is a special event that we organize in order to increase the awareness of this important research topic for the research community. We bring leaders from industry and academia to present the latest progresses on improving LLM's reasoning capability and enabling reasoning for different application development.
Jun Huan, Ye Xing, Wee Hyong Tok, Ruzica Piskac
KDD (2)4
2024 KDD 2024 Special Day - AI for Environment
abstract
Environmental problems such as air pollution monitoring and prevention, flood detection and prevention, land use, forest management, river water quality, wastewater treatment supervision, management of biodiversity etc. are more complex than typical realworld problems usually AI faces to.This added complexity rises from several aspects, such as the randomness shown by most of environmental processes involved, the 2D/3D nature of involved problems, the temporal aspects, the spatial aspects, the inexactness of the information, the multigranularity of the information etc.In fact, environmental problems belong to the most difficult problems with a lot of inexactness and uncertainty, and possibly conflicting objectives to be solved according to several classifications such as the one by Funtowicz Ravetz (Funtowicz Ravetz, 1999), which states that there are 3 kinds of problems.Also, they are non-structured problems in the classification proposed by H. Simon (Simon, 1966).All this complexity means that to effectively solve those problems a lot of knowledge is needed.This knowledge can be theoretical knowledge expressed in mechanistic models, such as the Gravidity Newton's Theory, or it can be empirical knowledge that can be expressed by means of empirical models, originated by some data and observations (data-driven knowledge) or by the expertise gathered by people when coping with such problems (model-driven knowledge, particularly expert-based knowledge).The KDD 2024 Special Day for AI for environment brings together researchers and practitioners to present their perspective on this very timely topic on how AI can be used for good, and improving the environment where we all live in.
Karina Gibert, Wee Hyong Tok, Miquel Sànchez-Marrè
KDD2
2024 The 2nd International Workshop: From Innovation to Scale (I2S) - Successfully Build, Commercialize, and Scale AI Innovations
abstract
In recent years, there have been exciting and accelerated developments in AI with novel developments in foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, the pace of adoption of these innovations driven by both academic and industry research labs has sped up with both big tech companies and startups looking to deliver value-differentiated products and services. With Generative AI (GenAI) garnering significant attention, the second edition of the I2S workshop focuses on two aspects: First, bringing together AI thought leaders from academia, big tech, and startups to discuss the opportunities, use-case themes, challenges, and risks of GenAI in various business verticals; and Second, bringing together startup founders to share experiences and lessons learned in commercializing GenAI innovations into successful enterprises highlighting challenges through the entire commercial journey - from productization to acquiring customers, building a team, and securing funding.
Ankur Teredesai, Michael Zeller, Mohak Shah, Shenghua Bao, Wee Hyong Tok, Linsey Pang
KDD5
2024 NL2Code-Reasoning and Planning with LLMs for Code Development
abstract
There is huge value in making software development more productive with AI. An important component of this vision is the capability to translate natural language to a programming language ("NL2Code") and thus to significantly accelerate the speed at which code is written.
Ye Xing, Jun Huan, Wee Hyong Tok, Cong Shen 0001, Johannes Gehrke, Katherine Lin, Arjun Guha, Omer Tripp, Murali Krishna Ramanathan
KDD3
2023 From Innovation to Scale (I2S) - Discuss and Learn How to Successfully Build, Commercialize, and Scale AI Innovations in Challenging Market Conditions
abstract
In recent years, the AI community has witnessed an exciting acceleration in innovation across foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, AI innovations driven by both academic and industry research labs have rapidly been adopted by big tech companies and startups to deliver value-differentiated products and services.
Ankur Teredesai, Michael Zeller, Shenghua Bao, Wee Hyong Tok, Linsey Pang
KDD4
2012 A Framework for Similarity Search of Time Series Cliques with Natural Relations
abstract
A Time Series Clique (TSC) consists of multiple time series which are related to each other by natural relations. The natural relations that are found between the time series depend on the application domains. For example, a TSC can consist of time series which are trajectories in video that have spatial relations. In conventional time series retrieval, such natural relations between the time series are not considered. In this paper, we formalize the problem of similarity search over a TSC database. We develop a novel framework for efficient similarity search on TSC data. The framework addresses the following issues. First, it provides a compact representation for TSC data. Second, it uses a multidimensional relation vector to capture the natural relations between the multiple time series in a TSC. Lastly, the framework defines a novel similarity measure that uses the compact representation and the relation vector. We conduct an extensive performance study, using both real-life and synthetic data sets. From the performance study, we show that our proposed framework is both effective and efficient for TSC retrieval.
Bin Cui 0001, Zhe Zhao 0001, Wee Hyong Tok
IEEE Trans. Knowl. Data Eng.3
2010 A Simple, Yet Effective and Efficient, Sliding Window Sampling Algorithm
Wee Hyong Tok, Chedy Raïssi, Stéphane Bressan
DASFAA (1)2
2010 Efficient similarity matching of Time Series Cliques with natural relations
abstract
A Time Series Clique (TSC) consists of multiple time series. In each TSC, the time series hold some natural relations with each other. In conventional time series retrieval methods, such natural relations are often ignored. In this paper, we formalize the problem of similarity search over TSC databases and develop a novel framework for similarity search on TSC data, which considers both time series patterns and relations. We conduct an extensive performance study, and the results show the effectiveness and efficiency of the proposed method.
Zhe Zhao 0001, Bin Cui 0001, Wee Hyong Tok, Jiakui Zhao
ICDE3
2009 Consistent Top-k Queries over Time
Mong-Li Lee, Wynne Hsu, Wee Hyong Tok
DASFAA4
2008 Twig'n Join: Progressive Query Processing of Multiple XML Streams
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA1
2008 A stratified approach to progressive approximate joins
abstract
Users often do not require a complete answer to their query but rather only a sample. They expect the sample to be either the largest possible or the most representative (or both) given the resources available. We call the query processing techniques that deliver such results 'approximate'. Processing of queries to streams of data is said to be 'progressive' when it can continuously produce results as data arrives. In this paper, we are interested in the progressive and approximate processing of queries to data streams when processing is limited to main memory. In particular, we study one of the main building blocks of such processing: the progressive approximate join. We devise and present several novel progressive approximate join algorithms. We empirically evaluate the performance of our algorithms and compare them with algorithms based on existing techniques. In particular we study the trade-off between maximization of throughput and maximization of representativeness of the sample.
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
EDBT1
2007 RRPJ: Result-Rate Based Progressive Relational Join
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA1
2007 Danaïdes: Continuous and Progressive Complex Queries on RSS Feeds
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA1
2007 Progressive High-Dimensional Similarity Join
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DEXA1
2006 Progressive Spatial Join
abstract
In spatial data exploration and analysis, the system would present a user with initial promising results and empower the user to modify runtime query parameters. The high degree of interactivity would significantly reduce users’ waiting time for results that are not useful, and then having to re-issue a new query. To support this level of interaction during query processing, it necessitates the study of adaptive and progressive spatial query processing techniques that can deliver initial results quickly and adapt to run-time fluctuations during the delivery of remote data. Our goal is to design a generic framework for adaptive and progressive spatial query processing. In this paper, we present our ongoing work on designing progressive spatial join algorithm as an initial step.
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
SSDBM1
2005 An adaptable distributed query processing architecture
Yongluan Zhou, Beng Chin Ooi, Kian-Lee Tan, Wee Hyong Tok
Data Knowl. Eng.4
2002 dbRouter - A Scaleable and Distributed Query Optimization and Processing Framework
Wee Hyong Tok, Stéphane Bressan
DEXA1
2002 Efficient and Adaptive Processing of Multiple Continuous Queries
Wee Hyong Tok, Stéphane Bressan
EDBT1
2002 Data Cleaning and XML: The DBLP Experience
abstract
With the increasing popularity of data-centric XML, data warehousing and mining applications are being developed for rapidly burgeoning XML data repositories. Data quality will no doubt be a critical factor for the success of such applications. Data cleaning, which refers to the processes used to improve data quality, has been well researched in the context of traditional databases. In earlier work we developed a knowledge-based framework for data cleaning relational databases. In this work, we present a novel attempt to apply this framework to XML databases. Our experimental dataset is the DBLP database, a popular online XML bibliography database used by many researchers.
Wai Lup Low, Wee Hyong Tok, Mong-Li Lee, Tok Wang Ling
ICDE2
2002 Predator-Miner: Ad hoc Mining of Associations Rules within a Database Management System
abstract
We present a prototype system, Predator-Miner, which extends Predator with an relational-like association rule mining operator to support data mining operations. Predator-Miner allows a user to combine association rule mining queries with SQL queries. This approach towards tight integration differs from existing techniques of using user-defined functions (UDFs), stored procedures, or re-expressing a mining query as several SQL queries in two aspects. First, by encapsulating the task of association rule mining in a relational operator, we allow association rule mining to be considered as part of the query plan, on which query optimization can be performed on the mining query holistically. Second, by integrating it as a relational operator, we can leverage on the mature field of relational database technology. We extend Predator to support a variant of DMQL, and allow SQL and DMQL to be intermixed in a query. We also demonstrate a cost-based mining query optimization framework.
Wee Hyong Tok, Twee-Hee Ong, Wai Lup Low, Indriyati Atmosukarto, Stéphane Bressan
ICDE1