VLDB 2026 Research / reviewers in the wild / expert
Ming-Chuan Wu
dblp:39/5024
· DBLP profile ↗
13ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0002-7537-4803ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Query processing and optimization · 32% Database system architecture and tuning · 26% Distributed and cloud data management · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% | |
| Artificial intelligence
1 paper |
Learning theory · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › join processing
distributed join |
0.2 | 1 | 2014 | Advanced Join Strategies for Large-Scale Distributed Computation · Proc. VLDB Endow. 2014 |
Query processing and optimization
join processing |
0.2 | 1 | 2014 | Advanced Join Strategies for Large-Scale Distributed Computation · Proc. VLDB Endow. 2014 |
Distributed and cloud data management › mapreduce
mapreduce query processing |
0.1 | 1 | 2012 | SCOPE: parallel databases meet MapReduce · VLDB J. 2012 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2012 | SCOPE: parallel databases meet MapReduce · VLDB J. 2012 |
Parallel and multicore computing
data-parallel programming |
0.1 | 1 | 2012 | Reoptimizing Data Parallel Computing · NSDI 2012 |
Machine learning › Learning theory › computational learning theory › replicable learning
reproducibility of machine learning experiments |
0.1 | 1 | 2019 | Data Platform for Machine Learning · SIGMOD Conference 2019 |
Query processing and optimization › parallel query processing
skew handling |
0.1 | 1 | 2014 | Advanced Join Strategies for Large-Scale Distributed Computation · Proc. VLDB Endow. 2014 |
Query processing and optimization › query optimization
cost-based optimization |
0.1 | 1 | 2005 | Distributed/Heterogeneous Query Processing in Microsoft SQL Server · ICDE 2005 |
Distributed and cloud data management
distributed query processing |
0.1 | 1 | 2005 | Distributed/Heterogeneous Query Processing in Microsoft SQL Server · ICDE 2005 |
Data integration and cleaning
heterogeneous query processing |
0.1 | 1 | 2005 | Distributed/Heterogeneous Query Processing in Microsoft SQL Server · ICDE 2005 |
Indexing and storage engines
bitmap index |
0.0 | 2 | 1999 | Query Optimization for Selections Using Bitmaps · SIGMOD Conference 1999 Encoded Bitmap Indexing for Data Warehouses · ICDE 1998 |
Indexing and storage engines › bitmap index
compressed bitmap index |
0.0 | 2 | 1999 | Query Optimization for Selections Using Bitmaps · SIGMOD Conference 1999 Encoded Bitmap Indexing for Data Warehouses · ICDE 1998 |
Compilers and program optimization
parallel program optimization |
0.0 | 1 | 2012 | Reoptimizing Data Parallel Computing · NSDI 2012 |
Query processing and optimization
view maintenance |
0.0 | 1 | 2003 | Statistics on Views · VLDB 2003 |
Data integration and cleaning
data source integration |
0.0 | 1 | 2005 | Distributed/Heterogeneous Query Processing in Microsoft SQL Server · ICDE 2005 |
Methods — techniques the papers use, named apart from their topics
distributed execution strategies · 0.2mapreduce · 0.1algebraic transformations · 0.1OLE DB interfaces · 0.1tree reduction · 0.0logical reduction · 0.0range-based indexing · 0.0bit slicing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Data Platform for Machine LearningabstractIn this paper, we present a purpose-built data management system, MLdp, for all machine learning (ML) datasets. ML applications pose some unique requirements different from common conventional data processing applications, including but not limited to: data lineage and provenance tracking, rich data semantics and formats, integration with diverse ML frameworks and access patterns, trial-and-error driven data exploration and evolution, rapid experimentation, reproducibility of the model training, strict compliance and privacy regulations, etc. Current ML systems/services, often named MLaaS, to-date focus on the ML algorithms, and offer no integrated data management system. Instead, they require users to bring their own data and to manage their own data on either blob storage or on file systems. The burdens of data management tasks, such as versioning and access control, fall onto the users, and not all compliance features, such as terms of use, privacy measures, and auditing, are available. MLdp offers a minimalist and flexible data model for all varieties of data, strong version management to guarantee re-producibility of ML experiments, and integration with major ML frameworks. MLdp also maintains the data provenance to help users track lineage and dependencies among data versions and models in their ML pipelines. In addition to table-stake features, such as security, availability and scalability, MLdp's internal design choices are strongly influenced by the goal to support rapid ML experiment iterations, which cycle through data discovery, data exploration, feature engineering, model training, model evaluation, and back to data discovery. The contributions of this paper are: 1) to recognize the needs and to call out the requirements of an ML data platform, 2) to share our experiences in building MLdp by adopting existing database technologies to the new problem as well as by devising new solutions, and 3) to call for actions from our communities on future challenges. Pulkit Agrawal 0002, Rajat Arya, Aanchal Bindal, Sandeep Bhatia, Anupriya Gagneja, Joseph Godlewski, Yucheng Low, Timothy Muss, Mudit Manu Paliwal, Sethu Raman, Vishrut Shah, Bochao Shen, Laura Sugden, Kaiyu Zhao, Ming-Chuan Wu |
SIGMOD Conference | 15 |
| 2014 | Advanced Join Strategies for Large-Scale Distributed ComputationabstractCompanies providing cloud-scale data services have increasing needs to store and analyze massive data sets (e.g., search logs, click streams, and web graph data). For cost and performance reasons, processing is typically done on large clusters of thousands of commodity machines by using high level scripting languages. In the recent past, there has been significant progress in adapting well-known techniques from traditional relational DBMSs to this new scenario. However, important challenges remain open. In this paper we study the very common join operation, discuss some unique challenges in the large-scale distributed scenario, and explain how to efficiently and robustly process joins in a distributed way. Specifically, we introduce novel execution strategies that leverage opportunities not available in centralized scenarios, and others that robustly handle data skew. We report experimental validations of our approaches on Scope production clusters, which power the Applications and Services Group at Microsoft. Nicolas Bruno, YongChul Kwon, Ming-Chuan Wu |
Proc. VLDB Endow. | 3 |
| 2012 | Reoptimizing Data Parallel Computing
Sameer Agarwal 0002, Srikanth Kandula, Nicolas Bruno, Ming-Chuan Wu, Ion Stoica, Jingren Zhou 0001 |
NSDI | 4 |
| 2012 | Recurring job optimization in scopeabstractNo abstract available. Nicolas Bruno, Sameer Agarwal 0002, Srikanth Kandula, Ming-Chuan Wu, Jingren Zhou 0001 |
SIGMOD Conference | 5 |
| 2012 | A Hot Query Bank approach to improve detection performance against SQL injection attacks
Yu-Chi Chung, Ming-Chuan Wu, Yih-Chang Chen, Wen-Kui Chang |
Comput. Secur. | 2 |
| 2012 | SCOPE: parallel databases meet MapReduce
Jingren Zhou 0001, Nicolas Bruno, Ming-Chuan Wu, Per-Åke Larson, Ronnie Chaiken, Darren Shakib |
VLDB J. | 3 |
| 2005 | Distributed/Heterogeneous Query Processing in Microsoft SQL ServerabstractThis paper presents an architecture overview of the distributed, heterogeneous query processor (DHQP) in the Microsoft SQL server database system to enable queries over a large collection of diverse data sources. The paper highlights three salient aspects of the architecture. First, the system introduces well-defined abstractions such as connections, commands, and rowsets that enable sources to plug into the system. These abstractions are formalized by the OLE DB data access interfaces. The generality of OLE DB and its broad industry adoption enables our system to reach a very large collection of diverse data sources ranging from personal productivity tools, to database management systems, to file system data. Second, the DHQP is built-in to the relational optimizer and execution engine of the system. This enables DH queries and updates to benefit from the cost-based algebraic transformations and execution strategies available in the system. Finally, the architecture is inherently extensible to support new data sources as they emerge as well as serves as a key extensibility point for the relational engine to add new features such as full-text search and distributed partitioned views. José A. Blakeley, Conor Cunningham, Nigel Ellis, Balaji Rathakrishnan, Ming-Chuan Wu |
ICDE | 5 |
| 2003 | Statistics on Views
César A. Galindo-Legaria, Milind Joshi, F. Michael Waas, Ming-Chuan Wu |
VLDB | 4 |
| 1999 | Query Optimization for Selections Using BitmapsabstractBitmaps are popular indexes for data warehouse (DW) applications and most database management systems offer them today. This paper proposes query optimization strategies for selections using bitmaps. Both continuous and discrete selection criteria are considered. Query optimization strategies are categorized into static and dynamic. Static optimization strategies discussed are the optimal design of bitmaps, and algorithms based on tree and logical reduction. The dynamic optimization discussed is the approach of inclusion and exclusion for both bit-sliced indexes and encoded bitmap indexes. Ming-Chuan Wu |
SIGMOD Conference | 1 |
| 1999 | Supporting Group-By and Pipelining in Bitmap-Enabled Query Processors
Alejandro P. Buchmann, Ming-Chuan Wu |
SOFSEM | 2 |
| 1998 | Encoded Bitmap Indexing for Data WarehousesabstractComplex query types, huge data volumes, and very high read/update ratios make the indexing techniques designed and tuned for traditional database systems unsuitable for data warehouses (DW). We propose an encoded bitmap indexing for DWs which improves the performance of known bitmap indexing in the case of large cardinality domains. A performance analysis and theorems which identify properties of good encodings for better performance are presented. We compare encoded bitmap indexing with related techniques, such as bit slicing, projection-, dynamic-, and range-based indexing. Ming-Chuan Wu, Alejandro P. Buchmann |
ICDE | 1 |
| 1996 | A Hyperrelational Approach to Integration and Manipulation of Data in Multidatabase SystemsabstractThe issue of interoperability among multiple autonomous databases has attracted a lot of attention from researchers in these years. The past research on this issue can be roughly divided into two main categories: the tightly-integrated approach that integrate databases by building an integrated schema and the loosely-integrated approach that achieves interoperability by using a multidatabase language. Most of the past efforts focused on the issues in the first approach. The problem with the first approach is, however, that it lacks a convenient representation of the integrated schema at the system level and a sound mathematical basis for data manipulation in a multidatabase system. In this paper, we propose to use hyperrelations as a powerful and succinct model for the global level representation of heterogeneous database schemas. A hyperrelation has the structure of a relation, but its contents are the schemas of the semantically equivalent local relations in the databases. With this representation, the metadata of the global database, local databases and the data of these databases are all representable by using the structure of a relation. The impact of such a representation is that all the elegant features of relational systems can be easily extended to multidatabase systems. A hyperrelational algebra is designed accordingly. This algebra is performed at the multidatabase systems (MDBS) level such that query transformation and optimization is supported on a sound mathematical basis. The major contributions of this paper include: (1) Local relations of various schemas (even though they retain information of the same semantics) can be systematically mapped to hyperrelations. As the structure of a hyperrelation is similar to that of a relation, data manipulation and management tasks (such as design of the global query language and the view mechanism) are greatly facilitated. (2) The hyperrelational algebra provides a sound basis for query transformation and optimization in a MDBS. Chiang Lee, Ming-Chuan Wu |
Int. J. Cooperative Inf. Syst. | 2 |
| 1994 | On concurrency control in multidatabase systemsabstractA multidatabase system is a system that supports data access to a collection of preexisting autonomous database systems. Concurrency control in such an environment is much more difficult in nature than that in centralized or distributed (homogeneous) database systems. Many factors, especially the autonomies, bring difficulties to it. We classify the multidatabase environment into value independent and value dependent ones, and discuss concurrency control mechanisms in these environments. We show that in the value independent environment, the task of concurrency control is resolvable by utilizing conventional concurrency control schemes. As for value dependent environment, we propose a general protocol which is deadlock free and preserves the local autonomy to solve the problems.> Ming-Chuan Wu, Chiang Lee |
COMPSAC | 1 |