EDBT 2026 Demo / reviewers in the wild / expert
Ning Zhang 0002
dblp:z/NingZhang0002
· DBLP profile ↗
10ranked-venue papers
6as first author
0since 2021 · last 2013
0000-0002-8781-4925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 6 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Information retrieval · 32% Query processing and optimization · 22% Indexing and storage engines · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
0.2 | 1 | 2013 | Unicorn: A System for Searching the Social Graph · Proc. VLDB Endow. 2013 |
Information retrieval
search engines |
0.2 | 1 | 2013 | Unicorn: A System for Searching the Social Graph · Proc. VLDB Endow. 2013 |
Information retrieval › web search › social search
social network search |
0.2 | 1 | 2013 | Unicorn: A System for Searching the Social Graph · Proc. VLDB Endow. 2013 |
Query processing and optimization
XML query processing |
0.1 | 4 | 2009 | BlossomTree: Evaluating XPaths in FLWOR Expressions · ICDE 2005 A Succinct Physical Storage Scheme for Efficient Evaluation of Path Queries in XML · ICDE 2004 Binary XML Storage and Query Processing in Oracle 11g · Proc. VLDB Endow. 2009 |
Data integration and cleaning
data warehouse |
0.1 | 1 | 2010 | Hive - a petabyte scale data warehouse using Hadoop · ICDE 2010 |
Distributed and cloud data management › mapreduce
mapreduce-based data management |
0.1 | 1 | 2010 | Hive - a petabyte scale data warehouse using Hadoop · ICDE 2010 |
Data models and query languages
XML data management |
0.1 | 1 | 2009 | Binary XML Storage and Query Processing in Oracle 11g · Proc. VLDB Endow. 2009 |
Query processing and optimization
approximate query processing |
0.1 | 1 | 2006 | XSEED: Accurate and Fast Cardinality Estimation for XPath Queries · ICDE 2006 |
Query processing and optimization
cardinality estimation |
0.1 | 1 | 2006 | XSEED: Accurate and Fast Cardinality Estimation for XPath Queries · ICDE 2006 |
Indexing and storage engines
feature-based indexing |
0.1 | 1 | 2006 | FIX: Feature-based Indexing Technique for XML Documents · VLDB 2006 |
Indexing and storage engines
synopsis structure |
0.1 | 1 | 2006 | XSEED: Accurate and Fast Cardinality Estimation for XPath Queries · ICDE 2006 |
Indexing and storage engines
XML indexing |
0.1 | 1 | 2006 | FIX: Feature-based Indexing Technique for XML Documents · VLDB 2006 |
Query processing and optimization › XML query processing
path expression evaluation |
0.1 | 1 | 2005 | BlossomTree: Evaluating XPaths in FLWOR Expressions · ICDE 2005 |
Information retrieval
pattern matching |
0.1 | 1 | 2005 | BlossomTree: Evaluating XPaths in FLWOR Expressions · ICDE 2005 |
Graph data management
path query |
0.0 | 1 | 2004 | A Succinct Physical Storage Scheme for Efficient Evaluation of Path Queries in XML · ICDE 2004 |
Methods — techniques the papers use, named apart from their topics
inverted indexing · 0.3mapreduce · 0.1securefiles · 0.1lightweight navigational index · 0.1NFA-based navigational algorithm · 0.1sampling · 0.1incremental synopsis construction · 0.1feature-based indexing · 0.1statistical learning · 0.1blossomtree · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Unicorn: A System for Searching the Social GraphabstractUnicorn is an online, in-memory social graph-aware indexing system designed to search trillions of edges between tens of billions of users and entities on thousands of commodity servers. Unicorn is based on standard concepts in information retrieval, but it includes features to promote results with good social proximity. It also supports queries that require multiple round-trips to leaves in order to retrieve objects that are more than one edge away from source nodes. Unicorn is designed to answer billions of queries per day at latencies in the hundreds of milliseconds, and it serves as an infrastructural building block for Facebook's Graph Search product. In this paper, we describe the data model and query language supported by Unicorn. We also describe its evolution as it became the primary backend for Facebook's search offerings. Michael Curtiss, Iain Becker, Tudor Bosman, Sergey Doroshenko, Lucian Grijincu, Sandhya Kunnatur, Søren B. Lassen, Philip Pronin, Sriram Sankar, Guanghao Shen, Gintaras Woss, Ning Zhang 0002 |
Proc. VLDB Endow. | 14 |
| 2010 | Hive - a petabyte scale data warehouse using HadoopabstractThe size of data sets being collected and analyzed in the industry for business intelligence is growing rapidly, making traditional warehousing solutions prohibitively expensive. Hadoop is a popular open-source map-reduce implementation which is being used in companies like Yahoo, Facebook etc. to store and process extremely large data sets on commodity hardware. However, the map-reduce programming model is very low level and requires developers to write custom programs which are hard to maintain and reuse. In this paper, we present Hive, an open-source data warehousing solution built on top of Hadoop. Hive supports queries expressed in a SQL-like declarative language - HiveQL, which are compiled into map-reduce jobs that are executed using Hadoop. In addition, HiveQL enables users to plug in custom map-reduce scripts into queries. The language includes a type system with support for tables containing primitive types, collections like arrays and maps, and nested compositions of the same. The underlying IO libraries can be extended to query data in custom formats. Hive also includes a system catalog - Metastore - that contains schemas and statistics, which are useful in data exploration, query optimization and query compilation. In Facebook, the Hive warehouse contains tens of thousands of tables and stores over 700TB of data and is being used extensively for both reporting and ad-hoc analyses by more than 200 users per month. Ashish Thusoo, Joydeep Sen Sarma, Namit Jain, Zheng Shao, Prasad Chakka, Ning Zhang 0002, Suresh Anthony, Hao Liu 0018, Raghotham Murthy |
ICDE | 6 |
| 2009 | Binary XML Storage and Query Processing in Oracle 11gabstractOracle RDBMS has supported XML data management for more than six years since version 9i. Prior to 11g, text-centric XML documents can be stored as-is in a CLOB column and schema-based data-centric documents can be shredded and stored in object-relational (OR) tables mapped from their XML Schema. However, both storage formats have intrinsic limitations---XML/CLOB has unacceptable query and update performance, and XML/OR requires XML schema. To tackle this problem, Oracle 11g introduces a native Binary XML storage format and a complete stack of data management operations. Binary XML was designed to address a wide range of real application problems encountered in XML data management---schema flexibility, amenability to XML indexes, update performance, schema evolution, just to name a few. In this paper, we introduce the Binary XML storage format based on Oracle SecureFiles System[21]. We propose a lightweight navigational index on top of the storage and an NFA-based navigational algorithm to provide efficient streaming processing. We further optimize query processing by exploiting XML structural and schema information that are collected in database dictionary. We conducted extensive experiments to demonstrate high performance of the native Binary XML in query processing, update, and space consumption. Ning Zhang 0002, Nipun Agarwal, Sivasankaran Chandrasekar, Sam Idicula, Vijay Medi, Sabina Petride, Balasubramanyam Sthanikam |
Proc. VLDB Endow. | 1 |
| 2007 | Compact access control labeling for efficient secure XML query evaluation
Huaxin Zhang, Ning Zhang 0002, Kenneth Salem, Donghui Zhuo |
Data Knowl. Eng. | 2 |
| 2006 | XSEED: Accurate and Fast Cardinality Estimation for XPath QueriesabstractWe propose XSEED, a synopsis of path queries for cardinality estimation that is accurate, robust, efficient, and adaptive to memory budgets. XSEED starts from a very small kernel, and then incrementally updates information of the synopsis. With such an incremental construction, a synopsis structure can be dynamically configured to accommodate different memory budgets. Cardinality estimation based on XSEED can be performed very efficiently and accurately. Extensive experiments on both synthetic and real data sets show that even with less memory, XSEED could achieve accuracy that is an order of magnitude better than that of other synopsis structures. The cardinality estimation time is under 2% of the actual querying time for a wide range of queries in all test cases. Ning Zhang 0002, M. Tamer Özsu, Ashraf Aboulnaga, Ihab F. Ilyas |
ICDE | 1 |
| 2006 | InterJoin: Exploiting Indexes and Materialized Views in XPath EvaluationabstractXML has become the standard for data exchange for a wide variety of applications, particularly in the scientific community. In order to efficiently process queries on XML representations of scientific data, we require specialized techniques for evaluating XPath expressions. Exploiting materialized views in query processing significantly enhances query processing performance. We propose a novel view definition that allows for intermediate (structural) join results to be stored and reused in XML query evaluation. Unlike current XML view proposals, our views do not require navigation in the original document or path-based pattern matching. Hence, they are evaluated significantly faster and are easily costed as part of a query plan. In general, current structural joins cannot exploit views efficiently when the view definition is not a prefix (or a suffix) of the XPath query. To increase the applicability of our proposed view definition, we propose a novel physical structural join operator called InterJoin. The InterJoin operator allows for joining interleaving XPath expressions, e.g., joining //A//C with //B to evaluate //A//B//C. InterJoin allows for more join alternatives in XML query plans. We propose several physical implementations for InterJoin, including a technique to exploit spatial indexes on the inputs. We give analytic cost models for the implementations so they can be costed in an existing XML query optimizer. Experiments on real and synthetic XML data show significant speed-ups of up to 200% using InterJoin, and speed-ups of up to 400% using our materialized views Derek Phillips, Ning Zhang 0002, Ihab F. Ilyas, M. Tamer Özsu |
SSDBM | 2 |
| 2006 | FIX: Feature-based Indexing Technique for XML Documents
Ning Zhang 0002, M. Tamer Özsu, Ihab F. Ilyas, Ashraf Aboulnaga |
VLDB | 1 |
| 2005 | BlossomTree: Evaluating XPaths in FLWOR ExpressionsabstractEfficient evaluation of path expressions has been studied extensively. However, evaluating more complex FLWOR expressions that contain multiple path expressions has not been well studied. In this paper, we propose a novel pattern matching approach, called BlossomTree, to evaluate a FLWOR expression that contains correlated path expressions. BlossomTree is a formalism to capture the semantics of the path expressions and their correlations. We propose a general algebraic framework (abstract data types and logical operators) to evaluate BlossomTree pattern matching that facilitates efficient evaluation and experimentation. We design efficient data structures and algorithms to implement the abstract data types and logical operators. Our experimental studies demonstrate that the BlossomTree approach can generate highly efficient query plans in different environments. Ning Zhang 0002, Shishir Agrawal, M. Tamer Özsu |
ICDE | 1 |
| 2005 | Statistical Learning Techniques for Costing XML Queries
Ning Zhang 0002, Peter J. Haas, Vanja Josifovski, Guy M. Lohman |
VLDB | 1 |
| 2004 | A Succinct Physical Storage Scheme for Efficient Evaluation of Path Queries in XMLabstractPath expressions are ubiquitous in XML processing languages. Existing approaches evaluate a path expression by selecting nodes that satisfies the tag-name and value constraints and then joining them according to the structural constraints. We propose a novel approach, next-of-kin (NoK) pattern matching, to speed up the node-selection step, and to reduce the join size significantly in the second step. To efficiently perform NoK pattern matching, we also propose a succinct XML physical storage scheme that is adaptive to updates and streaming XML as well. Our performance results demonstrate that the proposed storage scheme and path evaluation algorithm is highly efficient and outperforms the other tested systems in most cases. Ning Zhang 0002, Varun Kacholia, M. Tamer Özsu |
ICDE | 1 |