EDBT 2026 Demo / reviewers in the wild / expert
Liya Fan
dblp:99/5035
· DBLP profile ↗
20ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 since 2021Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Query processing and optimization · 80% Database system architecture and tuning · 14% Distributed and cloud data management · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 61% Parallel and multicore computing · 39% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 77% Compilers and program optimization · 23% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
incremental computation |
1.1 | 2 | 2023 | Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023 Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020 |
Query processing and optimization › incremental computation
incremental query processing |
0.9 | 2 | 2020 | Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020 Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020 |
Query processing and optimization › query optimization
cost-based optimization |
0.7 | 1 | 2023 | Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023 |
Query processing and optimization
query optimization |
0.4 | 1 | 2020 | Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.4 | 1 | 2020 | Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020 |
Parallel and multicore computing › parallel scheduling
resource-aware scheduling |
0.4 | 1 | 2020 | Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020 |
Program analysis › data flow analysis
value-flow analysis |
0.2 | 1 | 2016 | A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016 |
Cloud and datacenter computing
multi-tenancy |
0.2 | 1 | 2016 | A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016 |
Distributed and cloud data management › cloud database
cloud-native database |
0.2 | 1 | 2023 | Krypton: Real-time Serving and Analytical SQL Engine at ByteDance · Proc. VLDB Endow. 2023 |
Data integration and cleaning
data warehouse |
0.1 | 1 | 2020 | Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020 |
Query processing and optimization › view maintenance
incremental view maintenance |
0.1 | 1 | 2020 | Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020 |
Bioinformatics and computational biology › computational structural biology
single particle reconstruction |
0.1 | 1 | 2009 | A framework to refine particle clusters produced by EMAN · Bioinform. 2009 |
Compilers and program optimization
program transformation |
0.1 | 1 | 2016 | A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016 |
Bioinformatics and computational biology › structural biology
cryo-electron microscopy |
0.0 | 1 | 2009 | A framework to refine particle clusters produced by EMAN · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
incremental batch processing · 0.9hierarchical cache · 0.7cost-based optimizer framework · 0.7columnar storage · 0.7taint analysis · 0.5time-varying relations · 0.4rewrite rules · 0.4plan space exploration · 0.4value-flow graph · 0.2value flow graph · 0.2re-clustering · 0.1particle orientation determination · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bidirectional two-dimensional supervised multiset canonical correlation analysis for multi-view feature extraction
Liya Fan, Quan-Sen Sun, Xizhan Gao |
Pattern Recognit. Lett. | 2 |
| 2024 | A multi-rank two-dimensional CCA based on PDEs for multi-view feature extraction
Liya Fan, Quan-Sen Sun |
Expert Syst. Appl. | 2 |
| 2023 | Krypton: Real-time Serving and Analytical SQL Engine at ByteDanceabstractIn recent years, at ByteDance, we have started seeing more and more business scenarios that require performing real-time data serving besides complex Ad Hoc analysis over large amounts of freshly imported data. The serving workload requires performing complex queries over massive newly added data items with minimal delay. These systems are often used in mission-critical scenarios, whereas traditional OLAP systems cannot handle such use cases. To work around the problem, ByteDance products often have to use multiple systems together in production, forcing the same data to be ETLed into multiple systems, causing data consistency problems, wasting resources, and increasing learning and maintenance costs. To solve the above problem, we built a single Hybrid Serving and Analytical Processing (HSAP) system to handle both workload types. HSAP is still in its early stage, and very few systems are yet on the market. This paper demonstrates how to build Krypton, a competitive cloud-native HSAP system that provides both excellent elasticity and query performance by utilizing many previously known query processing techniques, a hierarchical cache with persistent memory, and a native columnar storage format. Krypton can support high data freshness, high data ingestion rates, and strong data consistency. We also discuss lessons and best practices we learned in developing and operating Krypton in production. Jianjun Chen 0001, Li Zhang 0132, Liya Fan, Mu Xiong, Benchao Dong, Kuankuan Guo, Yuanjin Lin, Zikang Wang, Yemeng Yang, Junda Zhao, Dongyan Zhou, Zhikai Zuo, Yuming Liang |
Proc. VLDB Endow. | 7 |
| 2023 | Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version)
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001 |
VLDB J. | 8 |
| 2022 | A novel supervised correlation analysis based on partial differential equations for multi-feature extraction and fusion
Liya Fan, Quan-Sen Sun |
Multim. Tools Appl. | 2 |
| 2020 | Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental ComputingabstractAs the primary approach to deriving decision-support insights, automated recurring routine analytic jobs account for a major part of cluster resource usages in modern enterprise data warehouses. These recurring routine jobs usually have stringent schedule and deadline determined by external business logic, and thus cause dreadful resource skew and severe resource over-provision in the cluster. In this paper, we present Grosbeak, a novel data warehouse that supports resource-aware incremental computing to process recurring routine jobs, smooths the resource skew, and optimizes the resource usage. Unlike batch processing in traditional data warehouses, Grosbeak leverages the fact that data is continuously ingested. It breaks an analysis job into small batches that incrementally process the progressively available data, and schedules these small-batch jobs intelligently when the cluster has free resources. In this demonstration, we showcase Grosbeak using real-world analysis pipelines. Users can interact with the data warehouse by registering recurring queries and observing the incremental scheduling behavior and smoothed resource usage pattern. Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001 |
SIGMOD Conference | 8 |
| 2020 | Tempura: A General Cost-Based Optimizer Framework for Incremental Data ProcessingabstractIncremental processing is widely-adopted in many applications, ranging from incremental view maintenance, stream computing, to recently emerging progressive data warehouse and intermittent query processing. Despite many algorithms developed on this topic, none of them can produce an incremental plan that always achieves the best performance, since the optimal plan is data dependent. In this paper, we develop a novel cost-based optimizer framework, called Tempura, for optimizing incremental data processing. We propose an incremental query planning model called TIP based on the concept of time-varying relations, which can formally model incremental processing in its most general form. We give a full specification of Tempura, which can not only unify various existing techniques to generate an optimal incremental plan, but also allow the developer to add their rewrite rules. We study how to explore the plan space and search for an optimal incremental plan. We evaluate Tempura in various incremental processing scenarios to show its effectiveness and efficiency. Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001 |
Proc. VLDB Endow. | 8 |
| 2016 | A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy EnabledabstractAs a popular technique in cloud computing, multi-tenancy (MT) can significantly ease software maintenance, and improve resource utilization. To make use of the MT technique, an application may need to be transformed to be MT-enabled. This process involves finding and processing a special kind of data entities named global isolation points (GIPs). Practically, finding all GIPs of an application is challenging. Traditional method involves manually browsing the application code, requiring a great deal of human effort. To solve this problem, we introduce a toolkit named Auto-MT to help find and process GIPs of an application. Auto-MT is able to find new GIPs based on their relations to known GIPs. To characterize the relation, a novel graph called value flow graph (VFG) is introduced, which models the value flows of data entities. It can also be used in other scenarios, like taint analysis. We have implemented Auto-MT as an Eclipse Plug-in, and applied it to transform Roller, a widely used Java application. Experimental results show that Auto-MT saves substantial human effort, and accelerates the process of transforming applications to be MT-enabled. Liya Fan, Zhihu Wang, Wenhao An, Yu Wang 0021 |
IEEE Trans. Serv. Comput. | 1 |
| 2015 | TBSTM: A Novel and Fast Nonlinear Classification Method for Image DataabstractA new classifier for image data classification named as linear twin bounded support tensor machine (linear TBSTM) is proposed by adding regularization terms in objective functions, which results in the realization of structural risk minimization avoids of the singularity of matrices. We know that up to now nonlinear classifiers based on STM for image data classification are not seen more. In order to remedy this limitation, a new matrix kernel function is introduced and based on which the nonlinear version of TBSTM is studied with a detailed theoretical derivation, and then a nonlinear classifier called as nonlinear TBSTM is suggested. In order to examine the effectiveness of the proposed classifiers, a series of comparative experiments with three linear classifiers STM, TSTM and PSTM are performed on 15 binary image classification problems taken from ORL, YALE and AR datasets. Experiment results show that the proposed classifiers are effective and efficient. Liya Fan, Xizhan Gao |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2015 | Scheduling for energy minimization on restricted parallel processors
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002 |
J. Parallel Distributed Comput. | 3 |
| 2015 | Projection twin SMMs for 2d image data classification
Liya Fan, Xizhan Gao |
Neural Comput. Appl. | 2 |
| 2014 | A Novel Indefinite Kernel Dimensionality Reduction Algorithm: Weighted Generalized Indefinite Kernel Discriminant Analysis
Liya Fan |
Neural Process. Lett. | 2 |
| 2013 | Energy-Efficient Scheduling with Time and Processors Eligibility Restrictions
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002 |
Euro-Par | 4 |
| 2013 | A fast calculation strategy of density function in ISAF reconstruction algorithm
Gongming Wang, Fa Zhang 0001, Qi Chu 0002, Liya Fan, Zhiyong Liu 0002 |
Sci. China Inf. Sci. | 4 |
| 2012 | A Cost-Effective Approach to Delivering Analytics as a ServiceabstractAnalytical solutions are considered as increasingly important for modern enterprises. Currently, systematical adoption of analytical solutions is limited to only a small set of large enterprises, as the deployment cost is high due to high performance hardware requirement and expensive analytics software. Moreover, such on-premises solutions are not suitable for the occasional analytics consumers. In order to accelerate the prevalence of analytical solutions, this paper explores the feasibility of leveraging SaaS (Software-as-a-Service) delivery model to provide analytics capabilities as services in a cost-effective way. The main contributions of our work include: (1) proposing a framework to enable enterprise tenants to consume analytics capabilities as services; (2) developing a method to enhance existing analytics platform to support multi-tenancy so that a single software instance can effectively support multiple concurrent tenants; (3) designing an SLA (Service Level Agreement) customization mechanism to satisfy the diverse analytics capability demands of tenants. A prototype system has been developed to evaluate the feasibility of our approach. Liya Fan, Wenhao An |
ICWS | 3 |
| 2012 | Feature extraction using fuzzy maximum margin criterion
Liya Fan |
Neurocomputing | 2 |
| 2012 | An effective approximation algorithm for the Malleable Parallel Task Scheduling problem
Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
J. Parallel Distributed Comput. | 1 |
| 2012 | A novel supervised dimensionality reduction algorithm: Graph-based Fisher analysis
Liya Fan |
Pattern Recognit. | 2 |
| 2009 | An effective scheduling algorithm for Linear Makespan Minimization on Unrelated Parallel MachinesabstractA simple yet common scheduling problem is identified, as a special case of the R||Cmaxproblem. We name it Linear Makespan Minimization on Unrelated Parallel Machines (LMMUPM). A novel algorithm, MOBSA (Multi-Objective Based Scheduling Algorithm), is presented to solve it. Two auxiliary problems are introduced as the basis of our algorithm. The first one can be reduced to a Multi-Objective Integer Program, while the second is constructed based on the solution of the first one. Results on random datasets revealed that MOBSA produced smaller and more stable makespans than other scheduling algorithms. Additionally, the makespan produced by MOBSA was within 1% of the optimum for every case. Presently, MOBSA has been applied to parallelize EMAN, one of the most popular software packages for cryo-electron microscopy single particle reconstruction. High speedups and ideal load balancing have been obtained. It is expected that MOBSA is also applicable to other similar applications. Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
HiPC | 1 |
| 2009 | A framework to refine particle clusters produced by EMANabstractMOTIVATION: EMAN is one of the most popular software packages for single particle reconstruction. But the particle clusters produced during its model refining stage are of low qualities. We attempt to refine the particle clusters by more accurately determining orientations of particles, and thereby achieving higher resolutions of consequent 3D structures. RESULTS: A particle reclustering framework (PRF) is introduced, which consists of three components. Each of them is responsible for one of the basic tasks of PRF: normalization, threshold determination and reclustering. Our implementation is also described and proved to meet the constraints proposed by PRF. Experiments revealed that our implementation improved resolutions of consequent structures for most cases, but only a little extra execution time was incurred. Therefore, it is practical to incorporate PRF in EMAN to improve qualities of generated 3D structures. AVAILABILITY AND IMPLEMENTATION: Implementation of our algorithm is available upon request from the authors. Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
Bioinform. | 1 |