Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Liya Fan

dblp:99/5035 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 since 2021Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Query processing and optimization · 80% Database system architecture and tuning · 14% Distributed and cloud data management · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 61% Parallel and multicore computing · 39%
Software engineering, system software, and programming languages
1 paper
Program analysis · 77% Compilers and program optimization · 23%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
incremental computation
1.122023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › incremental computation
incremental query processing
0.922020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › query optimization
cost-based optimization
0.712023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Query processing and optimization
query optimization
0.412020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.412020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Program analysis › data flow analysis
value-flow analysis
0.212016
A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016
Cloud and datacenter computing
multi-tenancy
0.212016
A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016
Distributed and cloud data management › cloud database
cloud-native database
0.212023
Krypton: Real-time Serving and Analytical SQL Engine at ByteDance · Proc. VLDB Endow. 2023
Data integration and cleaning
data warehouse
0.112020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › view maintenance
incremental view maintenance
0.112020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Bioinformatics and computational biology › computational structural biology
single particle reconstruction
0.112009
A framework to refine particle clusters produced by EMAN · Bioinform. 2009
Compilers and program optimization
program transformation
0.112016
A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled · IEEE Trans. Serv. Comput. 2016
Bioinformatics and computational biology › structural biology
cryo-electron microscopy
0.012009
A framework to refine particle clusters produced by EMAN · Bioinform. 2009

Methods — techniques the papers use, named apart from their topics

incremental batch processing · 0.9hierarchical cache · 0.7cost-based optimizer framework · 0.7columnar storage · 0.7taint analysis · 0.5time-varying relations · 0.4rewrite rules · 0.4plan space exploration · 0.4value-flow graph · 0.2value flow graph · 0.2re-clustering · 0.1particle orientation determination · 0.1
YearPublicationVenuePosition
2025 Bidirectional two-dimensional supervised multiset canonical correlation analysis for multi-view feature extraction
Liya Fan, Quan-Sen Sun, Xizhan Gao
Pattern Recognit. Lett.2
2024 A multi-rank two-dimensional CCA based on PDEs for multi-view feature extraction
Liya Fan, Quan-Sen Sun
Expert Syst. Appl.2
2023 Krypton: Real-time Serving and Analytical SQL Engine at ByteDance
abstract
In recent years, at ByteDance, we have started seeing more and more business scenarios that require performing real-time data serving besides complex Ad Hoc analysis over large amounts of freshly imported data. The serving workload requires performing complex queries over massive newly added data items with minimal delay. These systems are often used in mission-critical scenarios, whereas traditional OLAP systems cannot handle such use cases. To work around the problem, ByteDance products often have to use multiple systems together in production, forcing the same data to be ETLed into multiple systems, causing data consistency problems, wasting resources, and increasing learning and maintenance costs. To solve the above problem, we built a single Hybrid Serving and Analytical Processing (HSAP) system to handle both workload types. HSAP is still in its early stage, and very few systems are yet on the market. This paper demonstrates how to build Krypton, a competitive cloud-native HSAP system that provides both excellent elasticity and query performance by utilizing many previously known query processing techniques, a hierarchical cache with persistent memory, and a native columnar storage format. Krypton can support high data freshness, high data ingestion rates, and strong data consistency. We also discuss lessons and best practices we learned in developing and operating Krypton in production.
Jianjun Chen 0001, Li Zhang 0132, Liya Fan, Mu Xiong, Benchao Dong, Kuankuan Guo, Yuanjin Lin, Zikang Wang, Yemeng Yang, Junda Zhao, Dongyan Zhou, Zhikai Zuo, Yuming Liang
Proc. VLDB Endow.7
2023 Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version)
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
VLDB J.8
2022 A novel supervised correlation analysis based on partial differential equations for multi-feature extraction and fusion
Liya Fan, Quan-Sen Sun
Multim. Tools Appl.2
2020 Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing
abstract
As the primary approach to deriving decision-support insights, automated recurring routine analytic jobs account for a major part of cluster resource usages in modern enterprise data warehouses. These recurring routine jobs usually have stringent schedule and deadline determined by external business logic, and thus cause dreadful resource skew and severe resource over-provision in the cluster. In this paper, we present Grosbeak, a novel data warehouse that supports resource-aware incremental computing to process recurring routine jobs, smooths the resource skew, and optimizes the resource usage. Unlike batch processing in traditional data warehouses, Grosbeak leverages the fact that data is continuously ingested. It breaks an analysis job into small batches that incrementally process the progressively available data, and schedules these small-batch jobs intelligently when the cluster has free resources. In this demonstration, we showcase Grosbeak using real-world analysis pipelines. Users can interact with the data warehouse by registering recurring queries and observing the incremental scheduling behavior and smoothed resource usage pattern.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
SIGMOD Conference8
2020 Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing
abstract
Incremental processing is widely-adopted in many applications, ranging from incremental view maintenance, stream computing, to recently emerging progressive data warehouse and intermittent query processing. Despite many algorithms developed on this topic, none of them can produce an incremental plan that always achieves the best performance, since the optimal plan is data dependent. In this paper, we develop a novel cost-based optimizer framework, called Tempura, for optimizing incremental data processing. We propose an incremental query planning model called TIP based on the concept of time-varying relations, which can formally model incremental processing in its most general form. We give a full specification of Tempura, which can not only unify various existing techniques to generate an optimal incremental plan, but also allow the developer to add their rewrite rules. We study how to explore the plan space and search for an optimal incremental plan. We evaluate Tempura in various incremental processing scenarios to show its effectiveness and efficiency.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
Proc. VLDB Endow.8
2016 A Semi-Automatic Approach of Transforming Applications to be Multi-Tenancy Enabled
abstract
As a popular technique in cloud computing, multi-tenancy (MT) can significantly ease software maintenance, and improve resource utilization. To make use of the MT technique, an application may need to be transformed to be MT-enabled. This process involves finding and processing a special kind of data entities named global isolation points (GIPs). Practically, finding all GIPs of an application is challenging. Traditional method involves manually browsing the application code, requiring a great deal of human effort. To solve this problem, we introduce a toolkit named Auto-MT to help find and process GIPs of an application. Auto-MT is able to find new GIPs based on their relations to known GIPs. To characterize the relation, a novel graph called value flow graph (VFG) is introduced, which models the value flows of data entities. It can also be used in other scenarios, like taint analysis. We have implemented Auto-MT as an Eclipse Plug-in, and applied it to transform Roller, a widely used Java application. Experimental results show that Auto-MT saves substantial human effort, and accelerates the process of transforming applications to be MT-enabled.
Liya Fan, Zhihu Wang, Wenhao An, Yu Wang 0021
IEEE Trans. Serv. Comput.1
2015 TBSTM: A Novel and Fast Nonlinear Classification Method for Image Data
abstract
A new classifier for image data classification named as linear twin bounded support tensor machine (linear TBSTM) is proposed by adding regularization terms in objective functions, which results in the realization of structural risk minimization avoids of the singularity of matrices. We know that up to now nonlinear classifiers based on STM for image data classification are not seen more. In order to remedy this limitation, a new matrix kernel function is introduced and based on which the nonlinear version of TBSTM is studied with a detailed theoretical derivation, and then a nonlinear classifier called as nonlinear TBSTM is suggested. In order to examine the effectiveness of the proposed classifiers, a series of comparative experiments with three linear classifiers STM, TSTM and PSTM are performed on 15 binary image classification problems taken from ORL, YALE and AR datasets. Experiment results show that the proposed classifiers are effective and efficient.
Liya Fan, Xizhan Gao
Int. J. Pattern Recognit. Artif. Intell.2
2015 Scheduling for energy minimization on restricted parallel processors
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002
J. Parallel Distributed Comput.3
2015 Projection twin SMMs for 2d image data classification
Liya Fan, Xizhan Gao
Neural Comput. Appl.2
2014 A Novel Indefinite Kernel Dimensionality Reduction Algorithm: Weighted Generalized Indefinite Kernel Discriminant Analysis
Liya Fan
Neural Process. Lett.2
2013 Energy-Efficient Scheduling with Time and Processors Eligibility Restrictions
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002
Euro-Par4
2013 A fast calculation strategy of density function in ISAF reconstruction algorithm
Gongming Wang, Fa Zhang 0001, Qi Chu 0002, Liya Fan, Zhiyong Liu 0002
Sci. China Inf. Sci.4
2012 A Cost-Effective Approach to Delivering Analytics as a Service
abstract
Analytical solutions are considered as increasingly important for modern enterprises. Currently, systematical adoption of analytical solutions is limited to only a small set of large enterprises, as the deployment cost is high due to high performance hardware requirement and expensive analytics software. Moreover, such on-premises solutions are not suitable for the occasional analytics consumers. In order to accelerate the prevalence of analytical solutions, this paper explores the feasibility of leveraging SaaS (Software-as-a-Service) delivery model to provide analytics capabilities as services in a cost-effective way. The main contributions of our work include: (1) proposing a framework to enable enterprise tenants to consume analytics capabilities as services; (2) developing a method to enhance existing analytics platform to support multi-tenancy so that a single software instance can effectively support multiple concurrent tenants; (3) designing an SLA (Service Level Agreement) customization mechanism to satisfy the diverse analytics capability demands of tenants. A prototype system has been developed to evaluate the feasibility of our approach.
Liya Fan, Wenhao An
ICWS3
2012 Feature extraction using fuzzy maximum margin criterion
Liya Fan
Neurocomputing2
2012 An effective approximation algorithm for the Malleable Parallel Task Scheduling problem
Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002
J. Parallel Distributed Comput.1
2012 A novel supervised dimensionality reduction algorithm: Graph-based Fisher analysis
Liya Fan
Pattern Recognit.2
2009 An effective scheduling algorithm for Linear Makespan Minimization on Unrelated Parallel Machines
abstract
A simple yet common scheduling problem is identified, as a special case of the R||Cmaxproblem. We name it Linear Makespan Minimization on Unrelated Parallel Machines (LMMUPM). A novel algorithm, MOBSA (Multi-Objective Based Scheduling Algorithm), is presented to solve it. Two auxiliary problems are introduced as the basis of our algorithm. The first one can be reduced to a Multi-Objective Integer Program, while the second is constructed based on the solution of the first one. Results on random datasets revealed that MOBSA produced smaller and more stable makespans than other scheduling algorithms. Additionally, the makespan produced by MOBSA was within 1% of the optimum for every case. Presently, MOBSA has been applied to parallelize EMAN, one of the most popular software packages for cryo-electron microscopy single particle reconstruction. High speedups and ideal load balancing have been obtained. It is expected that MOBSA is also applicable to other similar applications.
Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002
HiPC1
2009 A framework to refine particle clusters produced by EMAN
abstract
MOTIVATION: EMAN is one of the most popular software packages for single particle reconstruction. But the particle clusters produced during its model refining stage are of low qualities. We attempt to refine the particle clusters by more accurately determining orientations of particles, and thereby achieving higher resolutions of consequent 3D structures. RESULTS: A particle reclustering framework (PRF) is introduced, which consists of three components. Each of them is responsible for one of the basic tasks of PRF: normalization, threshold determination and reclustering. Our implementation is also described and proved to meet the constraints proposed by PRF. Experiments revealed that our implementation improved resolutions of consequent structures for most cases, but only a little extra execution time was incurred. Therefore, it is practical to incorporate PRF in EMAN to improve qualities of generated 3D structures. AVAILABILITY AND IMPLEMENTATION: Implementation of our algorithm is available upon request from the authors.
Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002
Bioinform.1