Yali Zhu

dblp:10/7026 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
1since 2021 · last 2023
0009-0000-3254-4545ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
6 papers
Query processing and optimization · 73% Data stream processing · 27%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › transcriptomics
alternative splicing analysis
0.712023
FungiExp: a user-friendly database and analysis platform for exploring fungal gene expression and alternative splicing · Bioinform. 2023
Bioinformatics and computational biology
gene expression analysis
0.712023
FungiExp: a user-friendly database and analysis platform for exploring fungal gene expression and alternative splicing · Bioinform. 2023
Bioinformatics and computational biology › network bioinformatics › biological network analysis
gene co-expression network analysis
0.212023
FungiExp: a user-friendly database and analysis platform for exploring fungal gene expression and alternative splicing · Bioinform. 2023
Query processing and optimization
parallel query processing
0.212013
Adaptive and Big Data Scale Parallel Execution in Oracle · Proc. VLDB Endow. 2013
Query processing and optimization › query execution
SQL operators
0.212013
Adaptive and Big Data Scale Parallel Execution in Oracle · Proc. VLDB Endow. 2013
Data stream processing
continuous query processing
0.232006
Run-time operator state spilling for memory intensive long-running queries · SIGMOD Conference 2006
A Dynamically Adaptive Distributed System for Processing Complex Continuous Queries · VLDB 2005
Dynamic Plan Migration for Continuous Queries Over Data Streams · SIGMOD Conference 2004
Query processing and optimization › query execution
memory-constrained query processing
0.112006
Run-time operator state spilling for memory intensive long-running queries · SIGMOD Conference 2006
Cloud and datacenter computing
big data analytics
0.012013
Adaptive and Big Data Scale Parallel Execution in Oracle · Proc. VLDB Endow. 2013
Query processing and optimization
cost model
0.012004
Dynamic Plan Migration for Continuous Queries Over Data Streams · SIGMOD Conference 2004
Query processing and optimization › query optimization › statistics management
optimizer statistics
0.012008
Optimizer plan change management: improved stability and performance in Oracle 11g · Proc. VLDB Endow. 2008
Query processing and optimization
long-running queries
0.012006
Run-time operator state spilling for memory intensive long-running queries · SIGMOD Conference 2006
Query processing and optimization
query execution
0.012004
Dynamic Plan Migration for Continuous Queries Over Data Streams · SIGMOD Conference 2004

Methods — techniques the papers use, named apart from their topics

RNA-seq · 0.7runtime data distribution · 0.3multi-stage parallelization · 0.3simulation · 0.1experimental evaluation · 0.0cost modeling · 0.0
YearPublicationVenuePosition
2023 FungiExp: a user-friendly database and analysis platform for exploring fungal gene expression and alternative splicing
abstract
SUMMARY: Fungi form a large and heterogeneous group of eukaryotic organisms with diverse ecological niches. The high importance of fungi contrasts with our limited understanding of fungal lifestyle and adaptability to environment. Over the last decade, the high-throughput sequencing technology produced tremendous RNA-sequencing (RNA-seq) data. However, there is no comprehensive database for mycologists to conveniently explore fungal gene expression and alternative splicing. Here, we have developed FungiExp, an online database including 35 821 curated RNA-seq samples derived from 220 fungal species, together with gene expression and alternative splicing profiles. It allows users to query and visualize gene expression and alternative splicing in the collected RNA-seq samples. Furthermore, FungiExp contains several online analysis tools, such as differential/specific, co-expression network and cross-species gene expression conservation analysis. Through these tools, users can obtain new insights by re-analyzing public RNA-seq data or upload personal data to co-analyze with public RNA-seq data. AVAILABILITY AND IMPLEMENTATION: The FungiExp is freely available at https://bioinfo.njau.edu.cn/fungiExp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jinding Liu, Yaru Zhang, Yapin Shi, Yiqing Zheng, Yali Zhu, Zhuoran Guan, Danyu Shen, Daolong Dou
Bioinform.5
2013 Adaptive and Big Data Scale Parallel Execution in Oracle
abstract
This paper showcases some of the newly introduced parallel execution methods in Oracle RDBMS. These methods provide highly scalable and adaptive evaluation for the most commonly used SQL operations - joins, group-by, rollup/cube, grouping sets, and window functions. The novelty of these techniques is their use of multi-stage parallelization models, accommodation of optimizer mistakes, and the runtime parallelization and data distribution decisions. These parallel plans adapt based on the statistics gathered on the real data at query execution time. We realized enormous performance gains from these adaptive parallelization techniques. The paper also discusses our approach to parallelize queries with operations that are inherently serial. We believe all these techniques will make their way into big data analytics and other massively parallel database systems.
Srikanth Bellamkonda, Hua-Gang Li, Unmesh Jagtap, Yali Zhu, Vince Liang, Thierry Cruanes
Proc. VLDB Endow.4
2010 A new look at generating multi-join continuous query plans: A qualified plan generation problem
Yali Zhu, Venkatesh Raghavan, Elke A. Rundensteiner
Data Knowl. Eng.1
2008 Optimizer plan change management: improved stability and performance in Oracle 11g
abstract
Execution plans for SQL statements have a significant impact on the overall performance of database systems. New optimizer statistics, configuration parameter changes, software upgrades and hardware resource utilization are among a multitude of factors that may cause the query optimizer to generate new plans. While most of these plan changes are beneficial or benign, a few rogue plans can potentially wreak havoc on system performance or availability, affecting critical and time-sensitive business application needs. The normally desirable ability of a query optimizer to adapt to system changes may sometimes cause it to pick a sub-optimal plan compromising the stability of the system. In this paper, we present the new SQL Plan Management feature in Oracle 11g. It provides a comprehensive solution for managing plan changes to provide stable and optimal performance for a set of SQL statements. Two of its most important goals are preventing sub-optimal plans from being executed while allowing new plans to be used if they are verifiably better than previous plans. This feature is tightly integrated with Oracle's query optimizer. SQL Plan Management is available to users via both command-line and graphical interfaces. We describe the feature and then, using an industrial-strength application suite, present experimental results that show that SQL Plan Management provides stable and optimal performance for SQL statements with no performance regressions.
Mohamed Ziauddin, Dinesh Das, Hong Su, Yali Zhu, Khaled Yagoub
Proc. VLDB Endow.4
2006 Run-time operator state spilling for memory intensive long-running queries
abstract
Main memory is a critical resource when processing long-running queries over data streams with state intensive operators. In this work, we investigate state spill strategies that handle run-time memory shortage when processing such complex queries by selectively pushing operator states into disks. Unlike previous solutions which all focus on one single operator only, we instead target queries with multiple state intensive operators. We observe an interdependency among multiple operators in the query plan when spilling operator states. We illustrate that existing strategies, which do not take account of this interdependency, become largely ineffective in this query context. Clearly, a consolidated plan level spill strategy must be devised to address this problem. Several data spill strategies are proposed in this paper to maximize the run-time query throughput in memory constrained environments. The bottom-up state spill strategy is an operator-level strategy that treats all data in one operator state equally. More sophisticated partition-level data spill strategies are then proposed to take different characteristics of the input data into account, including the local output, the global output and the global output with penalty strategies. All proposed state spill strategies have been implemented in the D-CAPE continuous query system. The experimental results confirm the effectiveness of our proposed strategies. In particular, the global output strategy and the global output with penalty strategy have shown favorable results as compared to the other two more localized strategies.
Bin Liu 0005, Yali Zhu, Elke A. Rundensteiner
SIGMOD Conference2
2005 Modeling Diverse and Complex Interactions Enabled by Middleware as Connectors in Software Architectures
abstract
Middleware enables distributed components to interact with each others in diverse and complex manners. Such interactions should be modeled at architecture level for controlling the complexity of incorporating middleware into the target system. This paper extends a traditional architectural description language for describing the diverse and complex interactions enabled by middleware as complex connectors and constraints on them in a model driven process. Such functions and qualities of connectors that satisfy the requirements of the target system are modeled without any consideration of middleware at first. Then the connectors and constraints on them are refined by the characteristics induced by middleware. All information of connectors produced in the two-step process can be described at three levels, including the connection, coordination and context. The language and process are illustrated and evaluated by applying them into J2EE (Java 2 Platform Enterprise Edition) applications.
Yali Zhu
ICECCS1
2005 Modeling Architecture Based Development in UML
abstract
In this paper, an approach is presented to formally model architecture based software development process. The ability of UML in modeling software architecture is reinforced by defining a generic model of component and software architecture, and by integrating the model with UML class model and interaction model to unify software development process. In UML, software development is modeled in different views. With formal semantics, designer can keep consistency among these views. The paper demonstrates how to use the method to construct software.
Yali Zhu, Gang Huang 0001, Hong Mei 0001
ICECCS1
2005 An Adaptive Multi-Objective Scheduling Selection Framework for Continuous Query Processing
abstract
Adaptive operator scheduling algorithms for continuous query processing are usually designed to serve a single performance objective, such as minimizing memory usage or maximizing query throughput. We observe that different performance objectives may sometimes conflict with each other. Also due to the dynamic nature of streaming environments, the performance objective may need to change dynamically. Furthermore, the performance specification defined by users may itself be multi-dimensional. Therefore, utilizing a single scheduling algorithm optimized for a single objective is no longer sufficient. In this paper, we propose a novel adaptive scheduling algorithm selection framework named AMoS. It is able to leverage the strengths of existing scheduling algorithms to meet multiple performance objectives. AMoS employs a lightweight learning mechanism to assess the effectiveness of each algorithm. The learned knowledge can be used to select the algorithm that probabilistically has the best chance of improving the performance. In addition, AMoS has the flexibility to add and adapt to new scheduling algorithms, query plans and data sets during execution. Our experimental results show that AMoS significantly outperforms the existing scheduling algorithms with regard to satisfying both uni-objective and multi-objective performance requirements.
Timothy M. Sutherland, Yali Zhu, Luping Ding, Elke A. Rundensteiner
IDEAS2
2005 A Dynamically Adaptive Distributed System for Processing Complex Continuous Queries
Bin Liu 0005, Yali Zhu, Mariana Jbantova, Bradley Momberger, Elke A. Rundensteiner
VLDB2
2004 Quality Attribute Scenario Based Architectural Modeling for Self-Adaptation Supported by Architecture-Based Reflective Middleware
abstract
Reflective middleware is proposed for guaranteeing desired qualities of middleware based systems which reside in the extremely open and dynamic Internet. Current researches and practices focus on how to monitor and change the whole system through reflective mechanisms provided by middleware. However, they put little attention on why, when and what to monitor and change because it is very hard for middleware to collect enough knowledge which is usually specific to the whole system. Being an important artifact in software development, software architecture records plentiful design information, especially the considerations for quality attributes of the target system. It is a natural idea to provide reflective middleware with enough knowledge via software architecture. This paper presents a demonstration of the idea. In this demonstration, the self-adaptations can be analyzed in a quality attribute scenario based way and specified by an extended architecture description language. Such knowledge prescribed at the design phase can be used directly by an architecture based reflective middleware which then automatically adapts itself at runtime.
Yali Zhu, Gang Huang 0001, Hong Mei 0001
APSEC1
2004 Dynamic Plan Migration for Continuous Queries Over Data Streams
abstract
Dynamic plan migration is concerned with the on-the-fly transition from one continuous query plan to a semantically equivalent yet more efficient plan. Migration is important for stream monitoring systems where long-running queries may have to withstand fluctuations in stream workloads and data characteristics. Existing migration methods generally adopt a pause-drain-resume strategy that pauses the processing of new data, purges all old data in the existing plan, until finally the new plan can be plugged into the system. However, these existing strategies do not address the problem of migrating query plans that contain stateful operators, such as joins. We now develop solutions for online plan migration for continuous stateful plans. In particular, in this paper, we propose two alternative strategies, called the moving state strategy and the parallel track strategy, one exploiting reusability and the second employs parallelism to seamlessly migrate between continuous join plans without affecting the results of the query. We develop cost models for both migration strategies to analytically compare them. We embed these migration strategies into the CAPE [7], a prototype system of a stream query engine, and conduct a comparative experimental study to evaluate these two strategies for window-based join plans. Our experimental results illustrate that the two strategies can vary significantly in terms of output rates and intermediate storage spaces given distinct system configurations and stream workloads.
Yali Zhu, Elke A. Rundensteiner, George T. Heineman
SIGMOD Conference1
2004 CAPE: Continuous Query Engine with Heterogeneous-Grained Adaptivity
Elke A. Rundensteiner, Luping Ding, Timothy M. Sutherland, Yali Zhu, Bradford Pielech, Nishant K. Mehta
VLDB4
2003 A simulation-based methodology and tool for automating the modeling and analysis of voice-over-IP perceptual quality
Adrian E. Conway, Yali Zhu
Perform. Evaluation2
2002 Applying objective perceptual quality assessment methods in network performance modeling
abstract
A quantitative approach is developed for modeling and analyzing objectively the perceptual performance of computer-communication networks and computer systems that involve real-time human interaction and communication with real-time streamed signals such as audio, speech, or video. The proposed 'perceptual analysis' method is based on augmenting traditional system performance models with the incorporation of ITU-T recommended objective perceptual quality assessment methods such as PEAQ (perceived audio quality) or PESQ (perceptual evaluation of speech quality). The approach provides an avenue for including the quantitative evaluation of perceptual quality measures, such as MOS (mean opinion score), in system performance modeling. The analysis of a voice-over-IP network access link model with G.711 and G.729 multiplexed calls is presented to demonstrate a perceptual teletraffic engineering application.
Adrian E. Conway, Yali Zhu
ICCCN2