Yubin Bao

dblp:76/1733 · DBLP profile ↗
← Back
29ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-7350-0610ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 20 · 2 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 61% Parallel and multicore computing · 32% Storage systems · 4%
Databases, data mining, and information retrieval
3 papers
Graph data management · 38% Data models and query languages · 34% Indexing and storage engines · 20%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed graph processing
1.132021
HGraph: I/O-Efficient Distributed and Iterative Graph Computing by Hybrid Pushing/Pulling · IEEE Trans. Knowl. Data Eng. 2021
A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments · IEEE Trans. Parallel Distributed Syst. 2018
Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing · SIGMOD Conference 2016
Parallel and multicore computing › graph processing
iterative graph processing
0.512021
HGraph: I/O-Efficient Distributed and Iterative Graph Computing by Hybrid Pushing/Pulling · IEEE Trans. Knowl. Data Eng. 2021
Data models and query languages
multidimensional data model
0.412020
Haery: A Hadoop Based Query System on Accumulative and High-Dimensional Data Model for Big Data · IEEE Trans. Knowl. Data Eng. 2020
Distributed systems › fault tolerance › checkpointing
asynchronous checkpointing
0.312018
A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
asynchronous iterative methods
0.312018
A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments · IEEE Trans. Parallel Distributed Syst. 2018
Distributed systems › fault tolerance
failure recovery
0.312018
A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments · IEEE Trans. Parallel Distributed Syst. 2018
Distributed systems
fault tolerance
0.312018
A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments · IEEE Trans. Parallel Distributed Syst. 2018
Graph data management
distributed graph processing
0.212016
Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing · SIGMOD Conference 2016
Graph data management › distributed graph processing
iterative graph computation
0.212016
Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing · SIGMOD Conference 2016
Parallel and multicore computing › parallel programming models
message passing
0.212016
Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing · SIGMOD Conference 2016
Indexing and storage engines
key-value store
0.112020
Haery: A Hadoop Based Query System on Accumulative and High-Dimensional Data Model for Big Data · IEEE Trans. Knowl. Data Eng. 2020
Indexing and storage engines
partitioning
0.112020
Haery: A Hadoop Based Query System on Accumulative and High-Dimensional Data Model for Big Data · IEEE Trans. Knowl. Data Eng. 2020
Information retrieval
keyword search
0.012002
XBase: making your gigabyte disk queriable · SIGMOD Conference 2002
Information retrieval
personal information management
0.012002
XBase: making your gigabyte disk queriable · SIGMOD Conference 2002
Information retrieval
search engines
0.012002
XBase: making your gigabyte disk queriable · SIGMOD Conference 2002

Methods — techniques the papers use, named apart from their topics

priority scheduling · 0.5lightweight fault-tolerance · 0.5hybrid pushing/pulling · 0.5block-centric pulling · 0.5query algorithm · 0.4linearization algorithm · 0.4message pruning · 0.3load balancing · 0.3checkpointing · 0.3query language · 0.0XML metadata · 0.0
YearPublicationVenuePosition
2026 DIAL-KG: Schema-Free Incremental Knowledge Graph Construction via Dynamic Schema Induction and Evolution-Intent Assessment
Weidong Bao 0005, Ruyu Gao, Fangling Leng, Yubin Bao, Ge Yu 0001
DASFAA (6)5
2026 CL-ECDD: A contrastive learning framework for enterprise-level code defect detection
Liyuan Xue, Fangling Leng, Junhao Zou, Yubin Bao, Zikuan Zheng, Ge Yu 0001
Neurocomputing4
2022 Compress Blocks or Not: Tradeoffs for Energy Consumption of a Big Data Processing System
abstract
Currently, in addition to the performance, the energy consumption (hereinafter EC) of jobs running in a big data processing system is also of interest to academia and industry because it grows rapidly as an increasing amount of data is processed. Many studies focus on the EC optimization of jobs from the perspective of computation, which is specific to the algorithms in each job. However, the part of EC involved in I/O operations, which is general and universal, is mostly ignored in optimization. In this paper, we concentrate on the EC optimization of jobs from the perspective of I/O operations. To save energy, we argue that data compression could be exploited. On one hand, energy is saved by processing compressed data with less I/O cost. On the other hand, extra EC is incurred from the necessary data compression/decompression process, which may offset the saved energy. Therefore, there are tradeoffs to consider when determining whether to compress data for these jobs. In this paper, such tradeoffs and boundary conditions are studied. We first abstract a paradigm for the runtime environment of big data processing jobs. Then, we establish the power, jobs, compression, and I/O models in detail. Based on these models, we discuss the compression tradeoffs and derive the boundary conditions for EC optimization. Finally, we design and conduct experiments to validate our proposition. The experimental results confirm that the tradeoffs and boundary conditions exist for typical jobs in MapReduce and Spark. As explained, first, the EC of a job is reduced using data compression. Second, whether or not such optimization occurs is related to the specification of both the compression algorithm and the job and is determined by corresponding boundary conditions. Third, for a compression algorithm, the larger its compression/decompression speed and the better its compression ratio, the more likely it is to achieve EC optimization.
Jie Song 0001, Shengqiang Hu, Yubin Bao, Ge Yu 0001
IEEE Trans. Sustain. Comput.3
2021 HGraph: I/O-Efficient Distributed and Iterative Graph Computing by Hybrid Pushing/Pulling
abstract
In the big data era, distributed computation is becoming a preferred solution for iterative graph analysis. However, graphs are rapidly growing in size and more importantly, there exist a lot of messages across iterations. For better scalability, many distributed systems keep graph data and message data on disk. Now these systems solely employ either pushing or pulling mode to manage data, but neither can always work well during the entire computation. This is mainly because I/O access patterns are dynamic and complex. This article proposes a hybrid solution. It achieves the optimal performance in different scenarios by dynamically and adaptively switching modes between pushing and pulling. Specifically, we first devise a new block-centric pulling technique. It pulls messages much more I/O-efficiently than the existing vertex-centric pulling mode. We then combine pushing and pulling. For general-purpose, we categorize graph algorithms and accordingly present two seamless switching frameworks. We also design performance prediction components specialized to the two frameworks, to decide how and when we can switch modes. Some optimization strategies are also given to further enhance performance, such as priority scheduling and lightweight fault-tolerance. Extensive experiments against state-of-the-art solutions confirm the effectiveness of our proposals.
Zhigang Wang 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001, Jeffrey Xu Yu, Zhiqiang Wei 0002
IEEE Trans. Knowl. Data Eng.3
2020 Named Entity Recognition in Aircraft Design Field Based on Deep Learning
Yubin Bao, Yuanming An, Zhu Cheng, Rimeng Jiao, Fangling Leng, Ge Yu 0001
WISA1
2020 Haery: A Hadoop Based Query System on Accumulative and High-Dimensional Data Model for Big Data
abstract
Column-oriented stores, known for their scalability and flexibility, are a common NoSQL database implementation and are increasingly used in big data management. In column-oriented stores, a “full-scan” query strategy is inefficient and the search space can be reduced if data is well partitioned or indexed; however, there is no pre-defined schema for building and maintaining partitions and indexes at lower cost. We leverage an accumulative and high-dimensional data model, a sophisticated linearization algorithm, and an efficient query algorithm, to solve the challenge of how a pre-defined and well-partitioned data model can be applied to flexible and time-varied key-value data. We adapt a high-dimensional array as the data model to partition the key-value data without additional storage and massive calculation; improve the Z-order linearization algorithm, which map multidimensional data to one dimension while preserving locality of the data points, for flexibility; efficiently build an expansion mechanism for the data model to support time-varied data. The result is Haery, a column-oriented store, based on a distributed file system and computing framework. In experiments, Haery is compared with Hive, HBase, Cassandra, MongoDB, PostgresXL, and HyperDex in terms of query performance. With results indicating Haery on average performs 4.57x, 4.23x, 3.55x, 1.79x, 1.82x, and 120.6x faster, respectively.
Jie Song 0001, HongYan He, Richard Thomas 0002, Yubin Bao, Ge Yu 0001
IEEE Trans. Knowl. Data Eng.4
2019 TSH: Easy-to-be distributed partitioning for large-scale graphs
Ning Wang 0026, Zhigang Wang 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001
Future Gener. Comput. Syst.4
2018 A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments
abstract
Most graph algorithms are iterative in nature. They can be processed by distributed systems in memory in an efficient asynchronous manner. However, it is challenging to recover from failures in such systems. This is because traditional checkpoint fault-tolerant frameworks incur expensive barrier costs that usually offset the gains brought by asynchronous computations. Worse, surviving data are rolled back, leading to costly re-computations. This paper first proposes to leverage surviving data for failure recovery in an asynchronous system. Our framework guarantees the correctness of algorithms and avoids rolling back surviving data. Additionally, a novel asynchronous checkpointing solution is introduced to accelerate recovery at the price of nearly zero overheads. Some optimization strategies like message pruning, non-blocking recovery and load balancing are also designed to further boost the performance. We have conducted extensive experiments to show the effectiveness of our proposals using real-world graphs.
Zhigang Wang 0001, Lixin Gao 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001
IEEE Trans. Parallel Distributed Syst.4
2017 FSP: towards flexible synchronous parallel framework for expectation-maximization based algorithms on cloud
abstract
Myriad of parameter estimation algorithms can be performed by an Expectation-Maximization (EM) approach. Traditional synchronous frameworks can parallelize these EM algorithms on the cloud to accelerate computation while guaranteeing the convergence. However, expensive synchronization costs pose great challenges for efficiency. Asynchronous solutions have been recently designed to bypass high-cost synchronous barriers but at expense of potentially losing convergence guarantee.
Zhigang Wang 0001, Lixin Gao 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001
SoCC4
2017 An I/O-efficient and adaptive fault-tolerant framework for distributed graph computations
Zhigang Wang 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001, Lixin Gao 0001
Distributed Parallel Databases3
2016 A Fault-Tolerant Framework for Asynchronous Iterative Computations in Cloud Environments
abstract
Many graph algorithms are iterative in nature and can be supported by distributed memory-based systems in a synchronous manner. However, an asynchronous model has been recently proposed to accelerate iterative computations. Nevertheless, it is challenging to recover from failures in such a system, since a typical checkpointing based approach requires many expensive synchronization barriers that largely offset the gains of asynchronous computations.
Zhigang Wang 0001, Lixin Gao 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001
SoCC4
2016 Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing
abstract
Billion-node graphs are rapidly growing in size in many applications such as online social networks. Most graph algorithms generate a large number of messages during iterative computations. Vertex-centric distributed systems usually store graph data and message data on disk to improve scalability. Currently, these distributed systems with disk-resident data take a push-based approach to handle messages. This works well if few messages reside on disk. Otherwise, it is I/O-inefficient due to expensive random writes. By contrast, the existing memory-resident pull-based approach individually pulls messages for each vertex on demand. Although it can be used to avoid disk operations regarding messages, expensive I/O costs are incurred by random and frequent access to vertices.
Zhigang Wang 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001, Jeffrey Xu Yu
SIGMOD Conference3
2014 Label and Distance-Constraint Reachability Queries in Uncertain Graphs
Minghan Chen 0001, Yu Gu 0002, Yubin Bao, Ge Yu 0001
DASFAA (1)3
2014 Efficient Graph Similarity Join with Scalable Prefix-Filtering Using MapReduce
Jun Pang 0002, Yu Gu 0002, Jia Xu 0005, Yubin Bao, Ge Yu 0001
WAIM4
2012 Introducing SaaS Capabilities to Existing Web-Based Applications Automatically
Jie Song 0001, Zhenxing Yan, Yubin Bao, Zhiliang Zhu 0001
APWeb4
2012 Scalable Complex Event Processing on Top of MapReduce
Jiaxue Yang, Yu Gu 0002, Yubin Bao, Ge Yu 0001
APWeb3
2011 Requirement-Based Query and Update Scheduling in Real-Time Data Warehouses
Fangling Leng, Yubin Bao, Ge Yu 0001, Jingang Shi, Xiaoyan Cai
WAIM2
2010 Partitioned Dimension: Modeling the Numerical Dimension in Data Warehouse
abstract
In the traditional data warehouse modeling, a numerical attribute with continuous data are not suitable for modeling as a dimension, unless it is partitioned into a concept hierarchy according to some predefined mapping-rules. In some cases, such rules are flexible or unavailable, so the partitioned dimension is proposed in this paper as the solution of modeling the numerical dimension in these cases. Partitioned dimension is generated by clustering the frequent query conditions. The query and dynamic OLAP operations over partitioned dimensions are also introduced. At last, some experiments show that the proposed modeling approach is effective and efficient.
Jie Song 0001, Yubin Bao
APWeb2
2010 A Multilevel and Domain-Independent Duplicate Detection Model for Scientific Database
Jie Song 0001, Yubin Bao, Ge Yu 0001
WAIM2
2007 A Clustered Dwarf Structure to Speed Up Queries on Data Cubes
Fangling Leng, Yubin Bao, Daling Wang, Ge Yu 0001
DaWaK2
2006 An Efficient Indexing Technique for Computing High Dimensional Data Cubes
Fangling Leng, Yubin Bao, Ge Yu 0001, Daling Wang
WAIM2
2005 ESPClust: An Effective Skew Prevention Method for Model-Based Document Clustering
Ge Yu 0001, Daling Wang, Yubin Bao
CICLing4
2005 Evaluating Document-to-Document Relevance Based on Document Language Model: Modeling, Implementation and Performance Evaluation
Ge Yu 0001, Yubin Bao, Daling Wang
CICLing3
2005 An Optimized K-Means Algorithm of Reducing Cluster Intra-dissimilarity for Document Clustering
Daling Wang, Ge Yu 0001, Yubin Bao
WAIM3
2004 Performance Optimization of Fractal Dimension Based Feature Selection Algorithm
Yubin Bao, Ge Yu 0001, Huanliang Sun, Daling Wang
WAIM1
2004 CD-Trees: An Efficient Index Structure for Outlier Detection
Huanliang Sun, Yubin Bao, Faxin Zhao, Ge Yu 0001, Daling Wang
WAIM2
2003 Managing Very Large Document Collections Using Semantics
Guoren Wang, Hongjun Lu, Ge Yu 0001, Yubin Bao
J. Comput. Sci. Technol.4
2002 XBase: making your gigabyte disk queriable
abstract
With the rapid development of the Internet and the World Wide Web (WWW), very large amount of information is available and ready for downloading, most of which are free of charge. At the same time, hard disks with large capacity are available at affordable prices. Most of us nowadays often dump a large number of various types of documents into our computers without much thinking. On the other hand, file systems have not changed too much during the past decades. Most of them organize files in directories that form a tree structure, and a file is identified by its name and pathname in the directory tree. Remembering name of files created sometime ago and digging them out from a disk with dozen gigabytes of data in hundred thousands of files becomes never an easy task. Tools available for helping such a search are still far from satisfactory.Xbase (XML-based document BASE) is a prototype system aiming at addressing the above problem. By XML-based, we meant that XML is used to define the metadata. The current version of XBase stores text-based files, including semi-structured data such as XML, HTML, plain text documents (e.g., tex files, computer programs) and those files that can be converted into text (e.g., postscript files, PDF files). In XBase, file name is optional. Users can just load a file into XBase without giving a name and the directory where it should be stored. XBase will automatically associate it with attributes such as the time when the file was saved, its source, its size and type, and etc., To retrieve those files, XBase provides three access methods, explorative browsing, querying using query languages, and keyword based search.
Hongjun Lu, Guoren Wang, Ge Yu 0001, Yubin Bao, Jianhua Lv, Yaxin Yu
SIGMOD Conference4
2001 An Integrated Classification Rule Management System for Data Mining
Daling Wang, Yubin Bao, Xiao Ji, Guoren Wang, Baoyan Song
WAIM2