Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xinhui Tian

dblp:124/3447 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0003-3687-7923ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 25% Indexing and storage engines · 25% Database system architecture and tuning · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 25% Cloud and datacenter computing · 25% Distributed systems · 25%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management › cloud database
cloud-native database
0.912025
BlendHouse: A Cloud-Native Vector Database System in ByteHouse · ICDE 2025
Database system architecture and tuning
disaggregated storage and compute
0.912025
BlendHouse: A Cloud-Native Vector Database System in ByteHouse · ICDE 2025
Information retrieval › similarity search
nearest neighbor search
0.912025
BlendHouse: A Cloud-Native Vector Database System in ByteHouse · ICDE 2025
Indexing and storage engines
vector database
0.912025
BlendHouse: A Cloud-Native Vector Database System in ByteHouse · ICDE 2025
Cloud and datacenter computing › big data analytics
batch and stream processing
0.212016
Parallel Processing Systems for Big Data: A Survey · Proc. IEEE 2016
Performance modeling and evaluation › benchmarking › distributed system benchmarking
big data system benchmarking
0.212016
Parallel Processing Systems for Big Data: A Survey · Proc. IEEE 2016
Distributed systems
graph processing systems
0.212016
Parallel Processing Systems for Big Data: A Survey · Proc. IEEE 2016
Parallel and multicore computing › parallel computing
parallel data processing
0.212016
Parallel Processing Systems for Big Data: A Survey · Proc. IEEE 2016

Methods — techniques the papers use, named apart from their topics

survey · 0.2
YearPublicationVenuePosition
2025 BlendHouse: A Cloud-Native Vector Database System in ByteHouse
abstract
The rise of unstructured data retrieval in the AI era has created an urgent need for vector databases that manage high-dimensional vector embeddings and provide efficient vector search capabilities for AI applications. Performance, elasticity, and isolation are the key factors for vector databases to serve modern AI applications effectively. Disaggregation of storage and compute is widely recognized as the most effective approach in both academia and industry. Existing work either redesigns specialized vector databases according to the disaggregated architecture or integrates vector search into generalized databases that already use this architecture. However, challenges still remain in building elastic and efficient vector search systems within the disaggregated architecture, such as higher data fetching latency and the highly stateful nature of vector index, which hinder the system's ability to simultaneously achieve high performance, high elasticity and resource isolation. Additionally, a recent trend has emerged to integrate vector search into general-purpose databases, yet the extensibility and generality of integration methodologies have not been systematically studied. In this paper, we present BlendHouse, a cloud-native and generalized vector database system built on top of the disaggregated storage and computation architecture. BlendHouse achieves high performance, high elasticity and resource isolation simultaneously via a suite of optimizations specific to the vector search workload regarding the disaggregated architecture and the relational database. Experimental results demonstrate that BlendHouse outperforms Milvus and pgvector in terms of read and write performance. The integration methodology illustrated in this paper is extensible and general, paving the way for more powerful data management systems in the AI era.
Zhaojie Niu, Xinhui Tian, Xindong Peng, Xing Chen 0023
ICDE2
2021 AIBench Training: Balanced Industry-Standard AI Training Benchmarking
abstract
Earlier-stage evaluations of a new AI architecture/system need affordable AI benchmarks. Only using a few AI component benchmarks like MLPerf alone in the other stages may lead to misleading conclusions. Moreover, the learning dynamics are not well understood, and the benchmarks' shelf-life is short. This paper proposes a balanced benchmarking methodology. We use real-world benchmarks to cover the factors space that impacts the learning dynamics to the most considerable extent. After performing an exhaustive survey on Internet service AI domains, we identify and implement nineteen representative AI tasks with state-of-the-art models. For repeatable performance ranking (RPR subset) and workload characterization (WC subset), we keep two subsets to a minimum for affordability. We contribute by far the most comprehensive AI training benchmark suite. The evaluations show: (1) AIBench Training (v1.1) outperforms MLPerf Training (v0.7) in terms of diversity and representativeness of model complexity, computational cost, convergent rate, computation, and memory access patterns, and hotspot functions; (2) Against the AIBench full benchmarks, its RPR subset shortens the benchmarking cost by 64%, while maintaining the primary workload characteristics; (3) The performance ranking shows the single-purpose AI accelerator like TPU with the optimized TensorFlow framework performs better than that of GPUs while losing the latter's general support for various AI models. The specification, source code, and performance numbers are available from the AIBench homepage https://www.benchcouncil.org/aibench-training/index.html.
Fei Tang 0003, Wanling Gao, Jianfeng Zhan, Chuanxin Lan, Lei Wang 0004, Chunjie Luo, Zheng Cao 0003, Xingwang Xiong, Zihan Jiang 0006, Tianshu Hao, Fanda Fan, Fan Zhang 0047, Yunyou Huang, Jianan Chen 0003, Mengjia Du, Chen Zheng 0001, Daoyi Zheng, Haoning Tang, Kunlin Zhan, Defei Kong, Chongkang Tan, Xinhui Tian, Yatao Li, Junchao Shao, Xiaoyu Wang 0002, Jiahui Dai, Hainan Ye
ISPASS27
2020 Spark-based parallel calculation of 3D fourier shell correlation for macromolecule structure local resolution estimation
abstract
BACKGROUND: Resolution estimation is the main evaluation criteria for the reconstruction of macromolecular 3D structure in the field of cryoelectron microscopy (cryo-EM). At present, there are many methods to evaluate the 3D resolution for reconstructed macromolecular structures from Single Particle Analysis (SPA) in cryo-EM and subtomogram averaging (SA) in electron cryotomography (cryo-ET). As global methods, they measure the resolution of the structure as a whole, but they are inaccurate in detecting subtle local changes of reconstruction. In order to detect the subtle changes of reconstruction of SPA and SA, a few local resolution methods are proposed. The mainstream local resolution evaluation methods are based on local Fourier shell correlation (FSC), which is computationally intensive. However, the existing resolution evaluation methods are based on multi-threading implementation on a single computer with very poor scalability. RESULTS: This paper proposes a new fine-grained 3D array partition method by key-value format in Spark. Our method first converts 3D images to key-value data (K-V). Then the K-V data is used for 3D array partitioning and data exchange in parallel. So Spark-based distributed parallel computing framework can solve the above scalability problem. In this distributed computing framework, all 3D local FSC tasks are simultaneously calculated across multiple nodes in a computer cluster. Through the calculation of experimental data, 3D local resolution evaluation algorithm based on Spark fine-grained 3D array partition has a magnitude change in computing speed compared with the mainstream FSC algorithm under the condition that the accuracy remains unchanged, and has better fault tolerance and scalability. CONCLUSIONS: In this paper, we proposed a K-V format based fine-grained 3D array partition method in Spark to parallel calculating 3D FSC for getting a 3D local resolution density map. 3D local resolution density map evaluates the three-dimensional density maps reconstructed from single particle analysis and subtomogram averaging. Our proposed method can significantly increase the speed of the 3D local resolution evaluation, which is important for the efficient detection of subtle variations among reconstructed macromolecular structures.
Yongchun Lü, Xinhui Tian, Xiao Shi 0003, Xiaohui Zheng, Xin Gao 0001, Min Xu 0009
BMC Bioinform.3
2017 Towards memory and computation efficient graph processing on spark
abstract
Algorithms for large scale natural graph processing can be categorized into two types based on their value propagation behaviors: the unidirectional value propagation (UVP) algorithms and the bidirectional value propagation (BVP) algorithms. The behavior about how vertices interact with neighbors also differs between two algorithm types, which demands different system design choices. However, current distributed graph processing systems usually try to support both types in one general-purpose framework Such system design can not promise good performance and low resource consumption for both types. Especially, for UVP algorithms, current systems can not guarantee low memory footprint, computation efficiency and communication efficiency at the same time. In this paper, we propose a new graph processing engine on Spark, GraphV, which is specially designed for the unidirectional value propagation algorithms, and can satisfy all the above requirements for this type of algorithms. To retain the generalization for other algorithms, we also build a dual-engine framework by integrating GraphV with Spark's existing graph processing engine GraphX. The main design choices of GraphV include a cheap propagation-related partitioner, an one-step computation model, and a locality-aware local graph layout. According to the experiment results, GraphV is faster than GraphX by the factors of 1.2x-3.1x, with much less resource consumption. The source code of GraphV will be publicly available from http://prof.ict.ac.cn/GraphV.
Xinhui Tian, Yuanqing Guo, Jianfeng Zhan, Lei Wang 0004
IEEE BigData1
2016 Parallel Processing Systems for Big Data: A Survey
abstract
The volume, variety, and velocity properties of big data and the valuable information it contains have motivated the investigation of many new parallel data processing systems in addition to the approaches using traditional database management systems (DBMSs). MapReduce pioneered this paradigm change and rapidly became the primary big data processing system for its simplicity, scalability, and fine-grain fault tolerance. However, compared with DBMSs, MapReduce also arouses controversy in processing efficiency, low-level abstraction, and rigid dataflow. Inspired by MapReduce, nowadays the big data systems are blooming. Some of them follow MapReduce's idea, but with more flexible models for general-purpose usage. Some absorb the advantages of DBMSs with higher abstraction. There are also specific systems for certain applications, such as machine learning and stream data processing. To explore new research opportunities and assist users in selecting suitable processing systems for specific applications, this survey paper will give a high-level overview of the existing parallel data processing systems categorized by the data input as batch processing, stream processing, graph processing, and machine learning processing and introduce representative projects in each category. As the pioneer, the original MapReduce system, as well as its active variants and extensions on dataflow, data access, parameter tuning, communication, and energy optimizations will be discussed at first. System benchmarks and open issues for big data processing will also be studied in this survey.
Yunquan Zhang, Shigang Li 0002, Xinhui Tian, Haipeng Jia, Athanasios V. Vasilakos
Proc. IEEE4
2013 SecMon: A Secure Introspection Framework for Hardware Virtualization
abstract
With the fusion of cloud computing and virtualization technology, system security under virtualization becomes a key point in recent research. As a foundational technology to construct a secure system, virtual machine introspection receives more attention than ever. Almost all of the existing virtual machine monitors take the privileged virtual machine (Domain-0) as the monitoring machine, which ignore the threats brought by Domain-0 because of its huge code base of user-level tools. Besides, para-virtualized machines cannot provide the basic support for popular security applications of Windows operating system. This paper proposes a secure monitoring framework based on hardware virtualization. We use Windows operating system to build a monitoring virtual machine in hardware virtual machine domain, and set up monitoring mechanism in it. In addition, the security of the Windows monitoring machine itself is ensured all through its lifetime-bootstrap and runtime. The experiments show our secure monitoring system performs well in the secure monitoring process. The performance overhead it brings is considered to be acceptable.
Yunwei Gao, Xinhui Tian, Baiming Feng, Yuzhong Sun
PDP3