Zexin Xia

dblp:242/5161 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 55% Indexing and storage engines · 27% Data models and query languages · 18%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › indexing
search engine indexing
1.022023
A High-Performance Index for Real-Time Matrix Retrieval (Extended Abstract) · ICDE 2023
Aucher: Multi-modal Queries on Live Audio Streams in Real-Time · ICDE 2019
Indexing and storage engines
LSM-tree
1.022022
A High-Performance Index for Real-Time Matrix Retrieval · IEEE Trans. Knowl. Data Eng. 2022
Aucher: Multi-modal Queries on Live Audio Streams in Real-Time · ICDE 2019
Information retrieval › multimodal retrieval
multi-modal query processing
0.412019
Aucher: Multi-modal Queries on Live Audio Streams in Real-Time · ICDE 2019
Information retrieval › web search
real-time search
0.412019
Aucher: Multi-modal Queries on Live Audio Streams in Real-Time · ICDE 2019
Multimedia analysis and retrieval
audio retrieval
0.412019
Aucher: Multi-modal Queries on Live Audio Streams in Real-Time · ICDE 2019
Information retrieval › indexing
inverted index
0.212022
A High-Performance Index for Real-Time Matrix Retrieval · IEEE Trans. Knowl. Data Eng. 2022

Methods — techniques the papers use, named apart from their topics

voice search · 0.8log-structured merge-tree · 0.8embedding techniques · 0.7vector signature · 0.6hashing · 0.6LSM-tree · 0.6
YearPublicationVenuePosition
2023 A High-Performance Index for Real-Time Matrix Retrieval (Extended Abstract)
abstract
Embedding techniques can be used to represent words using word embedding [1] , images using image-to-vector techniques [2] , [3] and even database queries [4] . As a result, many more real-world objects can be represented by matrices. For example, a matrix can represent a document where each row (i.e., each vector) of the matrix stands for a word in the document. Figure 1 shows the key steps of representing an object (e.g., a document, a video or an audio stream) by a matrix. The intermediate step is to divide the object into small pieces and to convert the small pieces into vectors. The vectors of the object are then put together to form a matrix. These objects represented by matrices require new data management systems to support efficient indexing and retrieval.
Zeyi Wen, Mingyu Liang, Bingsheng He, Zexin Xia
ICDE4
2022 A High-Performance Index for Real-Time Matrix Retrieval
abstract
With the embedding techniques, many real-world objects can be represented using matrices. For example, a document can be represented by a matrix, where each row of the matrix represents a word. On the other hand, we have witnessed that many applications continuously generate new data represented by matrices and require real-time query answering on the data. These continuously generated matrices need to be well managed for efficient retrieval. In this paper, we propose an index for real-time matrix retrieval. Besides fast query response, the index also supports real-time insertion by exploiting the LSM-tree. Since the index is built for matrices, it consumes much more memory and requires much more time to search than the traditional index for information retrieval. To tackle the challenges, we power our proposed index with precise and fuzzy inverted lists, and propose a series of novel techniques to improve the memory consumption and the search efficiency of the index. The proposed techniques include vector signature, vector residual sorting, hashing based lookup, and dictionary initialization to guarantee the index quality. Comprehensive experimental results show that our proposed index can support real-time search on matrices and is more efficient than the state-of-the-art method.
Zeyi Wen, Mingyu Liang, Bingsheng He, Zexin Xia
IEEE Trans. Knowl. Data Eng.4
2019 Aucher: Multi-modal Queries on Live Audio Streams in Real-Time
abstract
This paper demonstrates a real-time search system called Aucher for live audio streams. Audio streaming services (e.g., Mixlr, Ximalaya, Lizhi and Facebook Live Audio) have become increasingly popular with the wide use of smart phones. Because of the popularity of audio broadcasting, the data volume of live audio streams is also ever increasing. Searching and indexing these audio streams is an important and challenging problem. Aucher is a system prototype which can support both voice search and keyword search on audio streams. We achieve the real-time response for queries by our novel index which exploits log structured merge-trees and supports multi-modal search. Moreover, our system can handle insertion about four times faster and more memory efficient than the state-of-the-art solution. We plan to demonstrate searching live audio streams by keywords and voice, illustrate the trade-off of freshness, popularity and relevance on query results, perform searching hot terms, and show the ability of searching live audio streams in real-time.
Zeyi Wen, Mingyu Liang, Bingsheng He, Zexin Xia, Bo Li 0001
ICDE4