Ermu Qiu

dblp:414/5278 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0003-5366-293XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 39% Data integration and cleaning · 30% Distributed and cloud data management · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
data lake
0.912025
LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search · ICDE 2025
Information retrieval › search engines › structured data search
table retrieval
0.912025
LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search · ICDE 2025
Data integration and cleaning › table discovery
table union search
0.912025
LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search · ICDE 2025
Information retrieval › similarity search
vector retrieval
0.312025
LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search · ICDE 2025

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 0.9pre-trained language model · 0.9pattern encoder · 0.9number encoder · 0.9cross-attention · 0.9
YearPublicationVenuePosition
2025 LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search
abstract
Table union search (TUS) represents a fundamental operation in data lakes to find tables unionable to the given one. Recent approaches to TUS mainly learn column representations for searching by introducing Pre-trained Language Models (PLMs), especially on columns with linguistic data. However, a significant amount of non-linguistic data, notably represented by domain-specific strings and numerical data in the data lake, are still under-explored in the existing methods. To address this issue, we propose LIFTus, an adaptive multi-aspect column representation for table unionable search, where aspect refers to a concept more flexible than data types, so that a single column can exhibit multiple aspects simultaneously. LIFTus aims at combining different aspects of a column (including both linguistic and non-linguistic aspects) to promote the effectiveness and generalization of TUS in a self-supervised manner. Specifically, besides employing PLMs to extract the linguistic aspects from an individual column, LIFTus trains a pattern encoder to learn possible character-level sequential patterns for the column, and builds a number encoder to capture numerical aspects of the column, including the distribution and magnitude features. LIFTus further utilizes a hierarchical cross-attention aided by aspect-relevant statistics to combine these aspects adaptively in producing the final column representations, which are indexed by vector retrieval techniques to achieve efficient search. Extensive experimental results demonstrate that LIFTus has outperformed the current state-of-the-art methods in terms of effectiveness, and achieved much better generalization capability to support unseen data.
Ermu Qiu, Jun Gao 0003, Yaofeng Tu, Jingru Yang
ICDE1