VLDB 2026 Research / reviewers in the wild / expert
Yudi Yang
dblp:05/5462
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Static Analysis for Efficient Streaming TokenizationabstractTokenization, also referred to as lexing or scanning, is the computational task of partitioning an input text into a sequence of substrings called tokens. Tokenization is one of the first stages of program compilation, it is used in natural language processing, and it is also useful for processing unstructured text or semi-structured data such as JSON, CSV, and XML. A tokenizer is typically specified as a list of regular expressions, which is called a tokenization grammar. Each regular expression describes a class of tokens (e.g., integer, floating-point number, variable identifier, string literal). The semantics of tokenization employs the longest match policy to disambiguate among the possible choices. This policy says that we should prefer a longer token over a shorter one. It is also known as the maximal munch policy. Angela W. Li, Yudi Yang, Konstantinos Mamouras |
ASPLOS (2) | 2 |
| 2026 | An Efficient Algorithm for Streaming BPE TokenizationabstractTokenization is an essential text preprocessing step in almost all large language models (LLMs), and byte-pair encoding (BPE) is a popular tokenization method used by models such as GPT, GPT-2, and RoBERTa. Since LLMs have many applications that require the fast processing of large amounts of data (e.g., real-time document summarization and analysis), offline tokenization algorithms may have high latency or prohibitive memory requirements. In this paper, we study BPE tokenization with a focus on providing a streaming implementation. A BPE tokenizer is specified with an ordered list of token merge rules, where each rule describes the merging of two adjacent tokens. We introduce the concept of delay for a list of BPE merge rules, which corresponds to the amount of lookahead needed before tokens can be finalized. We view BPE tokenization as a sequence-to-sequence transduction and show how to obtain a bound on delay. This bound enables streaming tokenization with a low memory footprint. We propose a novel streaming algorithm that uses a small amount of memory (independent of the input text) and has linear time complexity in the length of the input text. Our experimental evaluation shows that our algorithm performs well in comparison to existing BPE tokenizers. Konstantinos Mamouras, Angela W. Li, Yudi Yang |
Proc. ACM Program. Lang. | 3 |
| 2025 | A 23-29 GHz 3-stack Power Amplifier in 22nm FD-SOI CMOS TechnologyabstractA 23–29 GHz power amplifier (PA) based on a stacked field-effect transistor (FET) topology is proposed. To achieve higher output power and gain, the output stage employs a triple-stacked FET design. To improve stacking efficiency, neutralization capacitors are added between the differential common-source (CS) amplifiers to compensate for phase mismatches at the stack nodes. Round-table and vertical gate repetition techniques are applied to achieve a compact layout and enhance the PA’s power density. The proposed PA is fabricated using the 22nm FD-SOI process. Measurement results show that the proposed PA delivers a peak saturated power (Psat) of 19.6 dBm, with a core area of only 0.047 mm2, achieving a power density of 1936.2 mW/mm2. Jiewen Wang, Yudi Yang, Wenhua Chen 0002, Zhenghe Feng |
ISCAS | 3 |
| 2021 | Deep Multiple Auto-Encoder-Based Multi-view ClusteringabstractAbstract Multi-view clustering (MVC), which aims to explore the underlying structure of data by leveraging heterogeneous information of different views, has brought along a growth of attention. Multi-view clustering algorithms based on different theories have been proposed and extended in various applications. However, most existing MVC algorithms are shallow models, which learn structure information of multi-view data by mapping multi-view data to low-dimensional representation space directly, ignoring the nonlinear structure information hidden in each view, and thus, the performance of multi-view clustering is weakened to a certain extent. In this paper, we propose a deep multi-view clustering algorithm based on multiple auto-encoder, termed MVC-MAE, to cluster multi-view data. MVC-MAE adopts auto-encoder to capture the nonlinear structure information of each view in a layer-wise manner and incorporate the local invariance within each view and consistent as well as complementary information between any two views together. Besides, we integrate the representation learning and clustering into a unified framework, such that two tasks can be jointly optimized. Extensive experiments on six real-world datasets demonstrate the promising performance of our algorithm compared with 15 baseline algorithms in terms of two evaluation metrics. Guowang Du, Lihua Zhou, Yudi Yang, Kevin Lü 0001, Lizhen Wang 0001 |
Data Sci. Eng. | 3 |
| 2019 | Meta Path-Based Information Entropy for Modeling Social Influence in Heterogeneous Information NetworksabstractInfluence is a complex and subtle force that changes the behavior of involved users. Measuring influence can benefit to identify the influential users, and also benefit to provide important insights into the design of social platforms and applications. However, most existing work on social influence analysis has focused on homogeneous information networks. Few studies systematically investigate how to mine the strength of influence between nodes in heterogeneous information networks. In this paper, we present a meta path-based information entropy for modeling social influence in heterogeneous information networks (MPIE). Through setting meta paths, MPIE not only flexibly integrates heterogeneous information, but also obtains potential link information to measure the influence of nodes. Experiments on real data sets demonstrate the effectiveness of our proposed method. Yudi Yang, Lihua Zhou, Jinhua Yang |
MDM | 1 |
| 2007 | Metabolic network properties help assign weights to elementary modes to understand physiological flux distributionsabstractMOTIVATION: Elementary modes (EMs) analysis has been well established. The existing methodologies for assigning weights to EMs cannot be directly applied for large-scale metabolic networks, since the tremendous number of modes would make the computation a time-consuming or even an impossible mission. Therefore, developing more efficient methods to deal with large set of EMs is urgent. RESULT: We develop a method to evaluate the performance of employing a subset of the elementary modes to reconstruct a real flux distribution by using the relative error between the real flux vector and the reconstructed one as an indicator. We have found a power function relationship between the decrease of relative error and the increase of the number of the selecting EMs, and a logarithmic relationship between the increases of the number of non-zero weighted EMs and that of the number of the selecting EMs. Our discoveries show that it is possible to reconstruct a given flux distribution by a selected subset of EMs from a large metabolic network and furthermore, they help us identify the 'governing modes' to represent the cellular metabolism for such a condition. Qingzhao Wang, Yudi Yang, Hongwu Ma, Xue-Ming Zhao |
Bioinform. | 2 |