VLDB 2026 Research / reviewers in the wild / expert
Yifeng Jin
dblp:241/3021
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UIID: Unified intra-modal and inter-modal distillation for image-text retrieval
Yifeng Jin |
Neurocomputing | 5 |
| 2024 | Jade: A High-throughput Concurrent Copying Garbage CollectorabstractGarbage collection (GC) pauses are a notorious issue threatening the latency of applications. To mitigate this problem, state-of-the-art concurrent copying collectors allow GC threads to run simultaneously with application threads (mutators) in nearly all GC phases. However, the design of concurrent copying collectors does not always lead to low application latency. To this end, this work studies the behaviors of mainstream concurrent copying collectors in OpenJDK and mainly focuses on long application pauses under heavy workloads. By analyzing the design of those collectors, this work uncovers that lengthy pre-reclamation cycles (including GC phases before actual memory release), high GC frequency, and large metadata maintenance overhead are major factors for long pauses. Therefore, this work proposes Jade, a concurrent copying collector aiming to achieve both short pauses and high GC efficiency. Compared with existing collectors, Jade provides a group-wise collection mechanism to shorten pre-reclamation cycles while controlling GC frequency. It also embraces a generational heap layout and a single-phase algorithm to maximize young GC's throughput. The evaluation results on representative latency-critical applications show that Jade can reach sub-millisecond-level pauses even under heavy workloads and significantly improve applications' peak throughput compared with state-of-the-art concurrent collectors. Mingyu Wu 0001, Yude Lin, Yifeng Jin, Zhe Li 0037, Hongtao Lyu, Denghui Dong, Haibo Chen 0001, Binyu Zang |
EuroSys | 4 |
| 2024 | Energy-Efficient Fast Data Retrieval Strategy Based on RS Coded Placement in LEO ConstellationabstractLow earth orbit (LEO) constellation network is a crucial component and paradigm for big data applications in the future satellite Internet. However, the characteristic of multi-hop transmission significantly increases the data retrieval delay and energy consumption depending on the inter-satellite communications. To mitigate this issue, this paper explores encoding redundancy by reed-solomon (RS) codes to generate the data and parity blocks, and store them in the satellite nodes of LEO constellation. We propose a fast data retrieval strategy, involving three stages: data collection, parity block placement, and data retrieval. Through the derivations of delay and energy consumption, we find that both are intensively related to the placement of parity blocks. To reduce energy consumption during the data retrieval process, we formulate a minimum problem under the delay constraint, which is an integer nonlinear programming problem. Then, we design an energy-efficient parity block placement based on genetic algorithm (PBP-GA), which is heuristic with fast convergence property. Simulation results show that, PBP-GA achieves a comprehensive performance improvement in average data retrieval delay and total energy consumption, compared to other placement schemes, i.e., random placement, retrieval cluster priority placement, nearby placement and non-coding. Specifically, PBP-GA can find the optimal number of parity blocks in a practical scenario of data retrieval in LEO constellation. Zhineng Wu, Shushi Gu, Qinyu Zhang 0001, Yifeng Jin, Lei Zhang 0202, Wei Xiang 0001 |
VTC Spring | 5 |
| 2023 | Effective and Efficient Lexicographical Order Dependency DiscoveryabstractLexicographical order dependencies state relationships of order between lists of attributes. They naturally model the order-by clauses in SQL queries, and are proven useful in query optimizations concerning sorting. Despite their importance, order dependencies on a dataset are typically unknown and are too costly, if not impossible, to design or discover manually. Techniques for automatic order dependency discovery are recently studied. It is challenging for order dependency discovery to scale well, since it is by nature factorial in the number$m$of attributes and quadratic in the number$n$of tuples. In this article, we adopt a strategy that decouples the impact of$m$from that of$n$, and that still finds all minimal and valid lexicographical order dependencies. We present carefully designed data structures, a host of algorithms and optimizations, and an enhanced strategy combined with multithreaded parallelism, for an efficient implementation. Using a host of real-life and synthetic datasets, we experimentally verify our approach is up to orders of magnitude faster than the state-of-the-art methods, and can deliver better results with an improved definition of minimal attribute lists. Jixuan Chen, Yifeng Jin, Zijing Tan, Weidong Yang 0001, Shuai Ma 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Discovery of Approximate Lexicographical Order DependenciesabstractLexicographical order dependencies (LODs) specify orders between list of attributes, and are proven useful in optimizing SQL queries with order by clauses. To discover hidden dependencies from dirty data in practice, approximate dependency discoveries are actively studied, aiming at automatically discovering dependencies that hold on data with some exceptions. In this paper we study the discovery of approximate LODs. (1) We adapt two error measures, namely$g_1$and$g_3$, to LODs. We prove their desirable properties, present efficient algorithms for computing the measures and related lower and upper bounds, and study the relationship between the two measures. (2) We present an efficient approximate LOD discovery algorithm that is well suited to the two error measures, with a set of pruning rules, optimization techniques and ranking functions. (3) We study techniques for estimating$g_1$by sampling, with high accuracy and far less time. (4) We conduct extensive experiments to verify the effectiveness and scalability of our methods, using both real-life and synthetic data. Yifeng Jin, Zijing Tan, Jixuan Chen, Shuai Ma 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Subcarrier-pair based joint resource allocation for secure cooperative OFDM system with multiple untrusted AF relaysabstractAbstract In this paper, a cooperative orthogonal frequency division multiplexing system with multiple untrusted amplify‐and‐forward relays is considered. Since there exists no direct link between source and user because of the shadowing effect, a positive system secrecy rate cannot be obtained. Addressing this issue, a cooperative user‐aided jamming technique to improve the secrecy performance is employed. With the aim of maximising the system secrecy rate, a joint resource allocation problem of power allocation, relay assignment and subcarrier pairing is formulated. By using the signal‐to‐noise ratio based approach and the Lagrange dual method, this nonconvex problem is solved efficiently. To reduce the calculation complexity, a suboptimal algorithm, which decouples the power allocation with the relay assignment and subcarrier pairing is further proposed. Numerical results are presented to demonstrate the advantage of the two proposed algorithms and the impact of relays' location and number of subcarriers on the secrecy performance. Moreover, it is shown that the secrecy performance of the proposed scheme is decreased with respect to the number of untrusted relays. Yifeng Jin, Xunan Li, Guocheng Lv, Meihui Zhao |
IET Commun. | 1 |
| 2021 | Approximate Order Dependency DiscoveryabstractLexicographical order dependencies (ODs) specify orders between list of attributes, and are proven useful in optimizing SQL queries with order by clauses. To find hidden ODs from dirty data in practice, in this paper we make a first effort to study the approximate OD discovery problem, aiming at automatically discovering ODs that hold on the data with some exceptions. (1) We adapt two error measures to ODs, prove their desirable properties, and present efficient algorithms for computing the measures and related lower and upper bounds. (2) We present an efficient approximate OD discovery algorithm that is well suited to the two error measures, with a set of pruning rules and optimization techniques. (3) We conduct extensive experiments to verify the effectiveness and scalability of our methods, using real-life and synthetic data. Yifeng Jin, Zijing Tan, Weijun Zeng, Shuai Ma 0001 |
ICDE | 1 |
| 2020 | Efficient Bidirectional Order Dependency DiscoveryabstractBidirectional order dependencies state relationships of order between lists of attributes. They naturally model the order-by clauses in SQL queries, and are proved effective in query optimizations concerning sorting. Despite their importance, order dependencies on a dataset are typically unknown and are too costly, if not impossible, to design or discover manually. Techniques for automatic order dependency discovery are recently studied. It is challenging for order dependency discovery to scale well, since it is by nature factorial in the number m of attributes and quadratic in the number n of tuples. In this paper, we adopt a strategy that decouples the impact of m from that of n, and that still finds all minimal valid bidirectional order dependencies. We present carefully designed data structures, a host of algorithms and optimizations, for efficient order dependency discovery. With extensive experimental studies on both real-life and synthetic datasets, we verify our approach significantly outperforms state-of-the-art techniques, by orders of magnitude. Yifeng Jin, Zijing Tan |
ICDE | 1 |