VLDB 2026 Research / reviewers in the wild / expert
Yusik Kim
dblp:29/8799
· DBLP profile ↗
8ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4Artificial intelligence and machine learning · 3 · 2 since 2021Theory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionabstractWe introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon. Ahmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Peter W. J. Staar |
ICCV | 8 |
| 2024 | A Benchmark for Rule Induction in Automated Business Decisions
Hagen Völzer, Daniel Horn, Yusik Kim, Greger Ottosson |
RuleML+RR | 3 |
| 2019 | ExaPlan Archive: Data Placement and Provisioning for Large Storage Systems with Archival TiersabstractMany important big data use cases do not require data to be instantly available. Examples are video recordings in TV and film industry, surveillance videos and data from scientific experiments. Archiving such data to high-latency media storage, such as tape and optical disk libraries, results in significant cost savings. In this context, data is accessed by first staging it to low-latency media. However, archiving and staging operations incur additional device and bandwidth costs for both the active and archiving tiers, and might impact user data access performance. For instance, in terms of cost and performance, it is often suboptimal to archive all the data. This paper presents ExaPlan Archive, a scheme to determine the data placement and number of devices required in each tier of a multitiered storage system comprised of archival and active tiers that minimize the latency of the active tiers under budget and staging-time constraints. The efficiency of the proposed optimized archiving scheme is compared with an existing scheme that optimizes multitier storage with only direct-access tiers. The two schemes are evaluated using a staging workload of LOFAR radio telescopes long-term archive for astronomical observation data. Ilias Iliadis, Yusik Kim, Slavisa Sarafijanovic, Vinodh Venkatesan |
MASCOTS | 2 |
| 2017 | Data Prefetching for Large Tiered Storage SystemsabstractIn multi-tier storage systems with large amounts of data, most of the data is stored on inexpensive slower tiers such as cloud or tape to achieve cost savings. This also implies that retrieving the data from the slower storage tiers incurs high latency. Therefore, it would be beneficial to proactively prefetch data from slower tiers to faster tiers by predicting future data accesses. State-of-the-art access prediction methods typically record access history of individual files, data objects, or data segments. However, in systems with large amounts of infrequently accessed (or cold) data, file-level access history is often unavailable for much of the data due to the low frequency of access. In this paper, we extract information from file metadata to predict file accesses in a storage system. The proposed method relies on the hypothesis that users and applications access data stored in the system in a given context and that the context and, therefore, the set of files that are likely to be accessed can be identified by detecting access patterns in file metadata. As an application, we consider the LOFAR radio telescope's long term archive, where the access patterns are learned based on a rich set of metadata, and these patterns are then used to make predictions as to likely future accesses by the astronomers. Giovanni Cherubini, Yusik Kim, Mark A. Lantz, Vinodh Venkatesan |
ICDM | 2 |
| 2017 | ExaPlan: Efficient Queueing-Based Data Placement, Provisioning, and Load Balancing for Large Tiered Storage SystemsabstractMulti-tiered storage, where each tier consists of one type of storage device (e.g., SSD, HDD, or disk arrays), is a commonly used approach to achieve both high performance and cost efficiency in large-scale systems that need to store data with vastly different access characteristics. By aligning the access characteristics of the data, either fixed-sized extents or variable-sized files, to the characteristics of the storage devices, a higher performance can be achieved for any given cost. This article presents ExaPlan, a method to determine both the data-to-tier assignment and the number of devices in each tier that minimize the system’s mean response time for a given budget and workload. In contrast to other methods that constrain or minimize the system load, ExaPlan directly minimizes the system’s mean response time estimated by a queueing model. Minimizing the mean response time is typically intractable as the resulting optimization problem is both nonconvex and combinatorial in nature. ExaPlan circumvents this intractability by introducing a parameterized data placement approach that makes it a highly scalable method that can be easily applied to exascale systems. Through experiments that use parameters from real-world storage systems, such as CERN and LOFAR, it is demonstrated that ExaPlan provides solutions that yield lower mean response times than previous works. It supports standalone SSDs and HDDs as well as disk arrays as storage tiers, and although it uses a static workload representation, we provide empirical evidence that underlying dynamic workloads have invariant properties that can be deemed static for the purpose of provisioning a storage system. ExaPlan is also effective as a load-balancing tool used for placing data across devices within a tier, resulting in an up to 3.6-fold reduction of response time compared with a traditional load-balancing algorithm, such as the Longest Processing Time heuristic. Ilias Iliadis, Jens Jelitto, Yusik Kim, Slavisa Sarafijanovic, Vinodh Venkatesan |
ACM Trans. Storage | 3 |
| 2016 | Performance Evaluation of a Tape Library SystemabstractData with vastly different access characteristics is efficiently stored in multi-tiered storage systems. A cost-effective way to retain large volumes of infrequently accessed data is to store it on tape. Steady developments in tape technology deliver ever increasing storage capacities at low cost. This has established tape as a viable solution to cope with the extreme data growth in the context of Big Data. Assessing the performance of the various tiers is central to achieving appropriate tier dimensioning and storage provisioning. To that end, we develop an analytical model to evaluate the performance of a tape library system that considers various relevant aspects, such as the number of cartridges and tape drives as well as different mount/unmount policies. Closed-form expressions for the corresponding mean waiting times are derived. The validity of the model developed is confirmed by demonstrating that the predicted performance matches well with that obtained by simulation across a wide range of system parameter values. Ilias Iliadis, Yusik Kim, Slavisa Sarafijanovic, Vinodh Venkatesan |
MASCOTS | 2 |
| 2015 | ExaPlan: Queueing-Based Data Placement and Provisioning for Large Tiered Storage SystemsabstractMulti-tiered storage, where each tier comprises one type of storage device, e.g., SSD, HDD, is a commonly used approach to achieve both high performance and cost efficiency in large-scale systems that need to store data with vastly different access characteristics. By aligning the access characteristics of the data to the characteristics of the storage devices, higher performance can be achieved for any given cost. This article presents ExaPlan, a method to determine both the data-to-tier assignment and the number of devices in each tier that minimize the system's mean response time for a given budget and workload. In contrast to other methods that constrain or minimize the system load, ExaPlan directly minimizes the system's mean response time estimated by a queueing model. Minimizing the mean response time is typically intractable as the resulting optimization problem is both non-convex and combinatorial in nature. ExaPlan circumvents this intractability by introducing a parameterized data-placement approach that makes it a highly scalable method that can be easily applied to exascale systems. Through experiments that use parameters from real-world storage systems, such as CERN and LOFAR, it is demonstrated that ExaPlan provides solutions that yield lower mean response times than previous works. It is also capable of determining a data-to-tier assignment both at the level of files and at the level of fixed-size extents. For some of the workloads evaluated, file-level placement exhibited a significant performance improvement over extent-level placement. Ilias Iliadis, Jens Jelitto, Yusik Kim, Slavisa Sarafijanovic, Vinodh Venkatesan |
MASCOTS | 3 |
| 2011 | On Power-Law Distributed Balls in Bins and Its Applications to View Size Estimation
Ioannis Atsonios, Olivier Beaumont, Nicolas Hanusse, Yusik Kim |
ISAAC | 4 |