VLDB 2026 Research / reviewers in the wild / expert
Albert Kim
dblp:65/2692
· DBLP profile ↗
12ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 82% Cloud and datacenter computing · 18% | |
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 78% Data integration and cleaning · 12% Data models and query languages · 10% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query optimization › logical query optimization
disjunctive query optimization |
0.8 | 1 | 2024 | Optimizing Disjunctive Queries with Tagged Execution · Proc. ACM Manag. Data 2024 |
Query processing and optimization › query optimization › predicate optimization
predicate pushdown |
0.8 | 1 | 2024 | Optimizing Disjunctive Queries with Tagged Execution · Proc. ACM Manag. Data 2024 |
Cloud and datacenter computing
cloud storage |
0.7 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Storage systems › distributed storage
disaggregated storage |
0.7 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Storage systems › file systems
distributed file system |
0.7 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Storage systems
key-value storage |
0.7 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Storage systems › key-value storage
LSM-tree |
0.7 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Query processing and optimization › similarity join
set similarity join |
0.3 | 1 | 2017 | SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching Constraints · Proc. VLDB Endow. 2017 |
Data models and query languages › query language
visual query language |
0.2 | 1 | 2016 | Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016 |
Visualization and visual analytics › interactive data exploration
visual exploration |
0.2 | 1 | 2016 | Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016 |
Visualization and visual analytics
approximate visualization |
0.2 | 1 | 2015 | Rapid Sampling for Visualizations with Ordering Guarantees · Proc. VLDB Endow. 2015 |
Visualization and visual analytics › data reduction
sampling for visualization |
0.2 | 1 | 2015 | Rapid Sampling for Visualizations with Ordering Guarantees · Proc. VLDB Endow. 2015 |
Storage systems
crash recovery |
0.2 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Storage systems
storage reliability |
0.2 | 1 | 2023 | Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023 |
Query processing and optimization
query optimization |
0.1 | 1 | 2016 | Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016 |
Automata and formal languages
parsing |
0.0 | 1 | 1994 | Graded Unification: A Framework for Interactive Processing · ACL 1994 |
Logic in computer science
unification |
0.0 | 1 | 1994 | Graded Unification: A Framework for Interactive Processing · ACL 1994 |
Methods — techniques the papers use, named apart from their topics
tag generalization · 0.8append-only distributed file system · 0.7triangle inequality · 0.3signature-based filtering · 0.3theoretical analysis · 0.2sampling · 0.2unification · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Disjunctive Queries with Tagged ExecutionabstractDespite decades of research into query optimization, optimizing queries with disjunctive predicate expressions remains a challenge. Solutions employed by existing systems (if any) are often simplistic and lead to much redundant work being performed by the execution engine. To address these problems, we propose a novel form of query execution called tagged execution. Tagged execution groups tuples into subrelations based on which predicates in the query they satisfy (or don't satisfy) and tags them with that information. These tags then provide additional context for query operators to take advantage of during runtime, allowing them to eliminate much of the redundant work performed by traditional engines and realize predicate pushdown optimizations for disjunctive predicates. However, tagged execution brings its own challenges, and the question of what tags to create is a nontrivial one. Careless creation of tags can lead to an exponential blowup in the tag space, with the overhead outweighing the benefits. To address this issue, we present a technique called tag generalization to minimize the space of tags. We implemented the tagged execution model with tag generalization in our system Basilisk, and our evaluation showed an average 2.7x speedup in runtime over the traditional execution model with up to a 19x speedup in certain situations. Albert Kim, Samuel Madden 0001 |
Proc. ACM Manag. Data | 1 |
| 2023 | Disaggregating RocksDB: A Production ExperienceabstractAs in the general industry, there is a trend in Meta's data centers to migrate data from locally attached SSDs to cloud storage. We extended RocksDB [26], a widely used open-source storage engine designed and built for local SSDs, to leverage disaggregated storage. RocksDB's design, such as its data and log files' access patterns, makes an append-only distributed file system a desirable underlying storage. At Meta, we built disaggregated RocksDB using Tectonic File System [35], which so far had mainly been used for our data warehouse and blob storage stacks. We identified that metadata overhead and tail latencies were Tectonic's major performance gaps and addressed them accordingly. We improved the reliability, performance and other requirements with both general and customized optimizations to the core engine in RocksDB. We also took the time to deeply understand the common challenges presented by applications running on RocksDB and implemented enhancements to address them. This architecture enabled RocksDB to adapt to a more distributed architecture for performance enhancements. Siying Dong, Shiva Shankar P., Satadru Pan, Anand Ananthabhotla, Dhanabal Ekambaram, Shobhit Dayal, Nishant Vinaybhai Parikh, Yanqin Jin, Albert Kim, Sushil Patil, Jay Zhuang, Sam Dunster, Akanksha Mahajan 0001, Anirudh Chelluri, Chaitanya Datye, Lucas Vasconcelos Santana, Omkar Gawde |
Proc. ACM Manag. Data | 10 |
| 2022 | Characterization of Magnetic Communication Through Human BodyabstractBiomedical systems of implanted miniaturized sensors and actuators interconnected into an intra-body area net-work could revolutionize treatment options for chronic diseases afflicting internal organs. Considering the well-understood limitations of radio frequency (RF) propagation in the human body, we have explored magnetic resonance (MR) coupling for both communications and energy transfer through the body. In this paper, we have discussed the design and implementation of a software-defined prototype using Universal Software Radio Peripheral (USRP) boards. We have reported experimental results on the achieved packet error rates at different positions through-the-body distances and packet sizes. We have observed experimentally that the MR signal propagates through the body substantially better than in the air, and can provide a practical means for energy transfer and communications in intra-body networks. It also works better than the better understood galvanic coupling. Rajpreet Kaur Gulati, Sayemul Islam, Amitangshu Pal, Krishna Kant 0001, Albert Kim |
CCNC | 5 |
| 2017 | Fast-Forwarding to Desired Visualizations with Zenvisage
Tarique Siddiqui, John Lee 0005, Albert Kim, Edward Xue, Xiaofo Yu, Sean Zou, Lijin Guo, Changfeng Liu, Chaoran Wang, Karrie Karahalios, Aditya G. Parameswaran |
CIDR | 3 |
| 2017 | SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching ConstraintsabstractDetermining if two sets are related - that is, if they have similar values or if one set contains the other -- is an important problem with many applications in data cleaning, data integration, and information retrieval. For example, set relatedness can be a useful tool to discover whether columns from two different databases are joinable; if enough of the values in the columns match, it may make sense to join them. A common metric is to measure the relatedness of two sets by treating the elements as vertices of a bipartite graph and calculating the score of the maximum matching pairing between elements. Compared to other metrics which require exact matchings between elements, this metric uses a similarity function to compare elements between the two sets, making it robust to small dissimilarities in elements and more useful for real-world, dirty data. Unfortunately, the metric suffers from expensive computational cost, taking O ( n 3 ) time, where n is the number of elements in the sets, for each set-to-set comparison. Thus for applications that try to search for all pairings of related sets in a brute-force manner, the runtime becomes unacceptably large. To address this challenge, we developed S ilk M oth , a system capable of rapidly discovering related set pairs in collections of sets. Internally, S ilk M oth creates a signature for each set, with the property that any other set which is related must match the signature. S ilk M oth then uses these signatures to prune the search space, so only sets that match the signatures are left as candidates. Finally, S ilk M oth applies the maximum matching metric on remaining candidates to verify which of these candidates are truly related sets. An important property of S ilk M oth is that it is guaranteed to output exactly the same related set pairings as the brute-force method, unlike approximate techniques. Thus, a contribution of this paper is the characterization of the space of signatures which enable this property. We show that selecting the optimal signature in this space is NP-complete, and based on insights from the characterization of the space, we propose two novel filters which help to prune the candidates further before verification. In addition, we introduce a simple optimization to the calculation of the maximum matching metric itself based on the triangle inequality. Compared to related approaches, S ilk M oth is much more general, handling a larger space of similarity functions and relatedness metrics, and is an order of magnitude more efficient on real datasets. Dong Deng 0001, Albert Kim, Samuel Madden 0001, Michael Stonebraker |
Proc. VLDB Endow. | 2 |
| 2016 | Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics SystemabstractData visualization is by far the most commonly used mechanism to explore and extract insights from datasets, especially by novice data scientists. And yet, current visual analytics tools are rather limited in their ability to operate on collections of visualizations---by composing, filtering, comparing, and sorting them---to find those that depict desired trends or patterns. The process of visual data exploration remains a tedious process of trial-and-error. We propose zenvisage, a visual analytics platform for effortlessly finding desired visual patterns from large datasets. We introduce zenvisage's general purpose visual exploration language, ZQL ("zee-quel") for specifying the desired visual patterns, drawing from use-cases in a variety of domains, including biology, mechanical engineering, climate science, and commerce. We formalize the expressiveness of ZQL via a visual exploration algebra---an algebra on collections of visualizations---and demonstrate that ZQL is as expressive as that algebra. zenvisage exposes an interactive front-end that supports the issuing of ZQL queries, and also supports interactions that are "short-cuts" to certain commonly used ZQL queries. To execute these queries, zenvisage uses a novel ZQL graph-based query optimizer that leverages a suite of optimizations tailored to the goal of processing collections of visualizations in certain pre-defined ways. Lastly, a user survey and study demonstrates that data scientists are able to effectively use zenvisage to eliminate error-prone and tedious exploration and directly identify desired visualizations. Tarique Siddiqui, Albert Kim, John Lee 0005, Karrie Karahalios, Aditya G. Parameswaran |
Proc. VLDB Endow. | 2 |
| 2015 | Generalizing Over Uncertain Dynamics for Online Trajectory Generation
Albert Kim, Hongkai Dai, Leslie Pack Kaelbling, Tomás Lozano-Pérez |
ISRR (2) | 2 |
| 2015 | Rapid Sampling for Visualizations with Ordering GuaranteesabstractVisualizations are frequently used as a means to understand trends and gather insights from datasets, but often take a long time to generate. In this paper, we focus on the problem of rapidly generating approximate visualizations while preserving crucial visual properties of interest to analysts. Our primary focus will be on sampling algorithms that preserve the visual property of ordering ; our techniques will also apply to some other visual properties. For instance, our algorithms can be used to generate an approximate visualization of a bar chart very rapidly, where the comparisons between any two bars are correct. We formally show that our sampling algorithms are generally applicable and provably optimal in theory, in that they do not take more samples than necessary to generate the visualizations with ordering guarantees. They also work well in practice, correctly ordering output groups while taking orders of magnitude fewer samples and much less time than conventional sampling schemes. Albert Kim, Eric Blais, Aditya G. Parameswaran, Piotr Indyk, Samuel Madden 0001, Ronitt Rubinfeld |
Proc. VLDB Endow. | 1 |
| 2013 | The neural computation of scalar implicature
Joshua K. Hartshorne, Jesse Snedeker, Albert Kim |
CogSci | 3 |
| 2013 | Detecting miRNAs in deep-sequencing data: a software performance comparison and evaluationabstractDeep sequencing has become a popular tool for novel miRNA detection but its data must be viewed carefully as the state of the field is still undeveloped. Using three different programs, miRDeep (v1, 2), miRanalyzer and DSAP, we have analyzed seven data sets (six biological and one simulated) to provide a critical evaluation of the programs performance. We selected these software based on their popularity and overall approach toward the detection of novel and known miRNAs using deep-sequencing data. The program comparisons suggest that, despite differing stringency levels they all identify a similar set of known and novel predictions. Comparisons between the first and second version of miRDeep suggest that the stringency level of each of these programs may, in fact, be a result of the algorithm used to map the reads to the target. Different stringency levels are likely to affect the number of possible novel candidates for functional verification, causing undue strain on resources and time. With that in mind, we propose that an intersection across multiple programs be taken, especially if considering novel candidates that will be targeted for additional analysis. Using this approach, we identify and performed initial validation of 12 novel predictions in our in-house data with real-time PCR, six of which have been previously unreported. Vernell S. Williamson, Albert Kim, Bin Xie 0004, G. Omari McMichael, Vladimir I. Vladimirov |
Briefings Bioinform. | 2 |
| 2011 | Tessellation operating system: Building a real-time, responsive, high-throughput client OS for many-core architectures
Juan A. Colmenares, Sarah Bird, Gage Eads, Steven Hofmeyr, Albert Kim, Rohit Poddar, Hilfi Alkaff, Krste Asanovic, John Kubiatowicz |
Hot Chips Symposium | 5 |
| 1994 | Graded Unification: A Framework for Interactive ProcessingabstractAn extension to classical unification, called graded unification is presented. It is capable of combining contradictory information. An interactive processing paradigm and parser based on this new operator are also presented. Albert Kim |
ACL | 1 |