Albert Kim

dblp:65/2692 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 82% Cloud and datacenter computing · 18%
Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 78% Data integration and cleaning · 12% Data models and query languages · 10%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query optimization › logical query optimization
disjunctive query optimization
0.812024
Optimizing Disjunctive Queries with Tagged Execution · Proc. ACM Manag. Data 2024
Query processing and optimization › query optimization › predicate optimization
predicate pushdown
0.812024
Optimizing Disjunctive Queries with Tagged Execution · Proc. ACM Manag. Data 2024
Cloud and datacenter computing
cloud storage
0.712023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Storage systems › distributed storage
disaggregated storage
0.712023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Storage systems › file systems
distributed file system
0.712023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Storage systems
key-value storage
0.712023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Storage systems › key-value storage
LSM-tree
0.712023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Query processing and optimization › similarity join
set similarity join
0.312017
SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching Constraints · Proc. VLDB Endow. 2017
Data models and query languages › query language
visual query language
0.212016
Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016
Visualization and visual analytics › interactive data exploration
visual exploration
0.212016
Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016
Visualization and visual analytics
approximate visualization
0.212015
Rapid Sampling for Visualizations with Ordering Guarantees · Proc. VLDB Endow. 2015
Visualization and visual analytics › data reduction
sampling for visualization
0.212015
Rapid Sampling for Visualizations with Ordering Guarantees · Proc. VLDB Endow. 2015
Storage systems
crash recovery
0.212023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Storage systems
storage reliability
0.212023
Disaggregating RocksDB: A Production Experience · Proc. ACM Manag. Data 2023
Query processing and optimization
query optimization
0.112016
Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System · Proc. VLDB Endow. 2016
Automata and formal languages
parsing
0.011994
Graded Unification: A Framework for Interactive Processing · ACL 1994
Logic in computer science
unification
0.011994
Graded Unification: A Framework for Interactive Processing · ACL 1994

Methods — techniques the papers use, named apart from their topics

tag generalization · 0.8append-only distributed file system · 0.7triangle inequality · 0.3signature-based filtering · 0.3theoretical analysis · 0.2sampling · 0.2unification · 0.0
YearPublicationVenuePosition
2024 Optimizing Disjunctive Queries with Tagged Execution
abstract
Despite decades of research into query optimization, optimizing queries with disjunctive predicate expressions remains a challenge. Solutions employed by existing systems (if any) are often simplistic and lead to much redundant work being performed by the execution engine. To address these problems, we propose a novel form of query execution called tagged execution. Tagged execution groups tuples into subrelations based on which predicates in the query they satisfy (or don't satisfy) and tags them with that information. These tags then provide additional context for query operators to take advantage of during runtime, allowing them to eliminate much of the redundant work performed by traditional engines and realize predicate pushdown optimizations for disjunctive predicates. However, tagged execution brings its own challenges, and the question of what tags to create is a nontrivial one. Careless creation of tags can lead to an exponential blowup in the tag space, with the overhead outweighing the benefits. To address this issue, we present a technique called tag generalization to minimize the space of tags. We implemented the tagged execution model with tag generalization in our system Basilisk, and our evaluation showed an average 2.7x speedup in runtime over the traditional execution model with up to a 19x speedup in certain situations.
Albert Kim, Samuel Madden 0001
Proc. ACM Manag. Data1
2023 Disaggregating RocksDB: A Production Experience
abstract
As in the general industry, there is a trend in Meta's data centers to migrate data from locally attached SSDs to cloud storage. We extended RocksDB [26], a widely used open-source storage engine designed and built for local SSDs, to leverage disaggregated storage. RocksDB's design, such as its data and log files' access patterns, makes an append-only distributed file system a desirable underlying storage. At Meta, we built disaggregated RocksDB using Tectonic File System [35], which so far had mainly been used for our data warehouse and blob storage stacks. We identified that metadata overhead and tail latencies were Tectonic's major performance gaps and addressed them accordingly. We improved the reliability, performance and other requirements with both general and customized optimizations to the core engine in RocksDB. We also took the time to deeply understand the common challenges presented by applications running on RocksDB and implemented enhancements to address them. This architecture enabled RocksDB to adapt to a more distributed architecture for performance enhancements.
Siying Dong, Shiva Shankar P., Satadru Pan, Anand Ananthabhotla, Dhanabal Ekambaram, Shobhit Dayal, Nishant Vinaybhai Parikh, Yanqin Jin, Albert Kim, Sushil Patil, Jay Zhuang, Sam Dunster, Akanksha Mahajan 0001, Anirudh Chelluri, Chaitanya Datye, Lucas Vasconcelos Santana, Omkar Gawde
Proc. ACM Manag. Data10
2022 Characterization of Magnetic Communication Through Human Body
abstract
Biomedical systems of implanted miniaturized sensors and actuators interconnected into an intra-body area net-work could revolutionize treatment options for chronic diseases afflicting internal organs. Considering the well-understood limitations of radio frequency (RF) propagation in the human body, we have explored magnetic resonance (MR) coupling for both communications and energy transfer through the body. In this paper, we have discussed the design and implementation of a software-defined prototype using Universal Software Radio Peripheral (USRP) boards. We have reported experimental results on the achieved packet error rates at different positions through-the-body distances and packet sizes. We have observed experimentally that the MR signal propagates through the body substantially better than in the air, and can provide a practical means for energy transfer and communications in intra-body networks. It also works better than the better understood galvanic coupling.
Rajpreet Kaur Gulati, Sayemul Islam, Amitangshu Pal, Krishna Kant 0001, Albert Kim
CCNC5
2017 Fast-Forwarding to Desired Visualizations with Zenvisage
Tarique Siddiqui, John Lee 0005, Albert Kim, Edward Xue, Xiaofo Yu, Sean Zou, Lijin Guo, Changfeng Liu, Chaoran Wang, Karrie Karahalios, Aditya G. Parameswaran
CIDR3
2017 SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching Constraints
abstract
Determining if two sets are related - that is, if they have similar values or if one set contains the other -- is an important problem with many applications in data cleaning, data integration, and information retrieval. For example, set relatedness can be a useful tool to discover whether columns from two different databases are joinable; if enough of the values in the columns match, it may make sense to join them. A common metric is to measure the relatedness of two sets by treating the elements as vertices of a bipartite graph and calculating the score of the maximum matching pairing between elements. Compared to other metrics which require exact matchings between elements, this metric uses a similarity function to compare elements between the two sets, making it robust to small dissimilarities in elements and more useful for real-world, dirty data. Unfortunately, the metric suffers from expensive computational cost, taking O ( n 3 ) time, where n is the number of elements in the sets, for each set-to-set comparison. Thus for applications that try to search for all pairings of related sets in a brute-force manner, the runtime becomes unacceptably large. To address this challenge, we developed S ilk M oth , a system capable of rapidly discovering related set pairs in collections of sets. Internally, S ilk M oth creates a signature for each set, with the property that any other set which is related must match the signature. S ilk M oth then uses these signatures to prune the search space, so only sets that match the signatures are left as candidates. Finally, S ilk M oth applies the maximum matching metric on remaining candidates to verify which of these candidates are truly related sets. An important property of S ilk M oth is that it is guaranteed to output exactly the same related set pairings as the brute-force method, unlike approximate techniques. Thus, a contribution of this paper is the characterization of the space of signatures which enable this property. We show that selecting the optimal signature in this space is NP-complete, and based on insights from the characterization of the space, we propose two novel filters which help to prune the candidates further before verification. In addition, we introduce a simple optimization to the calculation of the maximum matching metric itself based on the triangle inequality. Compared to related approaches, S ilk M oth is much more general, handling a larger space of similarity functions and relatedness metrics, and is an order of magnitude more efficient on real datasets.
Dong Deng 0001, Albert Kim, Samuel Madden 0001, Michael Stonebraker
Proc. VLDB Endow.2
2016 Effortless Data Exploration with zenvisage: An Expressive and Interactive Visual Analytics System
abstract
Data visualization is by far the most commonly used mechanism to explore and extract insights from datasets, especially by novice data scientists. And yet, current visual analytics tools are rather limited in their ability to operate on collections of visualizations---by composing, filtering, comparing, and sorting them---to find those that depict desired trends or patterns. The process of visual data exploration remains a tedious process of trial-and-error. We propose zenvisage, a visual analytics platform for effortlessly finding desired visual patterns from large datasets. We introduce zenvisage's general purpose visual exploration language, ZQL ("zee-quel") for specifying the desired visual patterns, drawing from use-cases in a variety of domains, including biology, mechanical engineering, climate science, and commerce. We formalize the expressiveness of ZQL via a visual exploration algebra---an algebra on collections of visualizations---and demonstrate that ZQL is as expressive as that algebra. zenvisage exposes an interactive front-end that supports the issuing of ZQL queries, and also supports interactions that are "short-cuts" to certain commonly used ZQL queries. To execute these queries, zenvisage uses a novel ZQL graph-based query optimizer that leverages a suite of optimizations tailored to the goal of processing collections of visualizations in certain pre-defined ways. Lastly, a user survey and study demonstrates that data scientists are able to effectively use zenvisage to eliminate error-prone and tedious exploration and directly identify desired visualizations.
Tarique Siddiqui, Albert Kim, John Lee 0005, Karrie Karahalios, Aditya G. Parameswaran
Proc. VLDB Endow.2
2015 Generalizing Over Uncertain Dynamics for Online Trajectory Generation
Albert Kim, Hongkai Dai, Leslie Pack Kaelbling, Tomás Lozano-Pérez
ISRR (2)2
2015 Rapid Sampling for Visualizations with Ordering Guarantees
abstract
Visualizations are frequently used as a means to understand trends and gather insights from datasets, but often take a long time to generate. In this paper, we focus on the problem of rapidly generating approximate visualizations while preserving crucial visual properties of interest to analysts. Our primary focus will be on sampling algorithms that preserve the visual property of ordering ; our techniques will also apply to some other visual properties. For instance, our algorithms can be used to generate an approximate visualization of a bar chart very rapidly, where the comparisons between any two bars are correct. We formally show that our sampling algorithms are generally applicable and provably optimal in theory, in that they do not take more samples than necessary to generate the visualizations with ordering guarantees. They also work well in practice, correctly ordering output groups while taking orders of magnitude fewer samples and much less time than conventional sampling schemes.
Albert Kim, Eric Blais, Aditya G. Parameswaran, Piotr Indyk, Samuel Madden 0001, Ronitt Rubinfeld
Proc. VLDB Endow.1
2013 The neural computation of scalar implicature
Joshua K. Hartshorne, Jesse Snedeker, Albert Kim
CogSci3
2013 Detecting miRNAs in deep-sequencing data: a software performance comparison and evaluation
abstract
Deep sequencing has become a popular tool for novel miRNA detection but its data must be viewed carefully as the state of the field is still undeveloped. Using three different programs, miRDeep (v1, 2), miRanalyzer and DSAP, we have analyzed seven data sets (six biological and one simulated) to provide a critical evaluation of the programs performance. We selected these software based on their popularity and overall approach toward the detection of novel and known miRNAs using deep-sequencing data. The program comparisons suggest that, despite differing stringency levels they all identify a similar set of known and novel predictions. Comparisons between the first and second version of miRDeep suggest that the stringency level of each of these programs may, in fact, be a result of the algorithm used to map the reads to the target. Different stringency levels are likely to affect the number of possible novel candidates for functional verification, causing undue strain on resources and time. With that in mind, we propose that an intersection across multiple programs be taken, especially if considering novel candidates that will be targeted for additional analysis. Using this approach, we identify and performed initial validation of 12 novel predictions in our in-house data with real-time PCR, six of which have been previously unreported.
Vernell S. Williamson, Albert Kim, Bin Xie 0004, G. Omari McMichael, Vladimir I. Vladimirov
Briefings Bioinform.2
2011 Tessellation operating system: Building a real-time, responsive, high-throughput client OS for many-core architectures
Juan A. Colmenares, Sarah Bird, Gage Eads, Steven Hofmeyr, Albert Kim, Rohit Poddar, Hilfi Alkaff, Krste Asanovic, John Kubiatowicz
Hot Chips Symposium5
1994 Graded Unification: A Framework for Interactive Processing
abstract
An extension to classical unification, called graded unification is presented. It is capable of combining contradictory information. An interactive processing paradigm and parser based on this new operator are also presented.
Albert Kim
ACL1