VLDB 2026 Research / reviewers in the wild / expert
Jim Chen
dblp:60/5999
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-3255-9656ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 44% Database system architecture and tuning · 22% Data integration and cleaning · 17% | |
| Artificial intelligence
2 papers |
Video understanding and tracking · 42% Graph learning · 37% Segmentation and scene understanding · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 33% Storage systems · 33% Distributed systems · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 13 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network
graph transformer |
0.8 | 1 | 2024 | GTMGC: Using Graph Transformer to Predict Molecule's Ground-State Conformation · ICLR 2024 |
Bioinformatics and computational biology › structural bioinformatics › molecular structure prediction
molecular conformation prediction |
0.8 | 1 | 2024 | GTMGC: Using Graph Transformer to Predict Molecule's Ground-State Conformation · ICLR 2024 |
Bioinformatics and computational biology › molecular informatics
molecular representation learning |
0.8 | 1 | 2024 | GTMGC: Using Graph Transformer to Predict Molecule's Ground-State Conformation · ICLR 2024 |
Query processing and optimization
parallel query processing |
0.7 | 1 | 2023 | Progressive Partitioning for Parallelized Query Execution in Google's Napa · Proc. VLDB Endow. 2023 |
Storage systems › networked storage › storage networking
NVMe over Fabrics |
0.7 | 1 | 2023 | AIDTN: Towards a Real-Time AI Optimized DTN System With NVMeoF · IEEE Trans. Parallel Distributed Syst. 2023 |
Distributed systems › distributed communication
remote data access |
0.7 | 1 | 2023 | AIDTN: Towards a Real-Time AI Optimized DTN System With NVMeoF · IEEE Trans. Parallel Distributed Syst. 2023 |
High-performance computing › data transfer
wide-area data transfer |
0.7 | 1 | 2023 | AIDTN: Towards a Real-Time AI Optimized DTN System With NVMeoF · IEEE Trans. Parallel Distributed Syst. 2023 |
Data integration and cleaning
data warehouse |
0.5 | 1 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 |
Distributed and cloud data management
geo-distributed data management |
0.5 | 1 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 |
Query processing and optimization
view maintenance |
0.5 | 1 | 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at Google · Proc. VLDB Endow. 2021 |
Computer vision › Video understanding and tracking › video object segmentation
semi-supervised video object segmentation |
0.4 | 1 | 2020 | Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement · NeurIPS 2020 |
Computer vision › Video understanding and tracking
video object segmentation |
0.4 | 1 | 2020 | Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement · NeurIPS 2020 |
Network performance modeling › performance prediction
end-to-end performance prediction |
0.2 | 1 | 2023 | AIDTN: Towards a Real-Time AI Optimized DTN System With NVMeoF · IEEE Trans. Parallel Distributed Syst. 2023 |
Methods — techniques the papers use, named apart from their topics
self-attention · 1.5graph transformer · 1.5network feature modeling · 1.3machine learning prediction · 1.3load balancing · 0.7b-tree statistics · 0.7multi-datacenter replication · 0.5materialized view maintenance · 0.5feature bank · 0.4confidence loss · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GTMGC: Using Graph Transformer to Predict Molecule's Ground-State ConformationabstractThe ground-state conformation of a molecule is often decisive for its properties. However, experimental or computational methods, such as density functional theory (DFT), are time-consuming and labor-intensive for obtaining this conformation. Deep learning (DL) based molecular representation learning (MRL) has made significant advancements in molecular modeling and has achieved remarkable results in various tasks. Consequently, it has emerged as a promising approach for directly predicting the ground-state conformation of molecules. In this regard, we introduce GTMGC, a novel network based on Graph-Transformer (GT) that seamlessly predicts the spatial configuration of molecules in a 3D space from their 2D topological architecture in an end-to-end manner. Moreover, we propose a novel self-attention mechanism called Molecule Structural Residual Self-Attention (MSRSA) for molecular structure modeling. This mechanism not only guarantees high model performance and easy implementation but also lends itself well to other molecular modeling tasks. Our method has been evaluated on the Molecule3D benchmark dataset and the QM9 dataset. Experimental results demonstrate that our approach achieves remarkable performance and outperforms current state-of-the-art methods as well as the widely used open-source software RDkit. Guikun Xu, Yongquan Jiang, PengChuan Lei, Jim Chen |
ICLR | 5 |
| 2023 | Progressive Partitioning for Parallelized Query Execution in Google's NapaabstractNapa holds Google's critical data warehouses in log-structured merge trees for real-time data ingestion and sub-second response for billions of queries per day. These queries are often multi-key look-ups in highly skewed tables and indexes. In our production experience, only progressive query-specific partitioning can achieve Napa's strict query latency SLOs. Here we advocate good-enough partitioning that keeps the per-query partitioning time low without risking uneven work distribution. Our design combines pragmatic system choices and algorithmic innovations. For instance, B-trees are augmented with statistics of key distributions, thus serving the dual purpose of aiding lookups and partitioning. Furthermore, progressive partitioning is designed to be "good enough" thereby balancing partitioning time with performance. The resulting system is robust and successfully serves day-in-day-out billions of queries with very high quality of service forming a core infrastructure at Google. Jun'ichi Tatemura, Tao Zou 0002, Jagan Sankaranarayanan, Yanlai Huang, Jim Chen, Hao Zhang 0029, Gokul Nath Babu Manoharan, Goetz Graefe, Divyakant Agrawal, Brad Adelberg, Shilpa Kolhar, Indrajit Roy 0001 |
Proc. VLDB Endow. | 5 |
| 2023 | AIDTN: Towards a Real-Time AI Optimized DTN System With NVMeoFabstractLarge-scale data transport for data-intensive sciences is a complex multidimensional challenge. The challenge includes optimizing the end-to-end Big Data movement performance in real-time, supporting direct remote data access using NVMe over Fabrics (NVMeoF) and deploying to existing research platforms. AIDTN is the first effort to provide a unique AI system designed to incorporate NVMe over Fabrics (NVMeoF) and optimize coordination among multiple components supporting large-scale, multi-domain Wide Area Network (WAN) data-intensive science. AIDTN's research objective is to integrate next-generation storage architecture using NVMeoF, specialized network design using high-performance network appliances, Data Transfer Nodes (DTNs), catalysts in driving data transport, and a unique AI system explicitly designed for high-performance data movement challenges. AIDTN is the first system that uses network and system features to predict the end-to-end performance of high-performance data movement and further extends the model with NVMe-specific features for NVMeoF remote data access. As a result, AIDTN improves data movement performance by up to 284% while minimizing packet loss compared to other heuristics approaches. It also has a prediction error rate as low as 0.16 compared to AI models with the only network (error rate = 0.29) or network and system features (error rate = 0.19). Se-Young Yu, Qingyang Zeng, Jim Chen, Yan Chen 0004, Joe Mambretti |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Napa: Powering Scalable Data Warehousing with Robust Query Performance at GoogleabstractGoogle services continuously generate vast amounts of application data. This data provides valuable insights to business users. We need to store and serve these planet-scale data sets under the extremely demanding requirements of scalability, sub-second query response times, availability, and strong consistency; all this while ingesting a massive stream of updates from applications used around the globe. We have developed and deployed in production an analytical data management system, Napa, to meet these requirements. Napa is the backend for numerous clients in Google. These clients have a strong expectation of variance-free, robust query performance. At its core, Napa's principal technologies for robust query performance include the aggressive use of materialized views, which are maintained consistently as new data is ingested across multiple data centers. Our clients also demand flexibility in being able to adjust their query performance, data freshness, and costs to suit their unique needs. Robust query processing and flexible configuration of client databases are the hallmark of Napa design. Most of the related work in this area takes advantage of full flexibility to design the whole system without the need to support a diverse set of preexisting use cases. In comparison, a particular challenge we faced is that Napa needs to deal with hard constraints from existing applications and infrastructure, so we could not do a "green field" system, but rather had to satisfy existing constraints. These constraints led us to make particular design decisions and also devise new techniques to meet the challenges. In this paper, we share our experiences in designing, implementing, deploying, and running Napa in production with some of Google's most demanding applications. Ankur Agiwal, Gokul Nath Babu Manoharan, Indrajit Roy 0001, Jagan Sankaranarayanan, Hao Zhang 0029, Tao Zou 0002, Jim Chen, Thanh Do, Haoyan Geng, Raman Grover, Yanlai Huang, Adam Li, Jianyi Liang, Xi Mao, Maya Meng, Prashant Mishra, Rajesh Sr, Vijayshankar Raman, Sourashis Roy, Mayank Singh Shishodia, Tianhang Sun, Justin Tang, Jun'ichi Tatemura, Sagar Trehan, Ramkumar Vadali, Prasanna Venkatasubramanian, Joey Zhang, Zeleng Zhuang, Goetz Graefe, Divyakant Agrawal, Jeffrey F. Naughton, Sujata Kosalge, Hakan Hacigümüs |
Proc. VLDB Endow. | 8 |
| 2020 | Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region RefinementabstractThis paper presents a new matching-based framework for semi-supervised video object segmentation (VOS). Recently, state-of-the-art VOS performance has been achieved by matching-based algorithms, in which feature banks are created to store features for region matching and classification. However, how to effectively organize information in the continuously growing feature bank remains under-explored, and this leads to an inefficient design of the bank. We introduced an adaptive feature bank update scheme to dynamically absorb new features and discard obsolete features. We also designed a new confidence loss and a fine-grained segmentation module to enhance the segmentation accuracy in uncertain regions. On public benchmarks, our algorithm outperforms existing state-of-the-arts. Yongqing Liang 0001, Xin Li 0003, Navid H. Jafari, Jim Chen |
NeurIPS | 4 |
| 2000 | Reliability-Availability-Serviceability Characteristics of a Compressed-Memory SystemabstractNew compression innovations and high-density silicon technology enable us to introduce main-memory compression. This technology is able to achieve, in most cases, 2:1 or better compression without impacting performance. It provides an enormous cost/performance advantage, given the cost content of memory in modern enterprise servers. The complex and highly parallel data manipulations central to this compression implementation would, if unprotected by extensive error detection and error correction techniques, offer several potential data integrity exposures. This paper describes the memory subsystem of an enterprise class server with a compressed mainstore and the methods which have been employed to guarantee the integrity of the compressed data. These methods consist of a novel ECC algorithm which includes address information in the code words, the use of CRC codes for compressed data blocks, and various consistency checks on the memory management structures used in the management of a compressed mainstore. Jim Chen, David Har, Ken Mak, Charles O. Schulz, R. Brett Tremaine, Michael E. Wazlowski |
DSN | 1 |