Jens-S. Vöckler

dblp:55/200 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4Computer networks · 2Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 62% Knowledge representation and reasoning · 38%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 70% High-performance computing · 30%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
1.012026
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data · ACL (1) 2026
Data models and query languages
graph query language
1.012026
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.312026
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph querying
0.312026
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data · ACL (1) 2026
Distributed systems
grid computing
0.012004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Distributed systems › grid computing
data grid
0.012002
Applying Chimera virtual data concepts to cluster finding in the Sloan Sky Survey · SC 2002
High-performance computing › scientific workflow
scientific workflow management
0.012002
Applying Chimera virtual data concepts to cluster finding in the Sloan Sky Survey · SC 2002
Distributed systems › workflow management
workflow orchestration
0.012002
Applying Chimera virtual data concepts to cluster finding in the Sloan Sky Survey · SC 2002
High-performance computing
distributed computing infrastructure
0.012004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Computational science and engineering › astronomy
astronomical data analysis
0.012002
Applying Chimera virtual data concepts to cluster finding in the Sloan Sky Survey · SC 2002

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 2.0likelihood-based filtering · 2.0database-driven provenance tracking · 0.1
YearPublicationVenuePosition
2026 CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data
abstract
Knowledge Graph (KG) retrieval is a promising augmentation to address knowledge gaps and hallucinations in LLMs. As KGs in practice are stored in graph databases (e.g., Wikidata, Freebase), accurate retrieval requires translating natural language questions into structured queries (query generation). A key challenge of query generation is Text-to-Cypher, which generates Cypher queries for property graphs (e.g., Neo4j), a paradigm increasingly adopted in industry for their scalable architectures and expressive schemas. However, compared to other query generation tasks such as Text-to-SQL or Text-to-SPARQL, Text-to-Cypher remains underexplored due to scarce public KGs and datasets. Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress. To address this, we introduce CypherSmith, an instruction-tuning dataset over 12\times larger than prior public Text-to-Cypher datasets, spanning diverse domains to better support LLM fine-tuning. Our key distinction lies in fully leveraging open-source LLMs for large-scale synthetic data generation and introducing a novel likelihood-based filtering technique to ensure high-quality Text-to-Cypher data. Extensive experiments demonstrate the effectiveness of CypherSmith, achieving state-of-the-art LLM performance.
Kexuan Sun 0002, Jens-S. Vöckler, Thien Huu Nguyen, Thuy Vu
ACL (1)4
2008 Tracking provenance in a virtual data grid
abstract
Abstract The virtual data model allows data sets to be described prior to, and separately from, their physical materialization. We have implemented this model in a Virtual Data Language (VDL) and associated supporting tools, which provide for both the storage, query, and retrieval of virtual data set descriptions, and the automated, on‐demand materialization of virtual data sets. We use a standardized data provenance challenge exercise to illustrate the powerful queries that can be performed on the data maintained by these tools, which for a single virtual data set can include three elements: the computational procedure(s) that must be executed to materialize the data set, the runtime log(s) produced by the execution of the computation(s), and optional metadata annotation(s) that associate application semantics with data and procedures. Copyright © 2007 John Wiley & Sons, Ltd.
Ben Clifford, Ian T. Foster, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009
Concurr. Comput. Pract. Exp.3
2006 Virtual data Grid middleware services for data-intensive science
abstract
Abstract The GriPhyN virtual data system provides a suite of components and services for data‐intensive sciences that enables scientists to systematically and efficiently describe, discover, and share large‐scale data and computational resources. We describe the design and implementation of such middleware services in terms of a virtual data system interface called Chiron, and present virtual data integration examples from the QuarkNet education project and from functional‐MRI‐based neuroscience research. The Chiron interface also serves as an online ‘educator’ for virtual data applications. Copyright © 2005 John Wiley & Sons, Ltd.
Yong Zhao 0009, Michael Wilde, Ian T. Foster, Jens-S. Vöckler, James E. Dobson, Eric Gilbert, Thomas H. Jordan, Elizabeth Quigg
Concurr. Comput. Pract. Exp.4
2004 The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon
HPDC68
2003 The Virtual Data Grid: A New Model and Architecture for Data-Intensive Collaboration
Ian T. Foster, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009
CIDR2
2002 Applying Chimera virtual data concepts to cluster finding in the Sloan Sky Survey
abstract
In many scientific disciplines — especially long running, data- intensive collaborations — it is important to track all aspects of data capture, production, transformation, and analysis. In principle, one can then audit, validate, reproduce, and/or re-run with corrections various data transformations. We have recently proposed and prototyped the Chimera virtual data system, a new database-driven approach to this problem. We present here a major application study in which we apply Chimera to a challenging data analysis problem: the identification of galaxy clusters within the Sloan Digital Sky Survey. We describe the problem, its computational procedures, and the use of Chimera to plan and orchestrate the workflow of thousands of tasks on a data grid comprising hundreds of computers. This experience suggests that a general set of tools can indeed enhance the accuracy and productivity of scientific data reduction and that further development and application of this paradigm will offer great value.
James Annis, Yong Zhao 0009, Jens-S. Vöckler, Michael Wilde, Steve Kent, Ian T. Foster
SC3
2002 Chimera: AVirtual Data System for Representing, Querying, and Automating Data Derivation
abstract
A lot of scientific data is not obtained from measurements but rather derived from other data by the application of computational procedures. We hypothesize that explicit representation of these procedures can enable documentation of data provenance, discovery of available methods, and on-demand data generation (so-called "virtual data"). To explore this idea, we have developed the Chimera virtual data system, which combines a virtual data catalog for representing data derivation procedures and derived data, with a virtual data language interpreter that translates user requests into data definition and query operations on the database. We couple the Chimera system with distributed "data grid" services to enable on-demand execution of computation schedules constructed from database queries. We have applied this system to two challenge problems, the reconstruction of simulated collision event data from a high-energy physics experiment, and searching digital sky survey data for galactic clusters, with promising results.
Ian T. Foster, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009
SSDBM2
1998 Load and Traffic Balancing in Large Scale Cache Meshes
Christian Grimm, Jens-S. Vöckler, Helmut Pralle
Comput. Networks2
1998 Request Routing in Cache Meshes
Christian Grimm, Jens-S. Vöckler, Helmut Pralle
Comput. Networks2