Thomas Gschwind

dblp:83/2798 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0003-0212-4800ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 42% Question answering and dialogue systems · 21% Knowledge representation and reasoning · 21%
Software engineering, system software, and programming languages
3 papers
Services computing and microservices · 84% Requirements engineering and software design · 10% Software maintenance and evolution · 6%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
concept-based explanation
0.612022
Attention-based Interpretability with Concept Transformers · ICLR 2022
Machine learning › Trustworthy machine learning
interpretability
0.612022
Attention-based Interpretability with Concept Transformers · ICLR 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.612022
A Goal-Driven Natural Language Interface for Creating Application Integration Workflows · AAAI 2022
Natural language and speech › Question answering and dialogue systems
natural language interface
0.612022
A Goal-Driven Natural Language Interface for Creating Application Integration Workflows · AAAI 2022
Services computing and microservices › enterprise application integration
application integration
0.612022
A Goal-Driven Natural Language Interface for Creating Application Integration Workflows · AAAI 2022
Services computing and microservices › service composition
workflow composition
0.612022
A Goal-Driven Natural Language Interface for Creating Application Integration Workflows · AAAI 2022
Natural language and speech › Language models and text generation › code generation
natural language to code
0.312026
From Natural Language to Executable ETL Flows: The IBM DataStage Assistant · AAAI 2026
Machine learning › Deep learning architectures and training
attention mechanism
0.212022
Attention-based Interpretability with Concept Transformers · ICLR 2022
Software maintenance and evolution
program comprehension
0.112008
Extracting Interactions in Component-Based Systems · IEEE Trans. Software Eng. 2008
Requirements engineering and software design
software architecture
0.122008
The Vienna Component Framework Enabling Composition Across Component Models · ICSE 2003
Extracting Interactions in Component-Based Systems · IEEE Trans. Software Eng. 2008
Requirements engineering and software design › software architecture
component-based software engineering
0.012003
The Vienna Component Framework Enabling Composition Across Component Models · ICSE 2003
Requirements engineering and software design › software architecture › component-based software engineering
component-based systems
0.012008
Extracting Interactions in Component-Based Systems · IEEE Trans. Software Eng. 2008
Requirements engineering and software design › software architecture › component-based software engineering
component composition
0.012003
The Vienna Component Framework Enabling Composition Across Component Models · ICSE 2003
Distributed systems
replication
0.011999
NewsCache - A High-Performance Cache Implementation for Usenet News · USENIX ATC, General Track 1999

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0large language model · 2.0knowledge graph · 1.1abstract meaning representation · 1.1AI planning · 1.1concept transformer · 0.6finite state machine · 0.1dynamic analysis · 0.1
YearPublicationVenuePosition
2026 From Natural Language to Executable ETL Flows: The IBM DataStage Assistant
abstract
Modern ETL (Extract, Transform, Load) tools offer graphical, no-code interfaces for workflow creation but still require users to manually identify transformation functions and configure their properties, which is time-consuming and demands prior expertise. We present the research and engineering foundations of the IBM DataStage Assistant, a deployed capability that generates complete multi-stage ETL flows directly from natural language (NL) descriptions. Our framework infers transformation functions, their properties, and transformer expressions, enabling novices to discover relevant functions and allowing experts to bypass manual configuration. The proposed framework achieves a prediction accuracy of 96.4% for flow predictions, 87.0% for properties, and 83.6% for transformer expressions. We also show a document exploration module that uses retrieval-augmented generation (RAG) over product documentation to answer tool-specific questions in NL. Implemented in IBM DataStage, this approach supports iterative, in-environment workflow design and reduces context switching. In initial studies, it achieves up to 90% time savings for novices and 50% for experts.
Nitin Gupta 0005, Thomas Gschwind, Shramona Chakraborty, Sameep Mehta, Tristan Tyler, Shreya Sisodia, Ben Clermont
AAAI2
2023 Follow the Successful Herd: Towards Explanations for Improved Use and Mental Models of Natural Language Systems
abstract
While natural language systems continue improving, they are still imperfect. If a user has a better understanding of how a system works, they may be able to better accomplish their goals even in imperfect systems. We explored whether explanations can support effective authoring of natural language utterances and how those explanations impact users’ mental models in the context of a natural language system that generates small programs. Through an online study (n=252), we compared two main types of explanations: 1) system-focused, which provide information about how the system processes utterances and matches terms to a knowledge base, and 2) social, which provide information about how other users have successfully interacted with the system. Our results indicate that providing social suggestions of terms to add to an utterance helped users to repair and generate correct flows more than system-focused explanations or social recommendations of words to modify. We also found that participants commonly understood some mechanisms of the natural language system, such as the matching of terms to a knowledge base, but they often lacked other critical knowledge, such as how the system handled structuring and ordering. Based on these findings, we make design recommendations for supporting interactions with and understanding of natural language systems.
Michelle Brachman, Hyo Jin Do, Casey Dugan, Arunima Chaudhary, James M. Johnson, Priyanshu Rai, Tathagata Chakraborti, Thomas Gschwind, Jim Laredo, Christoph Miksovic, Paolo Scotton, Kartik Talamadupula, Gegi Thomas
IUI9
2022 A Goal-Driven Natural Language Interface for Creating Application Integration Workflows
abstract
Web applications and services are increasingly important in a distributed internet filled with diverse cloud services and applications, each of which enable the completion of narrowly defined tasks. Given the explosion in the scale and diversity of such services, their composition and integration for achieving complex user goals remains a challenging task for end-users and requires a lot of development effort when specified by hand. We present a demonstration of the Goal Oriented Flow Assistant (GOFA) system, which provides a natural language solution to generate workflows for application integration. Our tool is built on a three-step pipeline: it first uses Abstract Meaning Representation (AMR) to parse utterances; it then uses a knowledge graph to validate candidates; and finally uses an AI planner to compose the candidate flow. We provide a video demonstration of the deployed system as part of our submission.
Michelle Brachman, Christopher Bygrave, Tathagata Chakraborti, Arunima Chaudhary, Zhining Ding, Casey Dugan, Thomas Gschwind, James M. Johnson, Jim Laredo, Christoph Miksovic, Priyanshu Rai, Ramkumar Ramalingam, Paolo Scotton, Nagarjuna Surabathina, Kartik Talamadupula
AAAI8
2022 Attention-based Interpretability with Concept Transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, Paolo Scotton
ICLR4
2019 Fast Record Linkage for Company Entities
abstract
Record linkage is an essential part of nearly all real-world systems that consume structured and unstructured data coming from different sources. Typically no common key is available for connecting records. Massive data integration processes often have to be completed before any data analytics and further processing can be performed. In this work we focus on company entity matching, where company name, location and industry are taken into account. Our contribution is a highly scalable, enterprise-grade end-to-end system that uses rule-based linkage algorithms in combination with a machine learning approach to account for short company names. Linkage time is greatly reduced by an efficient decomposition of the search space using MinHash. Based on real-world ground truth datasets, we show that our approach reaches a recall of 91% compared to 73% for baseline approaches, while scaling linearly with the number of nodes used in the system.
Thomas Gschwind, Christoph Miksovic, Julian Minder, Katsiaryna Mirylenka, Paolo Scotton
IEEE BigData1
2018 Accurate Resampling of Top-k Monitoring Reports Using Unreported Residual Resource Consumption
Thomas Gschwind, Metin Feridun
IEEE Trans. Netw. Serv. Manag.1
2017 Using residual resource consumption to resample top-k monitoring reports
abstract
Top-k reports are compound metrics that provide useful information when diagnosing problems in a system, e.g., to identify persistent CPU usage by a process. In large systems, these reports are collected at regular intervals and need to be resampled to a coarser granularity to answer user queries for different sampling periods, or to save space and make it possible to keep historical data for long term performance analysis. However, resampling top-k reports, i.e., aggregating several reports collected for small time intervals into a single top-k report can introduce inaccuracies. For example, a process that consistently uses CPU over the aggregation interval but did not make it to the short term top-k reports will be missing from the aggregated report. In this paper, we present an algorithm that collects top-k reports at regular intervals and can aggregate them with little or no error. This is done by including residual resource consumption of unreported, but potentially significant entities in the top-k reports, and using these residual values during aggregation. We show different approaches to including residual resource consumption in individual top-k reports, analyze the error introduced, and demonstrate the effectiveness of the algorithm in real-world scenarios.
Thomas Gschwind, Metin Feridun
IM1
2009 Incremental Development of Model Transformation Chains Using Automated Testing
Jochen Malte Küster, Thomas Gschwind, Olaf Zimmermann
MoDELS2
2008 Applying Patterns during Business Process Modeling
Thomas Gschwind, Jana Koehler, Janette Wong
BPM1
2008 Combining Pattern Languages and Reusable Architectural Decision Models into a Comprehensive and Comprehensible Design Method
abstract
When constructing software systems, software architects must identify and evaluate many competing design options and document the rationale behind any selections made. Two supporting concepts are pattern languages and architectural decision models. Unfortunately, both concepts only provide partial support: Extensive upfront education is needed for practitioners to be in command of' the full pattern literature relevant in their field; retrospective architectural decision modeling is viewed as a painful extra responsibility without immediate gains. In this paper, we combine pattern languages and reusable architectural decision models into a design method that is both comprehensive and comprehensible. Our design method identifies the required decisions in requirements models systematically, gives domain-specific pattern selection advice, and provides traceability from platform-independent patterns to platform-specific decisions. We validate our approach by applying it to enterprise applications as an exemplary application genre and a SOA case study from the finance industry.
Olaf Zimmermann, Uwe Zdun, Thomas Gschwind, Frank Leymann
WICSA3
2008 Extracting Interactions in Component-Based Systems
abstract
Monitoring, analysing and understanding component based enterprise software systems are challenging tasks. These tasks are essential in solving and preventing performance and quality problems. Obtaining component level interactions which show the relationships between different software entities is a necessary prerequisite for such efforts. This paper focuses on component based Java applications, currently widely used by industry. They pose specific challenges while raising interesting opportunities for component level interaction extraction tools. We present a range of representative approaches for dynamically obtaining and using component interactions. For each approach we detail the needs it addresses, and the technical requirements for building an implementation of the approach. We also take a critical look at the different available implementations of the various techniques presented. We give performance and functional considerations and contrast them against each other by outlining their relative advantages and disadvantages. Based on this data, developers and system integrators can better understand the current state-of-the-art and the implications of choosing or implementing different dynamic interaction extraction techniques.
Trevor Parsons, Adrian Mos, Mircea Trofin, Thomas Gschwind, John Murphy 0001
IEEE Trans. Software Eng.4
2003 The Vienna Component Framework Enabling Composition Across Component Models
abstract
The Vienna Component Framework (VCF) supports the interoperability and composability of components across different component models, a facility that is lacking in existing component models. The VCF presents a unified component model-implemented by a facade component to the application programmer The programmer may write new components by composing components from different component models, accessed through the VCF The model supports common component features, namely, methods, properties, and events. To support a component model within the VCF a plugin component is needed that provides access to the component model. The paper presents the VCF's design, implementation issues, and evaluation. Performance measurements of VCF implementations of COM, Enterprise JavaBeans, CORBA distributed objects, and JavaBeans show that the overhead of accessing components through the VCF is negligible for distributed components.
Johann Oberleitner, Thomas Gschwind, Mehdi Jazayeri
ICSE2
1999 NewsCache - A High-Performance Cache Implementation for Usenet News
Thomas Gschwind, Manfred Hauswirth
USENIX ATC, General Track1