VLDB 2026 Research / reviewers in the wild / expert
Frank Schilder
dblp:21/2419
· DBLP profile ↗
21ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0001-8227-5099ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Unleashing the Power of Large Language Models for Legal ApplicationsabstractThe use of Large Language Models (LLMs) is revolutionizing the legal industry. In this technical talk, we would like to explore the various use cases of LLMs in legal tasks, discuss the best practices, investigate the available resources, examine the ethical concerns, and suggest promising research directions. Dell Zhang, Alina Petrova, Dietrich Trautmann, Frank Schilder |
CIKM | 4 |
| 2023 | Making a Computational AttorneyabstractThis “blue sky idea” paper outlines the opportunities and challenges in data mining and machine learning involving making a computational attorney — an intelligent software agent capable of helping human lawyers with a wide range of complex high-level legal tasks such as drafting legal briefs for the prosecution or defense in court. In particular, we discuss what a ChatGPT-like Large Legal Language Model (L3M) can and cannot do today, which will inspire researchers with promising short-term and long-term research objectives. Dell Zhang, Frank Schilder, Jack G. Conrad, Masoud Makrehchi, David von Rickenbach, Isabelle Moulinier |
SDM | 2 |
| 2022 | Multi-label legal document classification: A deep learning-based approach with label-attention and domain-specific pre-training
Dezhao Song, Andrew Vold, Kanika Madan, Frank Schilder |
Inf. Syst. | 4 |
| 2019 | Westlaw Edge AI Features Demo: KeyCite Overruling Risk, Litigation Analytics, and WestSearch PlusabstractWestlaw Edge, a new legal research platform from Westlaw, was launched in 2018. The three AI-enabled features that launched with Westlaw Edge are KeyCite Overruling Risk, Litigation Analytics, and WestSearch Plus. Keycite Overruling Risk uses NLP and machine learning to warn users when a point of law in a case may have been implicitly undermined based on a prior decision, when that prior citation has no direct citation relationship to the at-risk case. Litigation Analytics allows users to access valuable metadata extracted from legal dockets about the legal actions carried out by parties, lawyers, law firms, and judges presiding over cases. WestSearch Plus is a non-factoid Question Answering system that provides legally correct, jurisdictionally relevant, and conversationally responsive answers to user-entered questions in the legal domain. Added AI-enabled features are set to launch on the Westlaw Edge platform in 2019. It is available for demo at ICAIL 2019. Tonya Custis, Frank Schilder, Thomas Vacek, Gayle McElvain, Héctor Martínez Alonso |
ICAIL | 2 |
| 2019 | Building and Querying an Enterprise Knowledge GraphabstractInformation providers are faced with a critical challenge to process, retrieve and present information to their users in order to satisfy their complex information needs, because data has been increasing in an unprecedented manner, coming from diverse sources, and covering a variety of domains in heterogeneous formats. In this paper, we present Thomson Reuters' effort in developing a family of services for building and querying an enterprise knowledge graph in order to address this challenge. We first acquire data from various sources via different approaches. Furthermore, we mine useful information from the data by adopting a variety of techniques, including Named Entity Recognition and Relation Extraction; such mined information is further integrated with existing structured data (e.g., via Entity Linking techniques) in order to obtain relatively comprehensive descriptions of the entities. By modeling the data as an RDF graph model, we enable easy data management and the embedding of rich semantics in our data. Finally, in order to facilitate the querying of this mined and integrated data, i.e., the knowledge graph, we propose TR Discover, a natural language interface that allows users to ask questions of our knowledge graph in their own words; such natural language questions are then translated into executable queries for answer retrieval. We evaluate our services, i.e., named entity recognition, relation extraction, entity linking and natural language interface, on real-world datasets, and demonstrate and discuss their practicability and limitations. Dezhao Song, Frank Schilder, Shai Hertz, Giuseppe Saltini, Charese Smiley, Phani Nivarthi, Oren Hazai, Dudi Landau, Mike Zaharkin, Tom Zielund, Hugo Molina-Salgado, Chris Brew, Dan Bennett |
IEEE Trans. Serv. Comput. | 2 |
| 2018 | The E2E NLG Challenge: A Tale of Two SystemsabstractThis paper presents the two systems we entered into the 2017 E2E NLG Challenge: TemplGen, a templated-based system and SeqGen, a neural network-based system.Through the automatic evaluation, SeqGen achieved competitive results compared to the template-based approach and to other participating systems as well.In addition to the automatic evaluation, in this paper we present and discuss the human evaluation results of our two systems. Charese Smiley, Elnaz Davoodi, Dezhao Song, Frank Schilder |
INLG | 4 |
| 2017 | A sequence approach to case outcome detectionabstractWe describe a system to detect the outcome of U.S. Federal District Court cases based on PACER electronic dockets. We study the text processing components of the system and develop two model architectures in order to detect the outcome of a case per party (e.g., dismissed by Court or Verdict for Plaintiff). We conclude that modeling cases as a linear-chain graphical model (i.e., Conditional Random Field (CRF)) offers significantly better performance than modeling the case entry-by-entry (i.e., Logistic Regression (LR)). We in particular show that a first-order modeling of the CRF significantly outperforms the factorized model for the CRF architecture. Tom Vacek, Frank Schilder |
ICAIL | 2 |
| 2017 | Finding the "right" answers for customersabstractThis talk will present a few NLG systems developed within Thomson Reuters providing information to professionals such as lawyers, accountants or traders. Based on the experience developing these system, I will discuss the usefulness of automatic metrics, crowd-sourced evaluation, corpora studies and expert reviews. I will conclude with exploring the question of whether developers of NLG systems need to follow ethical guidelines and how those guidelines could be established. Frank Schilder |
INLG | 1 |
| 2017 | A Multidimensional Investigation of the Effects of Publication Retraction on Scholarly ImpactabstractDuring the past few decades, the rate of publication retractions has increased dramatically in academia. In this study, we investigate retractions from a quantitative perspective, aiming to answer two fundamental questions. One, how do retractions influence the scholarly impact of retracted papers, authors, and institutions? Two, does this influence propagate to the wider academic community through scholarly associations? Specifically, we analyzed a set of retracted articles indexed in Thomson Reuters Web of Science (WoS), and ran multiple experiments to compare changes in scholarly impact against a control set of nonretracted articles, authors, and institutions. We further applied the Granger Causality test to investigate whether different scientific topics are dynamically affected by retracted papers occurring within those topics. Our results show two key findings: first, the scholarly impact of retracted papers and authors significantly decreases after retraction, and the most severe impact decrease correlates with retractions based on proven, purposeful scientific misconduct; second, this retraction penalty does not seem to spread through the broader scholarly social graph, but instead has a limited and localized effect. Our findings may provide useful insights for scholars or science committees to evaluate the scholarly value of papers, authors, or institutions related to retractions. Xin Shuai, Jason Rollins, Isabelle Moulinier, Tonya Custis, Mathilda Edmunds, Frank Schilder |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2016 | When to Plummet and When to Soar: Corpus Based Verb Selection for Natural Language GenerationabstractFor data-to-text tasks in Natural Language Generation (NLG), researchers are often faced with choices about the right words to express phenomena seen in the data.One common phenomenon centers around the description of trends between two data points and selecting the appropriate verb to express both the direction and intensity of movement.Our research shows that rather than simply selecting the same verbs again and again, variation and naturalness can be achieved by quantifying writers' patterns of usage around verbs. Charese Smiley, Vassilis Plachouras, Frank Schilder, Hiroko Bretz, Jochen L. Leidner, Dezhao Song |
INLG | 3 |
| 2016 | Interacting with Financial Data using Natural LanguageabstractFinancial and economic data are typically available in the form of tables and comprise mostly of monetary amounts, numeric and other domain-specific fields. They can be very hard to search and they are often made available out of context, or in forms which cannot be integrated with systems where text is required, such as voice-enabled devices. This work presents a novel system that enables both experts in the finance domain and non-expert users to search financial data with both keyword and natural language queries. Our system answers the queries with an automatically generated textual description using Natural Language Generation (NLG). The answers are further enriched with derived information, not explicitly asked in the user query, to provide the context of the answer. The system is designed to be flexible in order to accommodate new use cases without significant development effort, thus allowing fast integration of new datasets. Vassilis Plachouras, Charese Smiley, Hiroko Bretz, Ola Taylor, Jochen L. Leidner, Dezhao Song, Frank Schilder |
SIGIR | 7 |
| 2015 | Natural Language Question Answering and Analytics for Diverse and Interlinked DatasetsabstractPrevious systems for natural language questions over complex linked datasets require the user to enter a complete and well-formed question, and present the answers as raw lists of entities. Using a feature-based grammar with a full formal semantics, we have developed a system that is able to support rich autosuggest, and to deliver dynamically generated analytics for each result that it returns. Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew |
HLT-NAACL | 2 |
| 2015 | TR Discover: A Natural Language Interface for Querying and Analyzing Interlinked Datasets
Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew, Tom Zielund, Hiroko Bretz, Robert Martin, Chris Dale, John Duprey, Johanna Harrison |
ISWC (2) | 2 |
| 2013 | A Statistical NLG Framework for Aggregated Planning and Realization
Ravikumar Kondadadi, Blake Howald, Frank Schilder |
ACL (1) | 3 |
| 2010 | Representation and Management of Narrative Information: Theoretical Principles and Implementation - Gian Piero Zarri, Springer Verlag, 2009, x+301 pp; ISBN 978-1-84800-077-3abstractZarri's book summarizes more than a decade of his research on knowledge representation for narrative text.The centerpiece of Zarri's work is the Narrative Knowledge Representation Language (NKRL), which he describes and compares to other competing theories.In addition, he discusses how to model the meaning of narrative text by giving many real-world examples.NKRL provides three different components or capabilities: (a) a representation system, (b) inferencing, and (c) an implementation.It is implemented via a Java-based system that shows how a representational theory can be applied to narrative texts.The book consists of five chapters and two appendices.Chapter 1 introduces the basic principles of NKRL.The chapter first defines the focus on nonfiction narratives by contrasting the domain with fictional narratives, for example, a novel.Zarri chooses n-ary predicates in order to represent events formally.He argues for a neo-Davidsonian knowledge representation following Schank (1980), Schubert (1976), and others, and at the same time he sets his approach apart from the knowledge representation proposals one can find in Semantic Web representation languages such as RDF and OWL.However, Zarri emphasizes that NKRL, despite its similarity to conceptual graphs (Sowa 1999), is more focused on practical applications.The chapter concludes by introducing so-called templates in an attempt to demonstrate the practical usefulness of NKRL.Chapter 2 provides an in-depth description of NKRL.Four connected components are introduced: r The definitional component provides a hierarchy of abstract concepts (e.g., artifact, company, activity) called HClass (hierarchy of classes).r The descriptive component is a hierarchy of event types called HTemp (hierarchy of templates) commonly found in the domain of non-fiction narratives (e.g., moving an object, producing a task or activity).r The factual component describes the concrete instantiation of an event.For example, the sentence Berlex Laboratories have performed an evaluation of a given compound would be represented as [ Frank Schilder |
Comput. Linguistics | 1 |
| 2009 | Query-based opinion summarization for legal blog entriesabstractWe present the first report of automatic sentiment summarization in the legal domain. This work is based on processing a set of legal questions with a system consisting of a semi-automatic Web blog search module and FastSum, a fully automatic extractive multi-document sentiment summarization system. We provide quantitative evaluation results of the summaries using legal expert reviewers. We report baseline evaluation results for query-based sentiment summarization for legal blogs: on a five-point scale, average responsiveness and linguistic quality are slightly higher than 2 (with human inter-rater agreement at k = 0.75). To the best of our knowledge, this is the first evaluation of sentiment summarization in the legal blogosphere. Jack G. Conrad, Jochen L. Leidner, Frank Schilder, Ravikumar Kondadadi |
ICAIL | 3 |
| 2007 | Opinion mining in legal blogsabstractWe perform a survey into the scope and utility of opinion mining in legal Weblogs (a.k.a. blawgs). The number of ‘blogs ’ in the legal domain is growing at a rapid pace and many potential applications for opinion detection and monitoring are arising as a result. We summarize current approaches to opinion mining before describing different categories of blawgs and their potential impact on the law and the legal profession. In addition to educating the community on recent developments in the legal blog space, we also conduct some introductory opinion mining trials. We first construct a Weblog test collection containing blog entries that discuss legal search tools. We subsequently examine the performance of a language modeling approach deployed for both subjectivity analysis (i.e., is the text subjective or objective?) and polarity analysis (i.e., is the text affirmative or negative towards its subject?). This work may thus help establish early baselines for these core opinion mining tasks. Jack G. Conrad, Frank Schilder |
ICAIL | 2 |
| 2004 | Extracting meaning from temporal nouns and temporal prepositionsabstractThis article provides a compositional semantics for temporal nouns and temporal prepositions that are annotated as temporal prepositional phrases or noun phrases by an automatic tagging system (e.g., last Monday, on Dec. 1 st , for three weeks or before Christmas ). Current temporal tagging systems rely on an ad-hoc-representation for temporal date and time expressions, but the more demanding tasks of temporal question-answering and automatic text summarization require a sound logical derivation and representation of temporal expressions. Our proposal draws from two formal accounts of temporal prepositional phrases by Pratt and Francez [2001] and von Stechow [2002b], and is realized within an automatic temporal tagging system for German newspaper articles. Frank Schilder |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2002 | Robust discourse parsing via discourse markers, topicality and positionabstractThis paper describes a simple discourse parsing and analysis algorithm that combines a formal underspecification utilising discourse grammar with Information Retrieval (IR) techniques. First, linguistic knowledge based on discourse markers is used to constrain a totally underspecified discourse representation. Then, the remaining underspecification is further specified by the computation of a topicality score for every discourse unit. This computation is done via the vector space model. Finally, the sentences in a prominent position (e.g. the first sentence of a paragraph) are given an adjusted topicality score. The proposed algorithm was evaluated by applying it to a text summarisation task. Results from a psycholinguistic experiment, indicating the most salient sentences for a given text as the ‘gold standard’, show that the algorithm performs better than commonly used machine learning and statistical approaches to summarisation. Frank Schilder |
Nat. Lang. Eng. | 1 |
| 1999 | Pointing to Events
Frank Schilder |
EACL | 1 |
| 1995 | Aspect and Discourse Structure: Is a Neutral Viewpoint Required?abstractWe apply Smith's theory of aspect (1991) to German -a language without any aspectual markers.In particular, we try to shed more light on the effects aspect can have on discourse structure and show how English and German behave differently in this respect.We furthermore Frank Schilder |
ACL | 1 |