Chris Clark

dblp:36/235 · DBLP profile ↗
← Back
2ranked-venue papers in the field
1as first author
2since 2021 · last 2023
0000-0001-9982-7849ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (1 first)
YearPublicationVenuePosition
2023 Exploring the Performance Impacts of Training Predictive Models with Inclusive Email Threads
abstract
Email threading is a commonly used tool by legal practitioners to streamline document review and classification in legal proceedings. Threading organizes component pieces of an email thread together to effectively reduce a dataset. The most inclusive threads and their associated document attachments are maintained, and non-inclusive or duplicative thread components are set aside. However, as data volumes continue to grow and outpace deadlines for legal proceedings, practitioners often look to incorporate multiple cost-effective and defensible technology solutions to further reduce or otherwise accelerate document review and classification. One such technology is predictive modeling – known in the legal industry as ‘predictive coding’ or ‘Technology Assisted Review (TAR)’ – which is a popular tool used to augment a manual document review and classification process. Like email threading, predictive coding has become more commonplace recently for its proven ability to minimize manual document classification, thus reducing the time and cost associated with this aspect of legal proceedings.In this study, we explore the performance impacts of layering predictive modeling onto an email threading reduction workflow. Generally, a predictive model is established first, and email threading is layered onto the scored output to further reduce and streamline document review. Our research evaluates a reversed workflow, where a population is initially reduced to its inclusive email threads, and this limited dataset is used to train and apply a predictive model to the larger population of documents for classification. Using classified data from four real-world legal proceedings, we compare the performance impact of email threading on predictive modeling by 1) training a model using all positive and negative examples, 2) training a model using positive and negative examples from only inclusive email threads, and 3) training a model using positive examples from only inclusive email threads, but negative examples from all emails. The results of our research provide thoughtful, empirical insights for legal practitioners to review when exploring the deployment of both email threading and predictive modeling into a single, cohesive document classification strategy for their legal proceedings.
Chris Clark, Han Qin, Nathaniel Huber-Fliflet, Adam Dabrowski, Jianping Zhang 0003
IEEE Big Data1
2023 Exploring Approaches to Optimize the Performance of Predictive Coding on Multilanguage Data Sets
abstract
Predictive modeling - known in the legal industry as ‘predictive coding’ or ‘Technology Assisted Review (TAR)’ - is a popular tool used by legal professionals to augment a historically manual document review and classification process in responding to data requests for various legal proceedings. It has become more commonplace over the past decade due to its proven ability to minimize manual document classification, thus reducing the time and cost associated with this aspect of legal proceedings. There is significant research supporting the effectiveness of this technology that includes topics, such as identifying the most performant machine learning algorithms (e.g., logistic regression) or establishing the best methodologies for selecting representative training and testing data. Primarily, this research has been performed without a focus on multilanguage data sets because many legal proceedings typically involve one primary, dominant language. As acceptance of predictive modeling technology grows, legal practitioners have developed independent preferences for handling multiple languages in data sets. These preferences have been primarily based on anecdotal experience rather than empirical assessments and have resulted in two modeling approaches. The first approach uses a single model, which is less complex, and more cost effective than the next approach. The second develops multiple language-specific models, which can create workflow complexity that may impact the overall cost savings that predictive coding seeks to achieve. Proponents of the second approach believe that creating models per language results in better performance over a single model approach. In this study, we empirically explore the performance differences between the two approaches - single model and multiple language-specific models. We hypothesize that a single model approach performs similarly to a language-specific modeling approach when classifying documents for relevance. We evaluate both approaches by comparing their resulting precision-recall curves across real-world test data from legal document review projects. Excitingly, our results demonstrate that in most scenarios, a single model (trained with multiple languages) approach will perform as well, or sometimes better than the approach that uses a group of language-specific models. The results of our research provide a roadmap for legal practitioners and other parties to thoughtfully engage in dialogue around the most effective way to deploy predictive coding within a multilanguage data set. This research will change how predictive models are deployed for multilanguage data sets across the legal industry and enable significant cost savings for clients and legal practitioners by utilizing the most efficient modeling method.
Christian J. Mahoney, Nathaniel Huber-Fliflet, Peter Gronvall, Chris Clark, Jianping Zhang 0003, Fusheng Wei, Qiang Mao
IEEE Big Data4