Jeremy Pickens

dblp:71/43 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
3since 2021 · last 2024
0000-0002-4447-035XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 Beyond the Bar: Generative AI as a Transformative Component in Legal Document Review
abstract
Review for responsiveness is a recall-oriented document classification task central to civil litigation. In large legal matters, it may involve the coding of millions of documents by teams of dozens to hundreds of contract attorneys. We describe a prototype document review system based on a large language model (LLM) for replacing the first level of attorney review. Our system accepts the same guidance—a written review protocol—that would be provided to a human review team. We tested our prototype in the context of a live legal matter, evaluating both human review and our LLM-based system against a gold standard coded by expert senior attorneys. Our prototype achieved an estimated 96% recall and 60% precision without matter-specific tuning, and has numerous avenues for further improvement.
Eugene Yang 0001, Roshanak Omrani, Evan Curtin, Tara Emory, Lenora Gray, Jeremy Pickens, Nathan Reff, Cristin Traylor, Sean Underwood, David D. Lewis, Aron J. Ahmadia
IEEE Big Data6
2024 High Recall Retrieval Via Technology-Assisted Review
abstract
High Recall Retrieval (HRR) tasks, including eDiscovery in the law, systematic literature reviews, and sunshine law requests focus on efficiently prioritizing relevant documents for human review.Technology-assisted review (TAR) refers to iterative human-in-the-loop workflows that combine human review with IR and AI techniques to minimize both time and manual effort while maximizing recall. This full-day tutorial provides a comprehensive introduction to TAR. The morning session presents an overview of the key technologies and workflow designs used, the basics of practical evaluation methods, and the social and ethical implications of TAR deployment. The afternoon session provides more technical depth on the implications of TAR workflows for supervised learning algorithm design, how generative AI is can be applied in TAR, more sophisticated statistical evaluation techniques, and a wide range of open research questions.
Lenora Gray, David D. Lewis, Jeremy Pickens, Eugene Yang 0001
SIGIR3
2022 ECIR 2022 Tutorial: Technology-Assisted Review for High Recall Retrieval
Eugene Yang 0001, Jeremy Pickens, David D. Lewis
ECIR (2)2
2018 Break up the Family: Protocols for Efficient Recall-Oriented Retrieval Under Legally-Necessitated Dual Constraints
abstract
In the legal domain, the primary objective of eDiscovery is to retrieve, from within a document population, as many individual relevant documents as possible given the expenditure of a reasonable review effort. The production of relevant documents to opposing counsel, however, is typically made on a family basis (an email and all its attachments), such that the entire family is produced as a collective unit. Counsel's ethical responsibility to prevent the disclosure of privileged information necessitates review of every document in every family being produced. Consequently, it is standard industry practice to batch entire families for review together. This work examines the overall review ordering and efficiency of alternative techniques for managing review leading to the production of relevant document families while fully protecting privilege. Batching, document sequencing, and therefore, training of the core supervised machine learning (AI) algorithm differ among each protocol, and lead to different overall efficiencies. Our empirical results support two conclusions. First, break up the family. In every instance, broken family retrieval protocols are more efficient than full family protocols. Second, a carefully designed and implemented dual phase workflow incorporating an initial, expedited relevance review can be more efficient than a single phase workflow, even if some documents are reviewed twice.
Jeremy Pickens, Tom Gricks, Andrew Bye
IEEE BigData1
2017 Second International Workshop On the Evaluation of Collaborative Information Seeking and Retrieval (Ecol'17)
abstract
The workshop on the evaluation of collaborative information retrieval and seeking (ECol) is held in conjunction with the ACM SIGIR Conference on Human Information Interaction & Retrieval (CHIIR) in Oslo, Norway. To make the workshop active and the participant pro-active, we released datasets and tools so as to help researchers contributing to the formalization of evaluation frameworks for challenging collaborative tasks. The workshop is split into two parts. First, a presentation session. Then, the afternoon is devoted to group discussion addressing challenges of evaluating and designing models for social and collaborative search.
Leif Azzopardi, Jeremy Pickens, Chirag Shah 0001, Laure Soulier, Lynda Tamine-Lechani
CHIIR2
2015 ECol 2015: First international workshop on the Evaluation on Collaborative Information Seeking and Retrieval
abstract
Collaborative Information Seeking/Retrieval (CIS/CIR) has given rise to several challenges in terms of search behavior analysis, retrieval model formalization as well as interface design. However, the major issue of evaluation in CIS/CIR is still underexplored. The goal of this workshop is to investigate the evaluation challenges in CIS/CIR with the hope of building standardized evaluation frameworks, methodologies, and task specifications that would foster and grow the research area (in a collaborative fashion).
Leif Azzopardi, Jeremy Pickens, Tetsuya Sakai, Laure Soulier, Lynda Tamine-Lechani
CIKM2
2013 Assessor disagreement and text classifier accuracy
abstract
Text classifiers are frequently used for high-yield retrieval from large corpora, such as in e-discovery. The classifier is trained by annotating example documents for relevance. These examples may, however, be assessed by people other than those whose conception of relevance is authoritative. In this paper, we examine the impact that disagreement between actual and authoritative assessor has upon classifier effectiveness, when evaluated against the authoritative conception. We find that using alternative assessors leads to a significant decrease in binary classification quality, though less so ranking quality. A ranking consumer would have to go on average 25% deeper in the ranking produced by alternative-assessor training to achieve the same yield as for authoritative-assessor training.
William Webber, Jeremy Pickens
SIGIR2
2011 3rd international workshop on collaborative information retrieval (CIR2011)
abstract
Synchronous, explicit search has some interesting characteristics that distinguish it from other types of interaction: there is much more emphasis on interaction, as the system has to not only communicate search results to the user, but also mediate some forms of communication and data sharing among its users. There are new algorithms that need to be invented that use inputs from multiple people to produce search results, and new evaluation metrics need to be invented that reflect the collaborative and interactive nature of the task. Finally, we need to integrate the expertise of library and information science researchers and practitioners by revisiting real-world information seeking situations with an eye for explicit, synchronous collaborative search.
Gene Golovchinsky, Juan M. Fernández-Luna, Juan F. Huete, Meredith Ringel Morris, Jeremy Pickens, Julio C. Rodríguez Cano
CIKM5
2011 Social and collaborative information seeking: panel
abstract
In recent years, information retrieval and information seeking have moved beyond their single-user roots and are becoming multi-user endeavors. However, there are multiple visions for how best to design multi-user interactions: social search versus collaborative search. The terms "social" and "collaborative" are overloaded with meaning, having been used to describe a wide variety of systems, user needs and goals, interaction styles, and algorithms. In this panel we adopt the following primary definitions: Information seeking tasks in which there are two or more people who lack the same information (share the same information need) and explicitly set out together to satisfy that need are known as collaborative. A collaborative information retrieval system provides mechanisms -- interfaces and mediation algorithms -- that allow the team to work together to find information that neither individual would have found when working alone. There is an inherent division of labor in collaborative work.
Jeremy Pickens
CIKM1
2011 DiG: a task-based approach to product search
abstract
While there are many commercial systems to help people browse and compare products, these interfaces are typically product centric. To help users identify products that match their needs more efficiently, we instead focus on building a task centric interface and system. Based on answers to initial questions about the situations in which they expect to use the product, the interface identifies products that match their needs, and exposes high-level product features related to their tasks, as well as low-level information including customer reviews and product specifications. We developed semi-automatic methods to extract the high-level features used by the system from online product data. These methods identify and group product features, mine and summarize opinions about those features, and identify product uses. User studies verified our focus on high-level features for browsing products and low-level features and specifications for comparing products.
Scott A. Carter, Francine Chen 0001, Aditi S. Muralidharan, Jeremy Pickens
IUI4
2010 Reverted indexing for feedback and expansion
abstract
Traditional interactive information retrieval systems function by creating inverted lists, or term indexes. For every term in the vocabulary, a list is created that contains the documents in which that term occurs and its relative frequency within each document. Retrieval algorithms then use these term frequencies alongside other collection statistics to identify the matching documents for a query. In this paper, we turn the process around: instead of indexing documents, we index query result sets. First, queries are run through a chosen retrieval system. For each query, the resulting document IDs are treated as terms and the score or rank of the document is used as the frequency statistic. An index of documents retrieved by basis queries is created. We call this index a reverted index. With reverted indexes, standard retrieval algorithms can retrieve the matching queries (as results) for a set of documents (used as queries). These recovered queries can then be used to identify additional documents, or to aid the user in query formulation, selection, and feedback.
Jeremy Pickens, Matthew Cooper 0002, Gene Golovchinsky
CIKM1
2010 Introduction to the special issue
Gene Golovchinsky, Meredith Ringel Morris, Jeremy Pickens
Inf. Process. Manag.3
2010 Role-based results redistribution for collaborative information retrieval
Chirag Shah 0001, Jeremy Pickens, Gene Golovchinsky
Inf. Process. Manag.2
2008 Ranked feature fusion models for ad hoc retrieval
abstract
We introduce the Ranked Feature Fusion framework for information retrieval system design. Typical information retrieval formalisms such as the vector space model, the best-match model and the language model first combine features (such as term frequency and document length) into a unified representation, and then use the representation to rank documents. We take the opposite approach: Documents are first ranked by the relevance of a single feature value and are assigned scores based on their relative ordering within the collection. A separate ranked list is created for every feature value and these lists are then fused to produce a final document scoring. This new "rank then combine" approach is extensively evaluated and is shown to be as effective as traditional "combine then rank" approaches. The model is easy to understand and contains fewer parameters than other approaches. Finally, the model is easy to extend (integration of new features is trivial) and modify. This advantage includes but is not limited to relevance feedback and distribution flattening.
Jeremy Pickens, Gene Golovchinsky
CIKM1
2008 Algorithmic mediation for collaborative exploratory search
abstract
We describe a new approach to information retrieval: algorithmic mediation for intentional, synchronous collaborative exploratory search. Using our system, two or more users with a common information need search together, simultaneously. The collaborative system provides tools, user interfaces and, most importantly, algorithmically-mediated retrieval to focus, enhance and augment the team's search and communication activities. Collaborative search outperformed post hoc merging of similarly instrumented single user runs. Algorithmic mediation improved both collaborative search (allowing a team of searchers to find relevant information more efficiently and effectively), and exploratory search (allowing the searchers to find relevant information that cannot be found while working individually).
Jeremy Pickens, Gene Golovchinsky, Chirag Shah 0001, Pernilla Qvarfordt, Maribeth Back
SIGIR1
2006 Term context models for information retrieval
abstract
At their heart, most if not all information retrieval models utilize some form of term frequency.The notion is that the more often a query term occurs in a document, the more likely it is that document meets an information need. We examine an alternative. We propose a model which assesses the presence of a term in a document not by looking at the actual occurrence of that term, but by a set of non-independent supporting terms, i.e. context. This yields a weighting for terms in documents which is different from and complementary to tf-based methods, and is beneficial for retrieval.
Jeremy Pickens, Andrew MacFarlane 0001
CIKM1
2003 Polyphonic music modeling with random fields
abstract
Recent interest in the area of music information retrieval and related technologies is exploding. However, very few of the existing techniques take advantage of recent developments in statistical modeling. In this paper we discuss an application of Random Fields to the problem of creating accurate yet flexible statistical models of polyphonic music. With such models in hand, the challenges of developing effective searching, browsing and organization techniques for the growing bodies of music collections may be successfully met. We offer an evaluation of these models in terms of perplexity and prediction accuracy, and show that random fields not only outperform Markov chains, but are much more robust in terms of overfitting.
Victor Lavrenko, Jeremy Pickens
ACM Multimedia2
2003 Music modeling with random fields
abstract
this report we discuss an application of Random Fields to the problem of statistical modeling of polyphonic music. With such models in hand, the challenges of developing effective searching, browsing, and organization techniques for the growing bodies of music collections may be successfully met
Victor Lavrenko, Jeremy Pickens
SIGIR2
2002 Harmonic models for polyphonic music retrieval
abstract
Most work in the ad hoc music retrieval field has focused on the retrieval of monophonic documents using monophonic queries. Polyphony adds considerably more complexity. We present a method by which polyphonic music documents may be retrieved by polyphonic music queries. A new harmonic description technique is given, wherein the information from all chords, rather than the most significant chord, is used. This description is then combined in a new and unique way with Markov statistical methods to create models of both documents and queries. Document models are compared to query models and then ranked by score. Though test collections for music are currently scarce, we give the first known recall-precision graphs for polyphonic music retrieval, and results are favorable.
Jeremy Pickens, Tim Crawford
CIKM1
2001 Feature Selection for Polyphonic Music Retrieval
abstract
No abstract available.
Jeremy Pickens
SIGIR1