Edward A. Fox

dblp:f/EdwardAFox · also Edward Alan Fox · DBLP profile ↗
← Back
65ranked-venue papers in the field
17as first author
10since 2021 · last 2024
0000-0003-1447-6870ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 52 (17 first)Big Data, Cloud & Distributed Data Systems · 6Knowledge Engineering, Semantic Web & Information Systems · 4Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2024 Automating Chapter-Level Classification for Electronic Theses and Dissertations
abstract
Traditional archival practices for describing electronic theses and dissertations (ETDs) rely on broad, high-level metadata schemes that fail to capture the depth, complexity, and interdisciplinary nature of these long scholarly works. The lack of detailed, chapter-level content descriptions impedes researchers’ ability to locate specific sections or themes, thereby reducing discoverability and overall accessibility. By providing chapter-level metadata information, we improve the effectiveness of ETDs as research resources. This makes it easier for scholars to navigate them efficiently and extract valuable insights. The absence of such metadata further obstructs interdisciplinary research by obscuring connections across fields, hindering new academic discoveries and collaboration. In this paper, we propose a machine learning and AI-driven solution to automatically categorize ETD chapters. This solution is intended to improve discoverability and promote understanding of chapters. Our approach enriches traditional archival practices by providing context-rich descriptions that facilitate targeted navigation and improved access. We aim to support interdisciplinary research and make ETDs more accessible. By providing chapter-level classification labels and using them to index in our developed prototype system, we make content in ETD chapters more discoverable and usable for a diverse range of scholarly needs. Implementing this AI-enhanced approach allows archives to serve researchers better, enabling efficient access to relevant information and supporting deeper engagement with ETDs. This will increase the impact of ETDs as research tools, foster interdisciplinary exploration, and reinforce the role of archives in scholarly communication within the data-intensive academic landscape.
Bipasha Banerjee, William A. Ingram, Edward A. Fox
IEEE Big Data3
2024 Nuclear Pore Segmentation in 3D FIB-SEM Images with Dynamic Cyclical Data Augmentation
abstract
We developed a deep learning approach to detect and segment nuclear pores in large high-resolution 3D FIB-SEM images. The approach extracts small blocks containing nuclear pore regions from different parts of a cell nucleus and applies data augmentation techniques combined with Random Block Sampling (RBS) to optimize training. By fine tuning the neural network and reducing the batch size to one, we enhanced the model’s ability to learn from individual image features, leading to more precise detection. Our Dynamic Cyclical Data Augmentation (DCDA) adapts in real-time with an automatic termination feature, reducing augmentation time by over 97.4% in our experiments, significantly improving efficiency without compromising model performance. Experimental results consistently demonstrate that our approach can quickly and accurately identify nuclear pores, achieving F1 scores from 0.7119 to 0.8555, with precision ranging from 0.7379 to 0.8516. Training time was also significantly reduced. These substantial improvements in both accuracy and efficiency underscore the method’s effectiveness for large-scale cellular analysis.
Chongyu He, Zhiwu Xie, Yinlin Chen, Edward A. Fox
IEEE Big Data4
2024 Agentic AI for Improving Precision in Identifying Contributions to Sustainable Development Goals
abstract
As research institutions increasingly commit to supporting the United Nations’ Sustainable Development Goals (SDGs), there is a pressing need to accurately assess their research output against these goals. Current approaches, primarily reliant on keyword-based Boolean search queries, conflate incidental keyword matches with genuine contributions, reducing retrieval precision and complicating benchmarking efforts. This study investigates the application of autoregressive Large Language Models (LLMs) as evaluation agents to identify relevant scholarly contributions to SDG targets in scholarly publications. Using a dataset of academic abstracts retrieved via SDG-specific keyword queries, we demonstrate that small, locally-hosted LLMs can differentiate semantically relevant contributions to SDG targets from documents retrieved due to incidental keyword matches, addressing the limitations of traditional methods. By leveraging the contextual understanding of LLMs, this approach provides a scalable framework for improving SDG-related research metrics and informing institutional reporting.
William A. Ingram, Bipasha Banerjee, Edward A. Fox
IEEE Big Data3
2024 Toward Automatically Improving Metadata Quality of Electronic Theses and Dissertations at Scale
abstract
Metadata is crucial for the accessibility, interoperability, and long-term usability of digital objects such as Electronic Theses and Dissertations (ETDs). In large-scale academic repositories, poor metadata quality can significantly impede the discovery and use of resources. This study addresses persistent issues of incomplete and inconsistent ETD metadata collected from U.S. university libraries. However, directly applying machine learning-based error detection and correction models may introduce unwanted errors due to the imperfection of these models. We propose an ETD metadata improvement system (ETDMIS) that mitigates the problem by integrating metadata validation and a version control mechanism. Our system was applied to a dataset of 100,000 U.S. ETDs, resulting in substantial improvements in metadata quality. Scalability was demonstrated by processing the entire dataset efficiently. The original and the enhanced metadata for the 100,000 ETDs are publicly accessible at https://github.com/lamps-lab/ETDMiner/tree/master/Meta100K.
Lamia Salsabil, Jian Wu 0006, William A. Ingram, Edward A. Fox
IEEE Big Data4
2024 Multi-dimensional Edge-Embedded GCNs for Arabic Text Classification
Ola Karajeh, Mohammed Al-Kabi, Edward A. Fox
TPDL (1)3
2023 Multi-view Graph-Based Text Representations for Imbalanced Classification
Ola Karajeh, Ismini Lourentzou, Edward A. Fox
TPDL3
2022 Applications of data analysis on scholarly long documents
abstract
Theses and dissertations record the work of graduate students and are typically a requirement at the culmination of the graduate degree. Thus, they contain important information that reflects a graduate student’s exploration of their research topic. Although print submission was commonplace early on, most universities now require students to submit an electronic version. The electronic document referred to as an ETD henceforth has become the primary way of submitting, storing, and distributing graduate work. Millions of such documents have been created in the past two decades. They are maintained and stored by university libraries, digital repositories, and other academic publishing companies. These online repositories have increased access to such documents. Nonetheless, these documents fail to meet the needs of researchers, who find it challenging to find and access knowledge from such long documents. The worldwide ETD collection has increased in volume to become what is known as ‘scholarly big data’. Apart from the text body, these documents contain a myriad of other pieces of knowledge like tables, figures, definitions, literature reviews, and references. There is a growing demand amongst researchers across various domains to make this collection of scholarly documents more computationally driven. We use ideas from natural language processing, information retrieval, and machine learning to excavate knowledge from this rich information source. In this paper, we examine some of the challenges we face, identify some key areas of exploration, and discuss our methods to mitigate the challenges.
Bipasha Banerjee, William A. Ingram, Jian Wu 0006, Edward A. Fox
IEEE Big Data4
2022 Improving Accessibility to Arabic ETDs Using Automatic Classification
Eman Abdelrahman, Edward A. Fox
TPDL2
2022 Differentially private synthetic medical data generation using convolutional GANs
Amirsina Torfi, Edward A. Fox, Chandan K. Reddy
Inf. Sci.2
2021 Building A Large Collection of Multi-domain Electronic Theses and Dissertations
abstract
In this work, we report our progress on building a collection containing over 450k Electronic Theses and Dissertations (ETDs), including full-text and metadata. Our goal is to close the gap of accessibility between long text and short text documents, and to create a new research opportunity for the scholarly community. For that, we developed an ETD Ingestion Framework (EIF) that automatically harvests metadata and PDFs of ETDs from university libraries. We faced multiple challenges and learned many lessons during the process, that led to proposed solutions to overcome/mitigate the limitations of the current data. We also described the data that we have collected. We hope our methods will be useful for building similar collections from university libraries and that the data can be used for research and education.
Sami Uddin, Bipasha Banerjee, Jian Wu 0006, William A. Ingram, Edward A. Fox
IEEE BigData5
2019 Spatio-Temporal Event Detection from Multiple Data Sources
Aman Ahuja, Ashish Baghudana, Wei Lu 0011, Edward A. Fox, Chandan K. Reddy
PAKDD (1)4
2019 Discovering Product Defects and Solutions from Online User Generated Contents
abstract
The recent increase in online user generated content (UGC) has led to the availability of a large number of posts about products and services. Often, these posts contain complaints that the consumers purchasing the products and services have. However, discovering and summarizing product defects and the related knowledge from large quantities of user posts is a difficult task. Traditional aspect opinion mining models, that aim to discover the product aspects and their corresponding opinions, are not sufficient to discover the product defect information from the user posts. In this paper, we propose the Product Defect Latent Dirichlet Allocation model (PDLDA), a probabilistic model that identifies domain-specific knowledge about product issues using interdependent three-dimensional topics: Component, Symptom, and Resolution. A Gibbs sampling based inference method for PDLDA is also introduced. To evaluate our model, we introduce three novel product review datasets. Both qualitative and quantitative evaluations show that the proposed model results in apparent improvement in the quality of discovered product defect information. Our model has the potential to benefit customers, manufacturers, and policy makers, by automatically discovering product defects from online data.
Xuan Zhang 0005, Zhilei Qiao, Aman Ahuja, Weiguo Fan, Edward A. Fox, Chandan K. Reddy
WWW5
2016 Automated arabic text classification with P-Stemmer, machine learning, and a tailored news article taxonomy
abstract
Arabic news articles in electronic collections are difficult to study. Browsing by category is rarely supported. Although helpful machine‐learning methods have been applied successfully to similar situations for English news articles, limited research has been completed to yield suitable solutions for Arabic news. In connection with a Qatar National Research Fund (QNRF)‐funded project to build digital library community and infrastructure in Qatar, we developed software for browsing a collection of about 237,000 Arabic news articles, which should be applicable to other Arabic news collections. We designed a simple taxonomy for Arabic news stories that is suitable for the needs of Qatar and other nations, is compatible with the subject codes of the International Press Telecommunications Council, and was enhanced with the aid of a librarian expert as well as five Arabic‐speaking volunteers. We developed tailored stemming (i.e., a new Arabic light stemmer called P‐Stemmer) and automatic classification methods (the best being binary Support Vector Machines classifiers) to work with the taxonomy. Using evaluation techniques commonly used in the information retrieval community, including 10‐fold cross‐validation and the Wilcoxon signed‐rank test, we showed that our approach to stemming and classification is superior to state‐of‐the‐art techniques.
Tarek Kanan, Edward A. Fox
J. Assoc. Inf. Sci. Technol.2
2014 Intelligent digital libraries and tailored services
Jonathan Leidig, Edward A. Fox
J. Intell. Inf. Syst.2
2013 Social Navigation Support for Groups in a Community-Based Educational Portal
Peter Brusilovsky, Chirayu Wongchokprasitti, Scott Britell, Lois M. L. Delcambre, Richard Furuta, Kartheek Chiluka, Lillian N. Cassel, Edward A. Fox
TPDL9
2012 Enhancing Digital Libraries and Portals with Canonical Structures for Complex Objects
Scott Britell, Lois M. L. Delcambre, Lillian N. Cassel, Edward A. Fox, Richard Furuta
TPDL4
2011 Digital Library 2.0 for Educational Resources
Monika Akbar, Weiguo Fan, Clifford A. Shaffer, Yinlin Chen, Lillian N. Cassel, Lois M. L. Delcambre, Dan Garcia 0001, Gregory W. Hislop, Frank M. Shipman III, Richard Furuta, B. Stephen Carpenter II, Hao-wei Hsieh, Bob Siegfried, Edward A. Fox
TPDL14
2011 Experiment and Analysis Services in a Fingerprint Digital Library for Collaborative Research
Sung Hee Park, Jonathan Leidig, Lin Tzy Li, Edward A. Fox, Nathan J. Short, Kevin E. Hoyle, A. Lynn Abbott, Michael S. Hsiao
TPDL4
2008 From concepts to implementation and visualization: tools from a team-based approach to ir
abstract
Researchers have been studying and developing teaching materials for information retrieval (IR), such as [3]. Toolkits also have been built that provide hands-on experience to students. For example, IR-Toolbox [4] is an effort to close the gap between the students' understanding of IR concepts and real-life indexing and search systems. Such tools might be good for helping students in non-technical areas such as in the Library and Information Science field to develop their conceptual model of search engines. However, they do not cover emerging topics and skills, such as content-based image retrieval (CBIR) and fusion search. Although there is open source software (such as those in http://www.searchtools.com/tools/tools-opensource.html) that can be used to teach basic and advanced IR topics, they require a student to have high-level technical knowledge and to spend a long time to gain a practical understanding of these topics.
Uma Murthy, Ricardo da Silva Torres, Edward A. Fox, Logambigai Venkatachalam, Seungwon Yang, Marcos André Gonçalves
SIGIR3
2008 Integration of complex archeology digital libraries: An ETANA-DL experience
Rao Shen, Naga Srinivas Vemuri, Weiguo Fan, Edward A. Fox
Inf. Syst.4
2007 "What is a good digital library?" - A quality model for digital libraries
Marcos André Gonçalves, Bárbara Lagoeiro Moreira, Edward A. Fox, Layne T. Watson
Inf. Process. Manag.3
2005 Connecting topics in document collections with stepping stones and pathways
abstract
In this paper, we present Stepping Stones and Pathways (SSP), an alternative model of building and presenting answers for the cases when queries on document collections cannot be answered just by a ranked list. Stepping Stones can handle questions like: What is the relation of topics X and Y? SSP addresses when the contents of a small set of related documents is needed as an answer rather than a single document, or when query splitting is required to satisfactorily explore a document space. Query results are networks of document groups representing topics, each group relating to and connecting (by documents) to other groups in the network. Thus, a network answers the user's information need. We devise new and more effective representations and techniques to visualize such answers, and to involve users as part of the answer-finding process. In order to verify the validity of our approach, and since the questions we aim to answer involve multiple topics, we performed a study involving a custom built broad collection of operating systems research papers, and evaluated the results with interested computer science students, using multiple measures.
Fernando Adrian Das Neves, Edward A. Fox
CIKM2
2005 A new framework to combine descriptors for content-based image retrieval
abstract
In this paper, we propose a novel framework using Genetic Programming to combine image database descriptors for content-based image retrieval (CBIR). Our framework is validated through several experiments involving two image databases and specific domains, where the images are retrieved based on the shape of their objects.
Ricardo da Silva Torres, Alexandre X. Falcão, Baoping Zhang, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Pável Calado
CIKM5
2005 Intelligent GP fusion from multiple sources for text classification
abstract
This paper shows how citation-based information and structural content (e.g., title, abstract) can be combined to improve classification of text documents into predefined categories. We evaluate different measures of similarity -- five derived from the citation information of the collection, and three derived from the structural content -- and determine how they can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our experiments with the ACM Computing Classification Scheme, using documents from the ACM Digital Library, indicate that GP can discover similarity functions superior to those based solely on a single type of evidence. Effectiveness of the similarity functions discovered through simple majority voting is better than that of content-based as well as combination-based Support Vector Machine classifiers. Experiments also were conducted to compare the performance between GP techniques and other fusion techniques such as Genetic Algorithms (GA) and linear fusion. Empirical results show that GP was able to discover better similarity functions than GA or other fusion techniques.
Baoping Zhang, Yuxin Chen 0003, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Marco Cristo, Pável Calado
CIKM4
2005 SimFusion: measuring similarity using unified relationship matrix
abstract
In this paper we use a Unified Relationship Matrix (URM) to represent a set of heterogeneous data objects (e.g., web pages, queries) and their interrelationships (e.g., hyperlinks, user click-through sequences). We claim that iterative computations over the URM can help overcome the data sparseness problem and detect latent relationships among heterogeneous data objects, thus, can improve the quality of information applications that require com- bination of information from heterogeneous sources. To support our claim, we present a unified similarity-calculating algorithm, SimFusion. By iteratively computing over the URM, SimFusion can effectively integrate relationships from heterogeneous sources when measuring the similarity of two data objects. Experiments based on a web search engine query log and a web page collection demonstrate that SimFusion can improve similarity measurement of web objects over both traditional content based algorithms and the cutting edge SimRank algorithm.
Wensi Xi, Edward A. Fox, Weiguo Fan, Benyu Zhang, Zheng Chen 0001, Jun Yan 0001, Dong Zhuang
SIGIR2
2005 Intelligent fusion of structural and citation-based evidence for text classification
abstract
This paper shows how different measures of similarity derived from the citation information and the structural content (e.g., title, abstract) of the collection can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our experiments with the ACM Computing Classification Scheme, using documents from the ACM Digital Library, indicate that GP can discover similarity functions superior to those based solely on a single type of evidence. Effectiveness of the similarity functions discovered through simple majority voting is better than that of content-based as well as combination-based Support Vector Machine classifiers. Experiments also were conducted to compare the performance between GP techniques and other fusion techniques such as Genetic Algorithms (GA) and linear fusion. Empirical results show that GP was able to discover better similarity functions than other fusion techniques.
Baoping Zhang, Yuxin Chen 0003, Weiguo Fan, Edward A. Fox, Marcos André Gonçalves, Marco Cristo, Pável Calado
SIGIR4
2005 An Asian digital libraries perspective
Edward A. Fox, Elisabeth Logan
Inf. Process. Manag.1
2004 MRSSA: an iterative algorithm for similarity spreading over interrelated objects
abstract
We introduce the Multiple Relationship Similarity Spreading Algorithm (MRSSA) to enhance IR effectiveness. This method has similarity computed in an iterative "spreading" fashion for multiple object types, combining both inter- and intra-object relationships. We demonstrate the value of this approach in the context of the WWW, where the key objects are web pages and queries, Relationships considered are derived from hyperlinks (in- and out-links) and click-through logs.
Gui-Rong Xue, Hua-Jun Zeng, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma, Wensi Xi, Edward A. Fox
CIKM7
2004 Combining structural and citation-based evidence for text classification
abstract
This paper discusses how citation-based information and structural content (e.g., title, abstract) can be combined to improve classification of text documents into predefined categories. We evaluate different measures of similarity derived from the citation structure and the structural content of the collection, and determine how they can be fused to improve classification effectiveness. To discover the best fusion framework, we apply Genetic Programming (GP) techniques. Our empirical experiments using documents from the ACM Digital Library and the ACM Computing Classification System show that we can discover similarity functions that work better than using evidence in isolation and whose combined performance through a simple majority voting is comparable to that of Support Vector Machine classifiers.
Baoping Zhang, Marcos André Gonçalves, Weiguo Fan, Yuxin Chen 0003, Edward A. Fox, Pável Calado, Marco Cristo
CIKM5
2004 Tuning before feedback: combining ranking discovery and blind feedback for robust retrieval
abstract
Both ranking functions and user queries are very important factors affecting a search engine's performance. Prior research has looked at how to improve ad-hoc retrieval performance for existing queries while tuning the ranking function, or modify and expand user queries using a fixed ranking scheme using blind feedback. However, almost no research has looked at how to combine ranking function tuning and blind feedback together to improve ad-hoc retrieval performance. In this paper, we look at the performance improvement for ad-hoc retrieval from a more integrated point of view by combining the merits of both techniques. In particular, we argue that the ranking function should be tuned first, using user-provided queries, before applying the blind feedback technique. The intuition is that highly-tuned ranking offers more high quality documents at the top of the hit list, thus offers a stronger baseline for blind feedback. We verify this integrated model in a large scale heterogeneous collection and the experimental results show that combining ranking function tuning and blind feedback can improve search performance by almost 30% over the baseline Okapi system.
Weiguo Fan, Ming Luo 0001, Li Wang 0007, Wensi Xi, Edward A. Fox
SIGIR5
2004 Link fusion: a unified link analysis framework for multi-type interrelated data objects
abstract
Web link analysis has proven to be a significant enhancement for quality based web search. Most existing links can be classified into two categories: intra-type links (e.g., web hyperlinks), which represent the relationship of data objects within a homogeneous data type (web pages), and inter-type links (e.g., user browsing log) which represent the relationship of data objects across different data types (users and web pages). Unfortunately, most link analysis research only considers one type of link. In this paper, we propose a unified link analysis framework, called "link fusion", which considers both the inter- and intra- type link structure among multiple-type inter-related data objects and brings order to objects in each data type at the same time. The PageRank and HITS algorithms are shown to be special cases of our unified link analysis framework. Experiments on an instantiation of the framework that makes use of the user data and web pages extracted from a proxy log show that our proposed algorithm could improve the search effectiveness over the HITS and DirectHit algorithms by 24.6% and 38.2% respectively.
Wensi Xi, Benyu Zhang, Zheng Chen 0001, Yizhou Lu, Shuicheng Yan, Wei-Ying Ma, Edward A. Fox
WWW7
2004 The effects of fitness functions on genetic programming-based ranking discovery forWeb search
abstract
Abstract Genetic‐based evolutionary learning algorithms, such as genetic algorithms (GAs) and genetic programming (GP), have been applied to information retrieval (IR) since the 1980s. Recently, GP has been applied to a new IR task—discovery of ranking functions for Web search—and has achieved very promising results. However, in our prior research, only one fitness function has been used for GP‐based learning. It is unclear how other fitness functions may affect ranking function discovery for Web search, especially since it is well known that choosing a proper fitness function is very important for the effectiveness and efficiency of evolutionary algorithms. In this article, we report our experience in contrasting different fitness function designs on GP‐based learning using a very large Web corpus. Our results indicate that the design of fitness functions is instrumental in performance improvement. We also give recommendations on the design of fitness functions for genetic‐based information retrieval experiments.
Weiguo Fan, Edward A. Fox, Praveen Pathak, Harris Wu
J. Assoc. Inf. Sci. Technol.2
2004 Recommender Systems Research: A Connection-Centric Survey
Saverio Perugini, Marcos André Gonçalves, Edward A. Fox
J. Intell. Inf. Syst.3
2004 Streams, structures, spaces, scenarios, societies (5s): A formal model for digital libraries
abstract
Digital libraries (DLs) are complex information systems and therefore demand formal foundations lest development efforts diverge and interoperability suffers. In this article, we propose the fundamental abstractions of Streams, Structures, Spaces, Scenarios, and Societies (5S), which allow us to define digital libraries rigorously and usefully. Streams are sequences of arbitrary items used to describe both static and dynamic (e.g., video) content. Structures can be viewed as labeled directed graphs, which impose organization. Spaces are sets with operations on those sets that obey certain constraints. Scenarios consist of sequences of events or actions that modify states of a computation in order to accomplish a functional requirement. Societies are sets of entities and activities and the relationships among them. Together these abstractions provide a formal foundation to define, relate, and unify concepts---among others, of digital objects, metadata, collections, and services---required to formalize and elucidate "digital libraries". The applicability, versatility, and unifying power of the 5S model are demonstrated through its use in three distinct applications: building and interpretation of a DL taxonomy, informal and formal analysis of case studies of digital libraries (NDLTD and OAI), and utilization as a formal basis for a DL description language.
Marcos André Gonçalves, Edward A. Fox, Layne T. Watson, Neill A. Kipp
ACM Trans. Inf. Syst.2
2002 Web-DL: an experience in building digital libraries from the web
abstract
The Web contains a huge volume of information, almost all unstructured and, therefore, difficult to manage. In Digital Libraries, however, information is explicitly organized, described, and managed. In this paper, we propose an architecture that allows the construction of digital libraries from the Web, using standard protocols and archival technologies, and incorporating powerful digital library and data extraction tools, thus benefiting from the breadth of the Web contents, but supporting services and organization available in digital libraries. The proposed architecture was applied to the Networked Digital Library of Theses and Dissertations, providing an important first step toward rapid construction of large DLs from the Web, as well as a large-scale solution for interoperability between independent digital libraries.
Pável Calado, Altigran S. da Silva, Berthier A. Ribeiro-Neto, Alberto H. F. Laender, Juliano Palmieri Lage, Davi de Castro Reis, Pablo A. Roberto, Monique V. Vieira, Marcos André Gonçalves, Edward A. Fox
CIKM10
2002 Java MARIAN: From an OPAC to a Modern Digital Library System
Marcos André Gonçalves, Paul Mather, Jun Wang 0125, Ming Luo 0001, Ryan Richardson, Rao Shen, Edward A. Fox
SPIRE9
2002 Machine Learning Approach for Homepage Finding Task
Wensi Xi, Edward A. Fox, Roy Patrick Tan, Jiang Shu
SPIRE2
2001 Building Interoperable Digital Library Services: MARIAN, Open Archives and NDLTD
abstract
In this demonstration, we present interoperable and personalized search services for the Networked Digital Library of Theses and Dissertations (NDLTD). Using standard protocols and software, including those specified by the Open Archives Initiative (OAI), distributed sites can share metadata easily. On top of these harvesting protocols, we implement a union collection of theses managed by the MARIAN digital library system. Our demonstration covers aspects of NDLTD, OAI, and MARIAN.
Edward A. Fox, Robert K. France, Marcos André Gonçalves, Hussein Suleman
SIGIR1
2000 Digital Libraries: Extending and Applying Library and Information Science and Technology
abstract
No abstract available.
Edward A. Fox
CIKM1
1999 Progress Toward Digital Libraries: Augmentation through Integration
Gary Marchionini, Edward A. Fox
Inf. Process. Manag.2
1996 Courseware, Training and Curriculum in Information Retrieval (Workshop Abstract)
abstract
No abstract available.
Edward A. Fox
SIGIR1
1996 Visualizing Search Results: Some Alternatives to Query-Document Similarity
abstract
A digital library of computer science literature, Envision provides powerful information visualization by displaying search results as a matrix of icons, with layout semantics under user control.Envision's Graphic View interacts with an Item Summary Window giving users access to bibliographic information, and XMosaic provides access to complete bibliographic information, abstracts, and full content.While many visualization interfaces for information retrieval systems depict ranked query-document similarity, Envision graphically presents a variety of document characteristics and supports an extensive range of user tasks.Formative usability evaluation results show great user satisfaction with Envision's style of presentation and the document characteristics visualized.
Lucy T. Nowell, Robert K. France, Deborah Hix, Lenwood S. Heath, Edward A. Fox
SIGIR5
1996 Gerald Salton, March 8, 1927 - August 28, 1995
Carolyn J. Crouch, Michael McGill, Michael E. Lesk, Karen Spärck Jones, Edward A. Fox, Donna K. Harman, Donald H. Kraft
J. Am. Soc. Inf. Sci.5
1995 SortTables: A Browser for a Digital Library
abstract
Much research in informationretrieval haafocused more on matching results to queries than on browsing those restdts.After briefly exploring browsing in physical and electronic libraries, we introduce SortTables, a new system that focuses on support for browsing.We explore the evolution of the system in light of early implementation experience and formative evaluation of the interface.Fh-mlly, we briefly review related work, and discuss future directions.
William C. Wake, Edward A. Fox
CIKM2
1995 Combining the Evidence of Multiple Query Representations for Information Retrieval
Nicholas J. Belkin, Paul B. Kantor, Edward A. Fox, Joseph A. Shaw
Inf. Process. Manag.3
1995 Incremental Clustering for Very Large Document Databases: Initial MARIAN Experience
Fazli Can, Edward A. Fox, Cory Snavely, Robert K. France
Inf. Sci.2
1993 An Object-Oriented Database for the Display Measurement and Analysis System
abstract
This paper describes development of an Display Measurement and Analysis System (DMAS).An object data model for DMAS has been specified, and a prototype database subsystem has been implemented using the ObjectStore ODBMS, accessed from a Windows graphical user interface.The DMAS database was designed to be fully object-oriented to support complex data modelling capabilities and easy extension.Research endeavors include exploring ODBMS in this unconventional engineering application domain and development of an object data model and a testbed for display measurement systems.
Yihong Qian, Edward A. Fox, Willard W. Farley
CIKM2
1993 Multiple Access and Retrieval of Information with Annotations (Demo)
abstract
No abstract available.
Edward A. Fox
SIGIR1
1993 Project Envision (Demo)
abstract
No abstract available.
Edward A. Fox
SIGIR1
1993 Development of a Modern OPAC: From REVTOLC to MARIAN
abstract
Since 1986 we have investigated the problems and possibilities of applying modern information retrieval methods to large online public access library catalogs (OPACs). In the Retrieval Experiment—Virginia Tech OnLine Catalog (REVTOLC) study we carried out a large pilot test in 1987 and a larger, controlled investigation in 1990, with 216 users and roughly 500,000 MARC records. Results indicated that a forms-based interface coupled with vector and relevance feedback retrieval methods would be well received. Recent efforts developing the Multiple Access and Retrieval of Information with Annotations (MARIAN) system have involved used of a specially developed object-oriented DBMS, construction of a client running under NeXTSTEP, programming of a distributed server with a thread assigned to each user session to increase concurrency on a small network of NeXTs, refinement of algorithms to use objects and stopping rules for greater efficiency, usability testing and iterative interface refinement.
Edward A. Fox, Robert K. France, Eskinder Sahle, Amjad M. Daoud, Ben E. Cline
SIGIR1
1993 Users, User Interfaces, and Objects: Envision, a Digital Library
abstract
Project Envision aims to build a “user-centered database from the computer science literature,” initially using the publications of the Association for Computing Machinery (ACM). Accordingly, we have interviewed potential users, as well as experts in library, information, and computer science—to understand their needs, to become aware of their perception of existing information systems, and to collect their recommendations. Design and formative usability evaluation of our interface have been based on those interviews, leading to innovative query formulation and search results screens that work well according to our usability testing. Our development of the Envision database, system software, and protocol for client-server communication builds upon work to identify and represent “objects” that will facilitate reuse and high-level communication of information from author to reader (user). All these efforts are leading not only to a usable prototype digital library but also to a set of nine principles for digital libraries, which we have tried to follow, covering issues of representation, architecture, and interfacing. © 1993 John Wiley & Sons, Inc.
Edward A. Fox, Deborah Hix, Lucy T. Nowell, Dennis J. Brueni, William C. Wake, Lenwood S. Heath, Durgesh Rao
J. Am. Soc. Inf. Sci.1
1992 A Faster Algorithm for Constructing Minimal Perfect Hash Functions
abstract
Our previous research on one-probe access to large collections of data indexed by alphanumeric keys has produced the first practical minimal perfect hash functions for this problem. Here, a new algorithm is described for quickly finding minimal perfect hash functions whose specification space is very close to the theoretical lower bound, i.e., around 2 bits per key. The various stages of processing are detailed, along with analytical and empirical results, including timing for a set of over 3.8 million keys that was processed on a NeXTstation in about 6 hours.
Edward A. Fox, Qi Fan Chen, Lenwood S. Heath
SIGIR1
1991 Order-Preserving Minimal Perfect Hash Functions and Information Retrieval
abstract
Rapid access to information is essential for a wide variety of retrieval systems and applications.Hashing has long been used when the fastest possible direct search is desired, but is generally not appropriate when sequential or range searches are also required.This paper describes a hashing method, developed for collections that are relatively static, that supports both direct and sequential access.The algorithms described give hash functions that are optimal in terms of time and hash table space utilization, and that preserve any a priori ordering desired, Furthermore, the resulting order-preserving minimal perfect hash functions (OPMPHFS) can be found using time and space that are linear in the number of keys involve~this is close to optimal,
Edward A. Fox, Qi Fan Chen, Amjad M. Daoud, Lenwood S. Heath
ACM Trans. Inf. Syst.1
1990 Order Preserving Minimal Perfect Hash Functions and Information Retrieval
abstract
Rapid access to information is essential for a wide variety of retrieval systems and applications. Hashing has long been used when the fastest possible direct search is desired, but is generally not appropriate when sequential or range searches are also required. This paper describes a hashing method, developed for collections that are relatively static, that supports both direct and sequential access. Indeed, the algorithm described gives hash functions that are optimal in terms of time and hash table space utilization, and that preserve any a priori ordering desired. Furthermore, the resulting order preserving minimal perfect hash functions (OPMPHFs) can be found using space and time that is on average linear in the number of keys involved.
Edward A. Fox, Qi Fan Chen, Amjad M. Daoud, Lenwood S. Heath
SIGIR1
1990 Using vector and extended Boolean matching in an expert system for selecting foster homes
abstract
Since there are many similarities between the fields of information retrieval and artificial intelligence, it may be appropriate to adapt methods developed in one field to applications usually thought of as belonging to the other. While recent attention in this regard has focused on the use of artificial intelligence methods in the information sciences, this project demonstrates the utility of information retrieval techniques in the expert systems area of artificial intelligence. In particular, it addresses the difficult task of building an expert system for the social sciences. FOCES (FOster Care Expert System) is a prototype assistant to social workers involved in selecting foster homes for children who are in need of placement. The design of FOCES called for use of GUESS (General-pUrpose Expert System Shell) and tailored Prolog routines for extended Boolean matching and vector correlation. Evaluation of the implemented system has shown that FOCES can perform its assigned task as well as trained social workers. © 1990 John Wiley & Sons, Inc.
Edward A. Fox, Sheila G. Winett
J. Am. Soc. Inf. Sci.1
1989 Using a frame-based language for information retrieval
abstract
With the advent of the information society, many researchers are turning to artificial intelligence techniques to provide effective retrieval over large bodies of textual information. In the CODER system, the mission of which is to provide an environment for experiments in applying AI to information retrieval, a factual representation language (FRL) serves as a tool for knowledge engineering and experimentation. the FRL is a hybrid AI language supporting strong typing for attribute values, a frame system, and Prolog-like relational structures. Inheritance is enforced throughout, and the semantics of type subsumption and object matching are formally defined. A collection of type and object managers called the knowledge administration complex implements this common language for storing knowledge and communicating it within the system. Storage of large numbers of complete knowledge objects (statements in the language) is supported by a system of external knowledge bases. Of the three types of knowledge structures in the language, the frame facility has proven most useful in the retrieval domain. This article discusses the frame construct itself, the implementation of the relevant portions of the knowledge administration complex and external knowledge bases, and the use of frames in retrieval research. It closes with a discussion of the utility of the FRL.
Marybeth T. Weaver, Robert K. France, Qi Fan Chen, Edward A. Fox
Int. J. Intell. Syst.4
1989 Research and development of information retrieval models and their application
Edward A. Fox
Inf. Process. Manag.1
1988 Coefficients for Combining Concept Classes in a Collection
abstract
This report considers combining information to improve retrieval. The vector space model has been extended so different classes of data are associated with distinct concept types and their respective subvectors. Two collections with multiple concept types are described, ISI-1460 and CACM-3204. Experiments indicate that regression methods can help predict relevance, given query-document similarity values for each concept type. After sampling and transformation of data, the coefficient of determination for the best model was .48 (.66) for ISI (CACM). Average precision for the two collections was 11% (31%) better for probabilistic feedback with all types versus with terms only. These findings may be of particular interest to designers of document retrieval or hypertext systems since the role of links is shown to be especially beneficial.
Edward A. Fox, Gary L. Nunn, Whay C. Lee
SIGIR1
1988 Practical enhanced Boolean retrieval: Experiences with the smart and sire systems
Edward A. Fox, Matthew B. Koll
Inf. Process. Manag.1
1987 Distributed Expert-Based Information Systems: An Interdisciplinary Approach
Nicholas J. Belkin, Christine L. Borgman, Helen M. Brooks, Tom Bylander, W. Bruce Croft, Penny J. Daniels, Scott C. Deerwester, Edward A. Fox, Peter Ingwersen, Roy Rada
Inf. Process. Manag.8
1987 Development of the coder system: A testbed for artificial intelligence methods in information retrieval
Edward A. Fox
Inf. Process. Manag.1
1985 Composite Document Extended Retrieval - An Overview
abstract
Experimental information retrieval (IR) systems, some dating back to the sixties, have demonstrated the viability of fully automatic document storage and retrieval methodologies with small to medium size bibliographic collections [72]. Many of these experimental systems utilize the vector space model in which each important term (such as a word stem) identifies a different dimension in a space, so that matrix methods and vector operations can be defined on queries and documents. Statistical techniques have been very effective, and probabilistic enhancements have given additional improvements [84]. However, the basic vector space model is oriented towards recording the essential information in the text of a title/abstract combination rather than describing more complex document structures. It is necessary to extend the model in order to handle composite documents.
Edward A. Fox
SIGIR1
1985 Advanced feedback methods in information retrieval
abstract
Abstract Automatic feedback methods may be used in online information retrieval to generate improved query statements based on information contained in previously retrieved documents. In this study automatic relevance feedback techniques are applied to Boolean query statements. The feedback operations are carried out using both the conventional Boolean logic, as well as an extended logic producing improved retrieval effectiveness. Experimental output is included to evaluate the automatic feedback operations.
Gerard Salton, Edward A. Fox, Ellen M. Voorhees
J. Am. Soc. Inf. Sci.2
1984 A comparison of two methods for boolean query relevancy feedback
Gerard Salton, Ellen M. Voorhees, Edward A. Fox
Inf. Process. Manag.3
1983 Automatic query formulations in information retrieval
abstract
Modern information retrieval systems are designed to supply relevant information in response to requests received from the user population. In most retrieval environments the search requests consist of keywords, or index terms, interrelated by appropriate Boolean operators. Since it is difficult for untrained users to generate effective Boolean search requests, trained search intermediaries are normally used to translate original statements of user need into useful Boolean search formulations. Methods are introduced in this study which reduce the role of the search intermediaries by making it possible to generate Boolean search formulations completely automatically from natural language statements provided by the system patrons. Frequency considerations are used automatically to generate appropriate term combinations as well as Boolean connectives relating the terms. Methods are covered to produce automatic query formulations both in a standard Boolean logic system, as well as in an extended Boolean system in which the strict interpretation of the connectives is relaxed. Experimental results are supplied to evaluate the effectiveness of the automatic query formulation process, and methods are described for applying the automatic query formulation process in practice.
Gerard Salton, Chris Buckley, Edward A. Fox
J. Am. Soc. Inf. Sci.3