EDBT 2026 Demo / reviewers in the wild / expert
Amit Singhal 0001
dblp:s/AmitSinghal
· DBLP profile ↗
28ranked-venue papers
11as first author
0since 2021 · last 2008
0000-0002-4010-6614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 7 first-authorArtificial intelligence and machine learning · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorComputer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Information retrieval · 92% Data mining · 8% | |
| Artificial intelligence
2 papers |
Segmentation and scene understanding · 57% Information extraction and text analysis · 29% Probabilistic and Bayesian machine learning · 14% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
web search |
0.1 | 2 | 2005 | Challenges in running a commercial search engine · SIGIR 2005 A case study in web search using TREC algorithms · WWW 2001 |
Natural language and speech › Information extraction and text analysis › event extraction
event classification |
0.1 | 1 | 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classification · CVPR 2008 |
Computer vision › Segmentation and scene understanding › image segmentation
saliency-based segmentation |
0.1 | 1 | 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classification · CVPR 2008 |
Information retrieval
adversarial retrieval |
0.1 | 1 | 2005 | Challenges in running a commercial search engine · SIGIR 2005 |
Information retrieval › search engines
commercial search engines |
0.1 | 1 | 2005 | Challenges in running a commercial search engine · SIGIR 2005 |
Information retrieval › web search
search engine spam |
0.1 | 1 | 2005 | Challenges in running a commercial search engine · SIGIR 2005 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.0 | 1 | 2003 | Probabilistic Spatial Context Models for Scene Content Understanding · CVPR (1) 2003 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.0 | 1 | 2003 | Probabilistic Spatial Context Models for Scene Content Understanding · CVPR (1) 2003 |
Computer vision › Segmentation and scene understanding › context modeling
spatial context modeling |
0.0 | 1 | 2003 | Probabilistic Spatial Context Models for Scene Content Understanding · CVPR (1) 2003 |
Information retrieval › ranking › text ranking
document ranking |
0.0 | 1 | 2001 | A case study in web search using TREC algorithms · WWW 2001 |
Data mining › text mining › information extraction
entity extraction |
0.0 | 1 | 2001 | A case study in web search using TREC algorithms · WWW 2001 |
Information retrieval › ranking
keyword ranking |
0.0 | 1 | 2001 | A case study in web search using TREC algorithms · WWW 2001 |
Visualization and visual analytics
visual saliency |
0.0 | 1 | 2000 | On Measuring Low-Level Saliency in Photographic Images · CVPR 2000 |
Information retrieval
evaluation |
0.0 | 2 | 2005 | Challenges in running a commercial search engine · SIGIR 2005 A case study in web search using TREC algorithms · WWW 2001 |
Information retrieval › document processing › document analysis › document representation
document expansion |
0.0 | 1 | 1999 | Document Expansion for Speech Retrieval · SIGIR 1999 |
Information retrieval › search interfaces
search interface design |
0.0 | 1 | 1999 | SCAN: Designing and Evaluating User Interfaces to Support Retrieval From Speech Archives · SIGIR 1999 |
Information retrieval › document retrieval
spoken document retrieval |
0.0 | 1 | 1999 | Document Expansion for Speech Retrieval · SIGIR 1999 |
Data mining › text mining
text classification |
0.0 | 1 | 1999 | ATTICS: A Software Platform for Online Text Classification (poster abstract) · SIGIR 1999 |
Information retrieval › query reformulation › query expansion
automatic query expansion |
0.0 | 1 | 1998 | Improving Automatic Query Expansion · SIGIR 1998 |
Information retrieval › relevance feedback
pseudo-relevance feedback |
0.0 | 1 | 1998 | Improving Automatic Query Expansion · SIGIR 1998 |
Information retrieval › query reformulation
query expansion |
0.0 | 1 | 1998 | Improving Automatic Query Expansion · SIGIR 1998 |
Information retrieval › information filtering
text filtering |
0.0 | 1 | 1998 | Boosting and Rocchio Applied to Text Filtering · SIGIR 1998 |
Information retrieval › information filtering
document routing |
0.0 | 1 | 1997 | Learning Routing Queries in a Query Zone · SIGIR 1997 |
Information retrieval › retrieval models › term weighting
document length normalization |
0.0 | 1 | 1996 | Pivoted Document Length Normalization · SIGIR 1996 |
Information retrieval
ranking |
0.0 | 1 | 1996 | Pivoted Document Length Normalization · SIGIR 1996 |
Information retrieval › indexing
document indexing |
0.0 | 1 | 1999 | Document Expansion for Speech Retrieval · SIGIR 1999 |
Information retrieval › ranking › search ranking
relevance ranking |
0.0 | 1 | 1999 | SCAN: Designing and Evaluating User Interfaces to Support Retrieval From Speech Archives · SIGIR 1999 |
Information retrieval
speech recognition |
0.0 | 1 | 1999 | Document Expansion for Speech Retrieval · SIGIR 1999 |
Information retrieval › query reformulation
query drift |
0.0 | 1 | 1998 | Improving Automatic Query Expansion · SIGIR 1998 |
Information retrieval › evaluation
retrieval effectiveness |
0.0 | 1 | 1998 | Improving Automatic Query Expansion · SIGIR 1998 |
Methods — techniques the papers use, named apart from their topics
structured prediction · 0.1hidden conditional random field · 0.1probabilistic spatial context modeling · 0.0link-based ranking · 0.0keyword-based ranking · 0.0spatial context modeling · 0.0user study · 0.0term cooccurrence · 0.0rocchio · 0.0proximity constraint · 0.0boosting · 0.0boolean filter · 0.0adaboost · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | Selective hidden random fields: Exploiting domain-specific saliency for event classificationabstractClassifying an event captured in an image is useful for understanding the contents of the image. The captured event provides context to refine models for the presence and appearance of various entities, such as people and objects, in the captured scene. Such contextual processing facilitates the generation of better abstractions and annotations for the image. Consider a typical set of consumer images with sports-related content. These images are taken mostly by amateur photographers, and often at a distance. In the absence of manual annotation or other sources of information such as time and location, typical recognition tasks are formidable on these images. Identifying the sporting event in these images provides a context for further recognition and annotation tasks. We propose to use the domain-specific saliency of the appearances of the playing surfaces, and ignore the noninformative parts of the image such as crowd regions, to discriminate among different sports. To this end, we present a variation of the hidden-state conditional random field that selects a subset of the observed features suitable for classification. The inferred hidden variables in this model represent a selection criteria desirable for the problem domain. For sports-related images, this selection criteria corresponds to the segmentation of the playing surface in the image. We demonstrate the utility of this model on consumer images collected from the Internet. Vidit Jain, Amit Singhal 0001, Jiebo Luo 0001 |
CVPR | 2 |
| 2008 | Web Search: Challenges and Directions
Amit Singhal 0001 |
ECIR | 1 |
| 2005 | Challenges in running a commercial search engineabstractThese are exciting times for Information Retrieval. Web search engines have brought IR to the masses. It now affects the lives of hundreds of millions of people, and growing, as Internet search companies launch ever more products based on techniques developed in These are exciting times for Information Retrieval. Web search engines have brought IR to the masses. It now affects the lives of hundreds of millions of people, and growing, as Internet search companies launch ever more products based on techniques developed in IR research.The real world poses unique challenges for search algorithms. They operate at unprecedented scales, and over a wide diversity of information. In addition, we have entered an unprecedented world of "Adversarial Information Retrieval". The lure of billions of dollars of commerce, guided by search engines, motivates all kinds of people to try all kinds of tricks to get their sites to the top of the search results.What techniques do people use to defeat IR algorithms? What are the evaluation challenges for a web search engine? How much impact has IR had on search engines? How does Google serve over 250 Million queries a day, often with sub-second response times? This talk will show that the world of algorithm and system design for commercial search engines can be described by two of Murphy's Laws: a) If anything can go wrong, it will, and b) If anything cannot go wrong, it will anyway. Amit Singhal 0001 |
SIGIR | 1 |
| 2005 | A Bayesian network-based framework for semantic image understanding
Jiebo Luo 0001, Andreas E. Savakis, Amit Singhal 0001 |
Pattern Recognit. | 3 |
| 2004 | A computational approach to determination of main subject regions in photographic images
Jiebo Luo 0001, Amit Singhal 0001, Stephen P. Etz, Robert T. Gray |
Image Vis. Comput. | 2 |
| 2003 | Probabilistic Spatial Context Models for Scene Content UnderstandingabstractScene content understanding facilitates a large number of applications, ranging from content-based image retrieval to other multimedia applications. Material detection refers to the problem of identifying key semantic material types (such as sky, grass, foliage, water, and snow in images). In this paper, we present a holistic approach to determining scene content, based on a set of individual material detection algorithms, as well as probabilistic spatial context models. A major limitation of individual material detectors is the significant number of misclassifications that occur because of the similarities in color and texture characteristics of various material types. We have developed a spatial context-aware material detection system that reduces misclassification by constraining the beliefs to conform to the probabilistic spatial context models. Experimental results show that the accuracy of materials detection is improved by 13% using the spatial context models over the individual material detectors themselves. Amit Singhal 0001, Jiebo Luo 0001, Weiyu Zhu |
CVPR (1) | 1 |
| 2003 | Natural object detection in outdoor scenes based on probabilistic spatial context modelsabstractNatural object detection in outdoor scenes, i.e., identifying key object types such as sky, grass, foliage, water, and snow, can facilitate content-based applications, ranging from image enhancement to other multimedia applications. A major limitation of individual object detectors is the significant number of misclassifications that occur because of the similarities in color and texture characteristics of various object types and lack of context information. We have developed a spatial context-aware object-detection system that first combines the output of individual object detectors to produce a composite belief vector for the objects potentially present in an image. Spatial context constraints, in the form of probability density functions obtained by learning, are subsequently used to reduce misclassification by constraining the beliefs to conform to the spatial context models. Experimental results show that the spatial context models improve the accuracy of natural object detection by 13% over the individual object detectors themselves. Jiebo Luo 0001, Amit Singhal 0001, Weiyu Zhu |
ICME | 2 |
| 2002 | Displaying images on mobile devices: capabilities, issues, and solutionsabstractWireless imaging is enabling visual communication "anytime anywhere" to become a reality. A key technical challenge is how to achieve best perceived image quality given the limited screen size and display bit depth of the mobile devices. We give an overview of the current capabilities of various mobile devices, highlight some of the technical issues, and present potential solutions. In addition, we present a review of some of the software products on the market and look ahead to the trend towards more capable devices. Amit Singhal 0001, Jiebo Luo 0001, Gustav J. Braun, Robert T. Gray, Nicolas Touchard |
ICIP (1) | 1 |
| 2002 | Terabit switching: a survey of techniques and current products
Amit Singhal 0001, Raj Jain |
Comput. Commun. | 1 |
| 2002 | Displaying images on mobile devices: capabilities, issues, and solutionsabstractAbstract Wireless imaging is enabling visual communication ‘anytime anywhere’ to become a reality. Apart from wireless communication issues, a key technical challenge is how to achieve the best‐perceived image quality given the limited screen size and display bit‐depth of the mobile devices. In this paper, we give an overview of the current capabilities of various mobile devices, highlight some of the technical issues, and present potential solutions. In addition, we present a review of some of the software products on the market and look ahead to the trend toward more capable devices. Copyright © 2002 John Wiley & Sons, Ltd. Jiebo Luo 0001, Amit Singhal 0001, Gustav J. Braun, Robert T. Gray, Nicolas Touchard, Olivier Seignol |
Wirel. Commun. Mob. Comput. | 2 |
| 2001 | Efficient Multicast Algorithms for Heterogeneous Switch-based Irregular Networks of WorkstationsabstractThis paper considers the problem of efficient multicast on worm-hole routed irregular heterogeneous networks of workstations, using multiple unicast messages. The fundamental issues are the avoidance of link contention and effective use of the faster nodes in the system for distributing the multicast message. Previously proposed schemes have either considered optimization for heterogeneity or elimination of contention, but not both together. We present two algorithms that addresses both issues and demonstrate their superiority through simulation studies. Amit Singhal 0001, Mohammad Banikazemi, P. Sadayappan, Dhabaleswar K. Panda 0001 |
IPDPS | 1 |
| 2001 | A case study in web search using TREC algorithmsabstractWeb search engines rank potentially relevant pages/sites for a user query. Ranking documents for user queries has also been at the heart of the Text REtrieval Conference (TREC in short) under the label ###### retrieval. The TREC community has developed document ranking algorithms that are known to be the best for searching the document collections used in TREC, which are mainly comprised of newswire text. However, the web search community has developed its own methods to rank web pages/sites, many of which use link structure on the web, and are quite dierentfrom the algorithms developed at TREC. This study evaluates the performance of a state-of-the-art keyword-based document ranking algorithm (coming out of TREC) on a popular web search task: nding the web page/site of an entity, #### companies, universities, organizations, individuals, etc. This form of querying is quite prevalentonthe web. The results from the TREC algorithms are compared to four commercial web search engines. Results show that for nding the web page/site of an entity, commercial web search engines are notably better than a state-of-the-art TREC algorithm. These results are in sharp contrast to results from several previous studies. Keywords Search engines, TREC ad-hoc, keyword-based ranking, linkbased ranking 1. Amit Singhal 0001, Marcin Kaszkiel |
WWW | 1 |
| 2001 | On measuring low-level self and relative saliency in photographic images
Jiebo Luo 0001, Amit Singhal 0001 |
Pattern Recognit. Lett. | 2 |
| 2000 | Boosting for Document RoutingabstractRankBoost is a recently proposed algorithm for learning ranking functions. It is simple to implement and has strong justifications from computational learning theory. We describe the algorithm and present experimental results on applying it to the document routing problem. The first set of results applies RankBoost to a text representation produced using modern term weighting methods. Performance of RankBoost is somewhat inferior to that of a state-of-the-art routing algorithm which is, however, more complex and less theoretically justified than RankBoost. RankBoost achieves comparable performance to the state-of-the-art algorithm when combined with feature or example selection heuristics. Our second set of results examines the behavior of RankBoost when it has to learn not only a ranking function but also all aspects of term weighting from raw data. Performance is usually, though not always, less good here, but the term weighting functions implicit in the resulting ranking functions are intriguing, and the approach could easily be adapted to mixtures of textual and nontextual data. Raj D. Iyer, David D. Lewis, Robert E. Schapire, Yoram Singer, Amit Singhal 0001 |
CIKM | 5 |
| 2000 | On Measuring Low-Level Saliency in Photographic ImagesabstractMeasuring perceptual saliency of regions in a scene is important for determining regions of interest. Color, texture and shape cues are good low-level features for detecting saliency. While self saliency refers to intrinsic attributes of a region, relative saliency as used to measure how salient a region is relative to its surrounding and, thus, needs to be defined within a spatial context. A few spatial context models are investigated in this study. In particular, we propose an auto-scaled, amorphous neighborhood as the context model to obtain reliable measurements of relative saliency features. Comparison of three context models has shown that the proposed model is capable of generating predicates more consistent with perceived saliency. Jiebo Luo 0001, Amit Singhal 0001 |
CVPR | 2 |
| 2000 | Quantitative Evaluation of Rank-Order Similarity of ImagesabstractRegion importance maps from image understanding algorithms and human observer studies are ordered rankings of the pixel locations. Kemeny and Snell's distance (d/sub KS/), an existing measure from ordinal ranking theory, can thus be used as a similarity measure between images. We address three problems with d/sub KS/: its high computational cost, its bias in favor of images with sparse histograms, and its image-size dependent range of values. We present a novel computationally efficient algorithm for computing d/sub KS/ between two images, and we derive a normalized form d/sub KS/ with no bias whose range is independent of image size. For evaluating an algorithm where the reference data and algorithm output are ordered rankings of pixels, d/sub KS/ is subjectively superior to the correlation coefficient as a figure of merit. Stephen P. Etz, Jiebo Luo 0001, Robert T. Gray, Amit Singhal 0001 |
ICIP | 4 |
| 2000 | On the Application of Bayes Networks to Semantic Understanding of Consumer PhotographsabstractBelief networks, such as Bayes nets, have emerged as an effective knowledge representation and inference engine in artificial intelligence and expert systems research. Their effectiveness is due to the ability to explicitly integrate domain knowledge in the network structure and to reduce a joint probability distribution to conditionally independence relationships. Current research in content-based image processing and analysis is largely limited to low-level feature extraction and classification. The ability to extract both low-level and semantic features and perform knowledge integration of different types of features would be very useful. We present a general knowledge integration framework that incorporates Bayes networks and has been used in two applications involving semantic understanding of consumer photographs. The first application aims at detecting main photographic subjects in an image and the second aims at selecting the most appealing image in an event. With these diverse examples, we demonstrate that effective inference engines can be built according to specific domain knowledge and available training data to solve inherently uncertain vision problems. Jiebo Luo 0001, Andreas E. Savakis, Stephen P. Etz, Amit Singhal 0001 |
ICIP | 4 |
| 1999 | ATTICS: A Software Platform for Online Text Classification (poster abstract)abstractNo abstract available. David D. Lewis, Daniel L. Stern, Amit Singhal 0001 |
SIGIR | 3 |
| 1999 | Document Expansion for Speech RetrievalabstractAdvances in automatic speech recognition allow us to search large speech collections using traditional information retrieval methods.The problem of \aboutness" for documents | is a document about a certain concept | has been at the core of document indexing for the entire history of IR.This problem is more dicult for speech indexing since automatic speech transcriptions often contain mistakes.In this study we s h o w t h a t document expansion can be successfully used to alleviate the eect of transcription mistakes on speech r etrieval.The loss of retrieval eectiveness due to automatic transcription errors can be reduced by document expansion from 15{27% relative t o r e t r i e v al from human transcriptions to only about 7{13%, even for automatic transcriptions with word error rates as high as 65%.For good automatic transcriptions (25% word error rate), retrieval eectiveness with document expansion is indistinguishable from retrieval from human transcriptions.This makes speech retrieval from automatic transcriptions, even poor ones, competitive with retrieval from perfect transcriptions. Amit Singhal 0001, Fernando Pereira 0003 |
SIGIR | 1 |
| 1999 | SCAN: Designing and Evaluating User Interfaces to Support Retrieval From Speech ArchivesabstractPrevious examinations of search in textual archives have assumed that users first retrieve a ranked set of documents relevant to their query, and then visually scan through these documents, to identify the information they seek.While document scanning is possible in text, it is much more laborious in speech archives, due to the inherently serial nature of speech.Yet, in developing tools for speech access, little attention has so far been paid to users' problems in scanning and extracting information from within "speech documents".We demonstrate the extent of these problems in two user studies.We show that users experience severe problems with local navigation in extracting relevant information from within "speech documents".Based on these results, we propose a new user interface (UI) design paradigm: What You See Is (Almost) What You Hear, (WYSIAWYH) -a multimodal method for accessing speech archives.This paradigm presents a visual analogue to the underlying speech, enabling visual scanning for effective local navigation.We empirically evaluate a UI based on this paradigm.We compare our WYSIAWYH UI with a visual "tape recorder", in relevance ranking, fact-finding, and summarization tasks involving broadcast news data.Our findings indicate that an interface supporting local navigation multimodally helps relevance ranking and fact-finding, but not summarization.We analyze the reasons for system success and identify outstanding research issues in UI design for speech archives. Steve Whittaker 0001, Julia Hirschberg, Donald Hindle, Fernando Pereira 0003, Amit Singhal 0001 |
SIGIR | 6 |
| 1998 | SCAN - speech content based audio navigator: a system overview
Donald Hindle, Julia Hirschberg, Ivan Magrin-Chagnolleau, Christine H. Nakatani, Fernando Pereira 0003, Amit Singhal 0001, Steve Whittaker 0001 |
ICSLP | 7 |
| 1998 | Improving Automatic Query ExpansionabstractMost casual users of IR systems type short queries. Recent research has shown that adding new words to these queries via blind feedback, without any input from the user, improves the performance of such queries. We investigate ways to improve this query expansion process by refining the set of documents used in feedback. We start by using manually formulated Boolean filters along with proximity constraints. Our approach is similar to the one proposed in [10]. Next, we investigate a completely automatic method that makes use of term cooccurrence information to estimate word correlation. Results show that refining the set of documents used in query expansion yields substantial improvements in retrieval effectiveness, both in terms of average precision and precision at top twenty documents. Such refinement often prevents the query drift caused by blind expansion. More importantly, the fully automatic approach developed in this study performs competitively with the best manual approach and... Mandar Mitra, Amit Singhal 0001, Chris Buckley |
SIGIR | 2 |
| 1998 | Boosting and Rocchio Applied to Text FilteringabstractWe discuss two learning algorithms for text filtering: modified Rocchio and a boosting algorithm called AdaBoost. We show how both algorithms can be adapted to maximize any general utility matrix that associates cost (or gain) for each pair of machine prediction and correct label. We first show that AdaBoost significantly outperforms another highly effective text filtering algorithm. We then compare AdaBoost and Rocchio over three large text filtering tasks. Overall both algorithms are comparable and are quite effective. AdaBoost produces better classifiers than Rocchio when the training collection contains a very large number of relevant documents. However, on these tasks, Rocchio runs much faster than AdaBoost. 1 Introduction With the explosion in the amount of information available electronically, information filtering systems that automatically send articles of potential interest to a user are becoming increasingly important. If users indicate their interests to a filtering system... Robert E. Schapire, Yoram Singer, Amit Singhal 0001 |
SIGIR | 3 |
| 1997 | Learning Routing Queries in a Query ZoneabstractWord usage is domain dependent. A common word in one domain can be quite infrequent in another. In this study we exploit this property of word usage to improve document routing. We show that routing queries (profiles) learned only from the documents in a query domain are better than the routing profiles learned when query domains are not used. We approximate a query domain by a query zone. Experiments show that routing profiles learned from a query zone are 8--12% more effective than the profiles generated when no query zoning is used. 1 Background Document routing is an important problem in the field of information retrieval. [12] When a user has marked several articles as relevant to his/her information need, a system should be able to automatically learn the user's "profile" and should be able to route (send) new, potentially interesting, articles to the user. This problem has also been called as selective dissemination of information or information filtering. [4] Most current st... Amit Singhal 0001, Mandar Mitra, Chris Buckley |
SIGIR | 1 |
| 1997 | Automatic Text Structuring and SummarizationabstractIn recent years, information retrieval techniques have been used for automatic generation of semantic hypertext links. This study applies the ideas from the automatic link generation research to attack another important problem in text processing—automatic text summarization. An automatic “general purpose” text summarization tool would be of immense utility in this age of information overload. Using the techniques used (by most automatic hypertext link generation algorithms) for inter-document link generation, we generate intra-document links between passages of a document. Based on the intra-document linkage pattern of a text, we characterize the structure of the text. We apply the knowledge of text structure to do automatic text summarization by passage extraction. We evaluate a set of fifty summaries generated using our techniques by comparing them to paragraph extracts constructed by humans. The automatic summarization methods perform well, especially in view of the fact that the summaries generated by two humans for the same article are surprisingly dissimilar. Gerard Salton, Amit Singhal 0001, Mandar Mitra, Chris Buckley |
Inf. Process. Manag. | 2 |
| 1996 | Pivoted Document Length NormalizationabstractArticle Free Access Share on Pivoted document length normalization Authors: Amit Singhal Department of Computer Science, Cornell University, Ithaca, NY Department of Computer Science, Cornell University, Ithaca, NYView Profile , Chris Buckley Department of Computer Science, Cornell University, Ithaca, NY Department of Computer Science, Cornell University, Ithaca, NYView Profile , Mandar Mitra Department of Computer Science, Cornell University, Ithaca, NY Department of Computer Science, Cornell University, Ithaca, NYView Profile Authors Info & Claims SIGIR '96: Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrievalAugust 1996 Pages 21–29https://doi.org/10.1145/243199.243206Online:18 August 1996Publication History 520citation2,173DownloadsMetricsTotal Citations520Total Downloads2,173Last 12 Months21Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Amit Singhal 0001, Chris Buckley, Mandar Mitra |
SIGIR | 1 |
| 1996 | Automatic Text Decomposition and StructuringabstractSophisticated text similarity measurements are used to determine relationships between natural-language texts and text excerpts. The resulting linked hypertext maps can be decomposed into text segments and text themes, and these decompositions are usable to identify different text types and text structures, leading to improved text access and utilization. Examples of text decomposition are given for expository and non-expository texts. Gerard Salton, James Allan 0001, Amit Singhal 0001 |
Inf. Process. Manag. | 3 |
| 1996 | Document Length NormalizationabstractIn the TREC collection—a large full-text experimental text collection with widely varying document lengths—we observe that the likelihood of a document being judged relevant by a user increases with the document length. We show that a retrieval strategy, such as the vector-space cosine match, that retrieves documents of different lengths with roughly equal chances, will not optimally retrieve useful documents from such a collection. We present a modified technique—pivoted cosine normalization—that attempts to match the likelihood of retrieving documents of all lengths to the likelihood of their relevance, and show that this technique yields significant improvements in retrieval effectiveness. Amit Singhal 0001, Gerard Salton, Mandar Mitra, Chris Buckley |
Inf. Process. Manag. | 1 |