Vittorio Castelli

dblp:c/VittorioCastelli · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorTheory of computation · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 CiteEval: Principle-Driven Citation Evaluation for Source Attribution
abstract
Yumo Xu, Peng Qi, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu, Bonan Min, Vittorio Castelli, Arshit Gupta, Zhiguo Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yumo Xu, Peng Qi 0003, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu 0004, Bonan Min, Vittorio Castelli, Arshit Gupta, Zhiguo Wang 0006
ACL (1)8
2024 RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
abstract
Rujun Han, Yuhao Zhang, Peng Qi, Yumo Xu, Jenyuan Wang, Lan Liu, William Yang Wang, Bonan Min, Vittorio Castelli. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Rujun Han, Yuhao Zhang 0004, Peng Qi 0003, Yumo Xu, Jenyuan Wang, Lan Liu 0004, William Yang Wang, Bonan Min, Vittorio Castelli
EMNLP9
2023 Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning
abstract
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang
ACL (1)10
2023 Taxonomy Expansion for Named Entity Recognition
abstract
Karthikeyan K, Yogarshi Vyas, Jie Ma, Giovanni Paolini, Neha John, Shuai Wang, Yassine Benajiba, Vittorio Castelli, Dan Roth, Miguel Ballesteros. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Karthikeyan K, Yogarshi Vyas, Jie Ma 0005, Giovanni Paolini, Neha Anna John, Yassine Benajiba, Vittorio Castelli, Dan Roth 0001, Miguel Ballesteros
EMNLP8
2023 Comparing Biases and the Impact of Multilingual Training across Multiple Languages
abstract
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Sharon Levy, Neha Anna John, Yogarshi Vyas, Jie Ma 0005, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth 0001
EMNLP8
2023 Pre-training Intent-Aware Encoders for Zero- and Few-Shot Intent Classification
abstract
Mujeen Sung, James Gung, Elman Mansimov, Nikolaos Pappas, Raphael Shu, Salvatore Romeo, Yi Zhang, Vittorio Castelli. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Mujeen Sung, James Gung, Elman Mansimov, Nikolaos Pappas 0004, Raphael Shu, Salvatore Romeo, Yi Zhang 0001, Vittorio Castelli
EMNLP8
2023 Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang
ICLR14
2022 Towards Robust Neural Retrieval with Source Domain Synthetic Pre-Finetuning
abstract
Research on neural IR has so far been focused primarily on standard supervised learning settings, where it outperforms traditional term matching baselines. Many practical use cases of such models, however, may involve previously unseen target domains. In this paper, we propose to improve the out-of-domain generalization of Dense Passage Retrieval (DPR) - a popular choice for neural IR - through synthetic data augmentation only in the source domain. We empirically show that pre-finetuning DPR with additional synthetic data in its source domain (Wikipedia), which we generate using a fine-tuned sequence-to-sequence generator, can be a low-cost yet effective first step towards its generalization. Across five different test sets, our augmented model shows more robust performance than DPR in both in-domain and zero-shot out-of-domain evaluation.
Revanth Gangi Reddy, Vikas Yadav, Md. Arafat Sultan, Martin Franz, Vittorio Castelli, Heng Ji 0001, Avirup Sil
COLING5
2021 KAAPA: Knowledge Aware Answers from PDF Analysis
abstract
We present KaaPa (Knowledge Aware Answers from Pdf Analysis), an integrated solution for machine reading comprehension over both text and tables extracted from PDFs. KaaPa enables interactive question refinement using facets generated from an automatically induced Knowledge Graph. In addition it provides a concise summary of the supporting evidence for the provided answers by aggregating information across multiple sources. KaaPa can be applied consistently to any collection of documents in English with zero domain adaptation effort. We showcase the use of KaaPa for QA on scientific literature using the COVID-19 Open Research Dataset.
Nicolas R. Fauceglia, Mustafa Canim, Alfio Massimiliano Gliozzo, Jennifer J. Liang, Nancy Xin Ru Wang, Douglas Burdick, Nandana Mihindukulasooriya, Vittorio Castelli, Guy Feigenblat, David Konopnicki, Yannis Katsis, Radu Florian, Yunyao Li 0001, Salim Roukos, Avirup Sil
AAAI8
2021 Synthetic Target Domain Supervision for Open Retrieval QA
abstract
Neural passage retrieval is a new and promising approach in open retrieval question answering. In this work, we stress-test the Dense Passage Retriever (DPR)---a state-of-the-art (SOTA) open domain neural retrieval model---on closed and specialized target domains such as COVID-19, and find that it lags behind standard BM25 in this important real-world setting. To make DPR more robust under domain shift, we explore its fine-tuning with synthetic training examples, which we generate from unlabeled target domain text using a text-to-text generator. In our experiments, this noisy but fully automated target domain supervision gives DPR a sizable advantage over BM25 in out-of-domain settings, making it a more viable model in practice. Finally, an ensemble of BM25 and our improved DPR model yields the best results, further pushing the SOTA for open retrieval QA on multiple out-of-domain test sets.
Revanth Gangi Reddy, Bhavani Iyer, Md. Arafat Sultan, Rong Zhang 0010, Avirup Sil, Vittorio Castelli, Radu Florian, Salim Roukos
SIGIR6
2020 The TechQA Dataset
abstract
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, Scott McCarley, Michael McCawley, Mohamed Nasr, Lin Pan, Cezar Pendus, John Pitrelli, Saurabh Pujar, Salim Roukos, Andrzej Sakrajda, Avi Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Vittorio Castelli, Rishav Chakravarti, Saswati Dana, Anthony Ferritto, Radu Florian, Martin Franz, Dinesh Garg, Dinesh Khandelwal, J. Scott McCarley, Mike McCawley, Mohamed Nasr, Lin Pan 0003, Cezar Pendus, John F. Pitrelli, Saurabh Pujar, Salim Roukos, Andrej Sakrajda, Avirup Sil, Rosario Uceda-Sosa, Todd Ward, Rong Zhang 0010
ACL1
2020 On the Importance of Diversity in Question Generation for QA
abstract
Automatic question generation (QG) has shown promise as a source of synthetic training data for question answering (QA).In this paper we ask: Is textual diversity in QG beneficial for downstream QA?Using top-p nucleus sampling to derive samples from a transformer-based question generator, we show that diversity-promoting QG indeed provides better QA training than likelihood maximization approaches such as beam search.We also show that standard QG evaluation metrics such as BLEU, ROUGE and METEOR are inversely correlated with diversity, and propose a diversity-aware intrinsic measure of overall QG quality that correlates well with extrinsic evaluation on QA.1
Md. Arafat Sultan, Shubham Chandel, Ramón Fernandez Astudillo, Vittorio Castelli
ACL4
2020 Multi-Stage Pre-training for Low-Resource Domain Adaptation
abstract
Rong Zhang, Revanth Gangi Reddy, Md Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avi Sil, Todd Ward. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Rong Zhang 0010, Revanth Gangi Reddy, Md. Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avirup Sil, Todd Ward
EMNLP (1)4
2016 A Joint Model for Answer Sentence Ranking and Answer Extraction
abstract
Answer sentence ranking and answer extraction are two key challenges in question answering that have traditionally been treated in isolation, i.e., as independent tasks. In this article, we (1) explain how both tasks are related at their core by a common quantity, and (2) propose a simple and intuitive joint probabilistic model that addresses both via joint computation but task-specific application of that quantity. In our experiments with two TREC datasets, our joint model substantially outperforms state-of-the-art systems in both tasks.
Md. Arafat Sultan, Vittorio Castelli, Radu Florian
Trans. Assoc. Comput. Linguistics2
2014 CLIR for Informal Content in Arabic Forum Posts
abstract
The field of Cross-Language Information Retrieval (CLIR) addresses the problem of finding documents in some language that are relevant to a question posed in a different language. Retrieving answers to questions written using formal vocabulary from collections of informal documents, as with many types of social media, is a largely unexplored subfield of CLIR. Because formal and informal content are often intermingled, CLIR systems that excel at finding formal content may tend to select formal over informal content. To measure this effect, a test collection annotated for both relevance and informality is needed. This paper describes the development of a small test collection for this task, with questions posed in formal English and the documents consisting of intermixed formal and informal Arabic. Experiments with this collection show that dialect classification can help to recognize informal content, thus improving precision. At the same time, the results indicate that neither dialect-tuned morphological analysis nor a lightweight CLIR approach that minimizes propagation of translation errors yet yield a reliable improvement in recall for informal content when compared to a straightforward document translation architecture.
Mossaab Bagdouri, Douglas W. Oard, Vittorio Castelli
CIKM3
2014 Query-Focused Opinion Summarization for User-Generated Content
Lu Wang 0008, Hema Raghavan, Claire Cardie, Vittorio Castelli
COLING4
2014 Joint question clustering and relevance prediction for open domain non-factoid question answering
abstract
Web searches are increasingly formulated as natural language questions, rather than keyword queries. Retrieving answers to such questions requires a degree of understanding of user expectations. An important step in this direction is to automatically infer the type of answer implied by the question, e.g., factoids, statements on a topic, instructions, reviews, etc. Answer Type taxonomies currently exist for factoid-style questions, but not for open-domain questions. Building taxonomies for non-factoid questions is a harder problem since these questions can come from a very broad semantic space. A few attempts have been made to develop taxonomies for non-factoid questions, but these tend to be too narrow or domain specific. In this paper, we address this problem by modeling the Answer Type as a latent variable that is learned in a data-driven fashion, allowing the model to be more adaptive to new domains and data sets. We propose approaches that detect the relevance of candidate answers to a user question by jointly 'clustering' questions according to the hidden variable, and modeling relevance conditioned on this hidden variable.
Snigdha Chaturvedi, Vittorio Castelli, Radu Florian, Ramesh Nallapati, Hema Raghavan
WWW2
2013 A Sentence Compression Based Framework to Query-Focused Multi-Document Summarization
Lu Wang 0008, Hema Raghavan, Vittorio Castelli, Radu Florian, Claire Cardie
ACL (1)3
2013 Finding What Matters in Questions
Xiaoqiang Luo, Hema Raghavan, Vittorio Castelli, Sameer Maskey, Radu Florian
HLT-NAACL3
2012 Distilling and exploring nuggets from a corpus
abstract
This paper describes a live and scalable system that automatically extracts information nuggets for entities/topics from a continuously updated corpus for effective exploration and analysis. A nugget is a piece of semantic information that (1) must be mapped semantically to the transitive closure of a pre-defined ontology, (2) is explicitly supported by text, and (3) has a natural language description that completely conveys its semantic to a user. Fig. 1 shows a type of nugget "involvement in events" for a person entity (Leon Panetta): each nugget has a short description ("meeting", "news conference") with a list of supporting passages.
Vittorio Castelli, Hema Raghavan, Radu Florian, Ding-Jung Han, Xiaoqiang Luo, Salim Roukos
SIGIR1
2010 Sheepdog, parallel collaborative programming-by-demonstration
Vittorio Castelli, Lawrence D. Bergman, Tessa A. Lau, Daniel Oblinger
Knowl. Based Syst.1
2007 Learning by Combining Observations and User Edits
Vittorio Castelli, Lawrence D. Bergman, Daniel Oblinger
AAAI1
2007 Distributed augmentation-based learning: a learning algorithm for distributed collaborative programming-by-demonstration
abstract
The learning algorithms used in Programming-by-Demonstration (PBD) are either on-line and incremental or off-line and batch. Neither category is entirely suitable for capturing know-how from demonstrations in a distributed, collaborative environment, where multiple experts can independently provide examples to improve the model.In this paper we describe Distributed Augmentation-Based Learning (DABL), the first real-time PBD learning algorithm suited for distributed know-how acquisition. DABL is an incremental learning algorithm that uses a version-control-like paradigm to combine independently constructed procedure models. An expert can check out a procedure model from a repository and modify it by means of new demonstrations or by manually editing it. The expert then reconciles the changes with those concurrently made by other experts and checked into the repository.DABL automatically merges the two procedures, learns new decision points based on reconcilable differences, and identifies conflicts where there are multiple valid ways of combining the changes or where the combination produces an invalid model, that is, one that does not lie in the search space of the learning algorithm.
Vittorio Castelli, Lawrence D. Bergman
IUI1
2007 Evaluating an Automated Tool to Assist Evolutionary Document Generation
abstract
While using how-to documents for guidance in performing computer-based tasks, users often run into problems due to inaccurate, out-of-date and incomplete documentation. These problems are often due to current documentation practices, which fail to keep how-to documents current, accurate, and complete. We believe that automated support for incremental update of how-to-documents, through the use of programming by demonstration and guided walkthrough techniques, is more effective than existing practice and produces documents that cause fewer problems for their users. In this paper, we present a study that evaluates this belief by comparing DocWizards, a tool utilizing these techniques, with a standard word processor. We show that more effective and efficient documentation can be generated by multiple authors using DocWizards in an incremental process, with effort comparable to that incurred using a traditional tool.
Gahgene Gweon, Lawrence D. Bergman, Vittorio Castelli, Rachel K. E. Bellamy
VL/HCC3
2007 Augmentation-Based Learning combining observations and user edits for Programming-by-Demonstration
Vittorio Castelli, Daniel Oblinger, Lawrence D. Bergman
Knowl. Based Syst.1
2006 An evaluation of using programming by demonstration and guided walkthrough techniques for authoring and utilizing documentation
abstract
Much existing documentation is informal and serves to communicate "how-to" knowledge among restricted working groups. Using current practices, such documentation is both difficult to maintain and difficult to use properly.In this paper, we propose a documentation system, called DocWizards, that uses programming by demonstration to support low-cost authoring and guided walkthrough techniques to improve document usability.We report a comparative study between the use of DocWizards and traditional techniques for authoring and following documentation. The study participants showed significant gains in efficiency and reduction in error rates when using DocWizards. In addition, they expressed a clear preference for using the DocWizards tool, both for authoring and for following documentation.
Madhu Prabaker, Lawrence D. Bergman, Vittorio Castelli
CHI3
2006 Augmentation-based learning: combining observations and user edits for programming-by-demonstration
abstract
In this paper we introduce a new approach to Programming-by-Demonstration in which the user is allowed to explicitly edit the procedure model produced by the learning algorithm while demonstrating the task. We describe a new algorithm, Augmentation-Based Learning, that supports this approach by considering both demonstrations and edits as constraints on the hypothesis space, and resolving con icts in favor of edits.
Daniel Oblinger, Vittorio Castelli, Lawrence D. Bergman
IUI2
2006 Structural Periodic Measures for Time-Series Data
Michail Vlachos, Philip S. Yu, Vittorio Castelli, Christopher Meek
Data Min. Knowl. Discov.3
2006 Bounds on expansion in LZ'77-like coding
abstract
We investigate the maximum increase in number of phrases that results from changing k consecutive symbols in a string x having length n parsed using an LZ'77-like algorithm. We consider a class of compression algorithms that partition a sequence into a collection y of nonoverlapping, variable-length phrases and encode them. Each phrase either is a singleton or matches a substring that starts to its left. We show that changing a single symbol of x in position i can yield an expansion that is of order O(n-i)/sup 2/3/ as (n-i)/spl rarr//spl infin/. Our lower bound requires an alphabet size of O(n-i)/sup 1/3/. We also show that changing k consecutive symbols starting from position i can yield an expansion having a similar but somewhat more involved form. The paper contains both analytically derived upper and lower bounds, and algorithms for numerically computing tighter bounds. While deriving the bounds, we provide a detailed analysis of how expansion can arise when changing consecutive symbols. This problem is motivated by management policies for computer systems, such as the IBM Memory eXpansion Technology (MXT) or the IBM iSeries compressed disks, that use LZ'77-like coding on small compression units, such as 1-4 kbyte, and store the compressed data in memory or on disk tracks. Here, when a change of a portion of the compression unit occurs, for example, an L2 cache line, or a 512-byte disk sector, the data is recompressed and potentially stored in a different location. Knowing the maximum expansion, rather than the average expansion, is an important factor for designing policies for allocation and management of memory or disk space.
Vittorio Castelli, Luis A. Lastras
IEEE Trans. Inf. Theory1
2006 Near sufficiency of random coding for two descriptions
abstract
We give a single-letter outer bound for the two-descriptions problem for independent and identically distributed (i.i.d.) sources that is universally close to the El Gamal and Cover (EGC) inner bound. The gaps for the sum and individual rates using a quadratic distortion measure are upper-bounded by 1.5 and 0.5 bits/sample, respectively, and are universal with respect to the source being encoded and the desired distortion levels. Variants of our basic ideas are presented, including upper and lower bounds on the second channel's rate when the first channel's rate is arbitrarily close to the rate-distortion function; these bounds differ, in the limit as the code block length goes to infinity, by not more than 2 bits/sample. An interesting aspect of our methodology is the manner in which the matching single-letter outer bound is obtained, as we eschew common techniques for constructing single-letter bounds in favor of new ideas in the field of rate loss bounds. We expect these techniques to be generally applicable to other settings of interest.
Luis A. Lastras, Vittorio Castelli
IEEE Trans. Inf. Theory2
2005 Near Tightness of the El Gamal and Cover Region for Two Descriptions
abstract
We give a single letter outer bound for the two descriptions problem for iid sources that is universally close to the El Gamal and Cover (EGC) inner bound. The gaps in the quadratic distortion case for the sum and individual rates are upper bounded by 1.5 and 0.5 bits/sample, respectively. These constant bounds are universal with respect to the source being encoded, provided that its variance is finite. They are also universal with respect to the desired distortion levels, under the assumption that, after normalizing the source to have unit variance, D/sub i/ /spl isin/ (0,1) for i /spl isin/ {0,1,2} and D/sub 0/ /spl les/ (D/sub 1//sup -1/ + D/sub 2//sup -1/ - 1)/sup -1/.
Luis A. Lastras, Vittorio Castelli
DCC2
2005 Similarity-Based Alignment and Generalization
Daniel Oblinger, Vittorio Castelli, Tessa A. Lau, Lawrence D. Bergman
ECML2
2005 A Multi-metric Index for Euclidean and Periodic Matching
Michail Vlachos, Zografoula Vagena, Vittorio Castelli, Philip S. Yu
PKDD3
2005 On Periodicity Detection and Structural Periodic Similarity
abstract
This work motivates the need for more flexible structural similarity measures between time-series sequences, which are based on the extraction of important periodic features. Specifically, we present non-parametric methods for accurate periodicity detection and we introduce new periodic distance measures for time-series sequences. The goal of these tools and techniques are to assist in detecting, monitoring and visualizing structural periodic changes. It is our belief that these methods can be directly applicable in the manufacturing industry for preventive maintenance and in the medical sciences for accurate classification and anomaly detection.
Michail Vlachos, Philip S. Yu, Vittorio Castelli
SDM3
2005 DocWizards: a system for authoring follow-me documentation wizards
abstract
Traditional documentation for computer-based procedures is difficult to use: readers have trouble navigating long complex instructions, have trouble mapping from the text to display widgets, and waste time performing repetitive procedures. We propose a new class of improved documentation that we call follow-me documentation wizards. Follow-me documentation wizards step a user through a script representation of a procedure by highlighting portions of the text, as well application UI elements. This paper presents algorithms for automatically capturing follow-me documentation wizards by demonstration, through observing experts performing the procedure. We also present our DocWizards implementation on the Eclipse platform. We evaluate our system with an initial user study that showing that most users have a marked preference for this form of guidance over traditional documentation.
Lawrence D. Bergman, Vittorio Castelli, Tessa A. Lau, Daniel Oblinger
UIST2
2004 Bounds on expansion in LZ'77-like coding
abstract
This paper investigates the maximum increase in number of phrases that results from changing one symbol in a string that has been parsed using an LZ'77-like algorithm. We provide upper and lower bounds to the maximum expansion as a function of the position of the changed symbol and of the string length.
Vittorio Castelli, Luis A. Lastras
ISIT1
2004 Sheepdog: learning procedures for technical support
abstract
Technical support procedures are typically very complex. Users often have trouble following printed instructions describing how to perform these procedures, and these instructions are difficult for support personnel to author clearly. Our goal is to learn these procedures by demonstration, watching multiple experts performing the same procedure across different operating conditions, and produce an executable procedure that runs interactively on the user's desktop. Most previous programming by demonstration systems have focused on simple programs with regular structure, such as loops with fixed-length bodies. In contrast, our system induces complex procedure structure by aligning multiple execution traces covering different paths through the procedure. This paper presents a solution to this alignment problem using Input/Output Hidden Markov Models. We describe the results of a user study that examines how users follow printed directions. We present Sheepdog, an implemented system for capturing, learning, and playing back technical support procedures on the Windows desktop. Finally, we empirically evalute our system using traces gathered from the user study and show that we are able to achieve 73% accuracy on a network configuration task using a procedure trained by non-experts.
Tessa A. Lau, Lawrence D. Bergman, Vittorio Castelli, Daniel Oblinger
IUI3
2003 CSVD: Clustering and Singular Value Decomposition for Approximate Similarity Search in High-Dimensional Spaces
abstract
Nearest-neighbor search of high-dimensionality spaces is critical for many applications, such as content-based retrieval from multimedia databases, similarity search of patterns in data mining, and nearest-neighbor classification. Unfortunately, even with the aid of the commonly used indexing schemes, the performance of nearest-neighbor (NN) queries deteriorates rapidly with the number of dimensions. We propose a method, called Clustering with Singular Value Decomposition (CSVD), which supports efficient approximate processing of NN queries, while maintaining good precision-recall characteristics. CSVD groups homogeneous points into clusters and separately reduces the dimensionality of each cluster using SVD. Cluster selection for NN queries relies on a branch-and-bound algorithm and within-cluster searches can be performed with traditional or in-memory indexing methods. Experiments with texture vectors extracted from satellite images show that CSVD achieves significantly higher dimensionality reduction than plain SVD for the same normalized mean squared error (NMSE), which translates into a higher efficiency in processing approximate NN queries.
Vittorio Castelli, Alexander Thomasian, Chung-Sheng Li
IEEE Trans. Knowl. Data Eng.1
2001 Solarspire: querying temporal solar imagery by content
abstract
In this paper, we describe a novel content-based retrieval application which permits astrophysicists to search large image sequence archives for solar phenomenon, such as solar flares, based on the spatio-temporal behavior of the solar phenomenon. Specifically, images are preprocessed to identify bright and dark spots based on their relative intensity with respect to their neighboring regions. Temporally persistent objects are then extracted from the collection of spots, and their spatio-temporal behavior represented as intensity and size time series. Users define a query in terms of a model of spatio-temporal behaviors through a Web-based interface. The stored intensity and size time series are searched, and series segments that match the specified specified spatio-temporal behavior are returned. The benchmark results based on 2500 satellite images show that the proposed methodology demonstrated better than 85% accuracy on a solar phenomenon previously identified by astrophysicists.
Matthew L. Hill, Vittorio Castelli, Chung-Sheng Li, Yuan-Chi Chang, Lawrence D. Bergman, John R. Smith, Barbara J. Thompson
ICIP (1)2
2000 The Onion Technique: Indexing for Linear Optimization Queries
abstract
This paper describes the Onion technique, a special indexing structure for linear optimization queries. Linear optimization queries ask for top-N records subject to the maximization or minimization of linearly weighted sum of record attribute values. Such query appears in many applications employing linear models and is an effective way to summarize representative cases, such as the top-50 ranked colleges. The Onion indexing is based on a geometric property of convex hull, which guarantees that the optimal value can always be found at one or more of its vertices. The Onion indexing makes use of this property to construct convex hulls in layers with outer layers enclosing inner layers geometrically. A data record is indexed by its layer number or equivalently its depth in the layered convex hull. Queries with linear weightings issued at run time are evaluated from the outmost layer inwards. We show experimentally that the Onion indexing achieves orders of magnitude speedup against sequential linear scan when N is small compared to the cardinality of the set. The Onion technique also enables progressive retrieval, which processes and returns ranked results in a progressive manner. Furthermore, the proposed indexing can be extended into a hierarchical organization of data to accommodate both global and local queries.
Yuan-Chi Chang, Lawrence D. Bergman, Vittorio Castelli, Chung-Sheng Li, Ming-Ling Lo, John R. Smith
SIGMOD Conference3
2000 SPIRE: A Progressive Content-Based Spatial Image Retrieval Engine
abstract
In this demo, we will show the implementation of a content-based SPatial Image Retrieval Engine (SPIRE) for multimodal unstructured data. This architecture provides a framework for retrieving multi-modal data including image, image sequence, time series and parametric data from large archives. Dramatic speedup (from a factor of 4 to 35) has been achieved for many search operations such as template matching, texture feature extraction. This framework has been applied and validated in solar flares and petroleum exploration in which spatial and spatial-temporal phenomena are located.
Chung-Sheng Li, Lawrence D. Bergman, Vittorio Castelli, John R. Smith
SIGMOD Conference3
1999 GUEST EDITORS' INTRODUCTION: Content-Based Access of Image and Video Libraries
Alberto Del Bimbo, Vittorio Castelli, Shih-Fu Chang, Chung-Sheng Li
Comput. Vis. Image Underst.2
1999 Scan: A Hierarchical Algorithm for Similarity Search in Databases Consisting of Long Sequences
Chung-Sheng Li, Philip S. Yu, Vittorio Castelli
Knowl. Inf. Syst.3
1998 MALM: A Framework for Mining Sequence Database at Multiple Abstraction Levels
abstract
Article MALM: a framework for mining sequence database at multiple abstraction levels Share on Authors: Chung-Sheng Li IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NY IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NYView Profile , Philip S. Yu IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NY IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NYView Profile , Vittorio Castelli IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NY IBM Thomas J.Watson Research Center, P.O. Box 704, Yorktown Heights, NYView Profile Authors Info & Claims CIKM '98: Proceedings of the seventh international conference on Information and knowledge managementNovember 1998 Pages 267–272https://doi.org/10.1145/288627.288666Online:01 November 1998Publication History 38citation409DownloadsMetricsTotal Citations38Total Downloads409Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Chung-Sheng Li, Philip S. Yu, Vittorio Castelli
CIKM3
1998 Clustering and Singular Value Decomposition for Approximate Indexing in High Dimensional Spaces
abstract
High-dimensionality indexing of feature spaces is critical for many data-intensive applications such as content-based retrieval of images or video from multimedia databases and similarity retrieval of patterns in data mining.Unfortunately, the performance of nearest neighbor (NN) queries, which are required for similarity search, deteriorates rapidly with the increase in the number of dimensions.We propose the Clustering with Singular Value Decomposition (CSVD) method, which combines clustering and singular value decomposition (SVD) to reduce the numb e r o f i n d e x d i m e nsions, while maintaining a reasonably high precision for a given value of recall.In the proposed CSVD method, homogeneous points are grouped into clusters such that the points in each cluster are more amenable to dimensionality reduction than the original dataset.Experiments with texture vectors extracted from satellite images show that CSVD achieves signi cantly higher dimensionality reduction than SVD for the same fraction of total variance preserved.Conversely, for the same compression ratio CSVD results in an increase in preserved total variance with respect to SVD (e.g., a 70% increase for a 20:1 compression ratio).This translates to a higher eciency in processing approximate NN queries, as quanti ed through experimental results.
Alexander Thomasian, Vittorio Castelli, Chung-Sheng Li
CIKM2
1998 Dynamic Assembly of Views in Data Cubes
abstract
In this paper, we present a method for dynamically assembling views in multi-dimensional data cubes in order to more e#ciently support data analysis and querying involving aggregations. The proposed method decomposes the data cubes into an indexed hierarchy of view elements. The view elements di#er from traditional data cube cells in that they correspond to partial and residual aggregations of the data cube. The view elements provide highly granular building blocks for synthesizing the aggregated and rangeaggregated views of the data cubes. We propose a strategy for selecting and materializing the view elements based on the frequency of view access. This allows the dynamic adaptation of the view element sets to patterns of retrieval. We present a fast and optimal algorithm for selecting non-expansive view element sets that minimize the processing costs for generating a population of aggregated views. We also present a greedy algorithm for selecting redundant view element sets in order...
John R. Smith, Chung-Sheng Li, Vittorio Castelli, Anant Jhingran
PODS3
1997 Deriving Texture Feature Set for Content-Based Retrieval of Satellite Image Database
abstract
In this paper, the performance of similarity retrieval from satellite image databases by using different sets of spatial and transformed-based texture features is evaluated and compared. A benchmark consisting of 37 satellite image clips from various satellite instruments is devised for the experiments. We show that although the proposed feature set perform only slightly better with the Brodatz set, its performance is far superior for the satellite images. The result indicates that more than 25% of the benchmark patterns can be retrieved with more than 80% accuracy by using normalized Euclidean distance. In contrast, less than 10% of the patterns are retrieved with more than 80% accuracy by using transformed-based feature sets (such as those based on Gabor filter or quadrature mirror filter (QMF)).
Chung-Sheng Li, Vittorio Castelli
ICIP (1)2
1997 MMAP: Modified Maximum A Posteriori Algorithm for Image Segmentation in Large Image/Video Databases
abstract
Block-based feature extraction and clustering algorithms usually have to trade off between resolution and accuracy, as larger image block tends to generate more representative features at the expense of clustering resolution. We propose a new postprocessing technique for optimally combining the labeling results from overlapping image regions. This technique, the modified maximum a posteriori (MMAP) method, utilizes both local and global information from the neighborhood of the image region under consideration. Consequently, the resolution of the clustering becomes independent of the accuracy of the feature. Experimental results show dramatic improvement of the classification accuracy over methods that do not postprocess the clustering labels with the MMAP algorithm.
Norbert Strobel, Chung-Sheng Li, Vittorio Castelli
ICIP (1)3
1996 Progressive classification in the compressed domain for large EOS satellite databases
abstract
We introduce a new framework for classifying large images (in the EOS; Earth Observing System) that is more accurate and less computationally expensive than the classical pixel-by-pixel approach. This approach, called progressive classification, is well suited for analyzing large images, such as multispectral satellite scenes, compressed with wavelet-based or block-transform-based transformations. These transformations produce a multiresolution pyramid representation of the data. A progressive classifier analyses the image at the coarsest resolution level, and it decides whether each coefficient corresponds to a homogeneous block of pixels in the original image or to a heterogeneous block. In the first case it labels the block, in the second case it recursively analyzes the region of the image at the immediately finer resolution level. Computational efficiency, compared to the classical approach, results from examining a much smaller number of coefficients than the number of pixels in the original image. Thus, progressive classification is a prime candidate as a content-based search operator for remotely-sensed data.
Vittorio Castelli, Chung-Sheng Li, John Turek, Ioannis Kontoyiannis
ICASSP1
1996 HierarchyScan: A Hierarchical Similarity Search Algorithm for Databases of Long Sequences
abstract
We present a hierarchical algorithm, HierarchyScan, that efficiently locates one-dimensional subsequences within a collection of sequences of arbitrary length. The subsequences identified by HierarchyScan match a given template pattern in a scale- and phase-independent fashion. The idea is to perform correlation between the stored sequences and the template in the transformed domain hierarchically. Only those subsequences whose maximum correlation value is higher than a predefined threshold will be selected. The performance of this approach is compared to the sequential scanning and an order-of-magnitude speedup is observed.
Chung-Sheng Li, Philip S. Yu, Vittorio Castelli
ICDE3
1996 The relative value of labeled and unlabeled samples in pattern recognition with an unknown mixing parameter
abstract
We observe a training set Q composed of l labeled samples {(X/sub 1/,/spl theta//sub 1/),...,(X/sub l/, /spl theta//sub l/)} and u unlabeled samples {X/sub 1/',...,X/sub u/'}. The labels /spl theta//sub i/ are independent random variables satisfying Pr{/spl theta//sub i/=1}=/spl eta/, Pr{/spl theta//sub i/=2}=1-/spl eta/. The labeled observations X/sub i/ are independently distributed with conditional density f/sub /spl theta/i/(/spl middot/) given /spl theta//sub i/. Let (X/sub 0/,/spl theta//sub 0/) be a new sample, independently distributed as the samples in the training set. We observe X/sub 0/ and we wish to infer the classification /spl theta//sub 0/. In this paper we first assume that the distributions f/sub 1/(/spl middot/) and f/sub 2/(/spl middot/) are given and that the mixing parameter is unknown. We show that the relative value of labeled and unlabeled samples in reducing the risk of optimal classifiers is the ratio of the Fisher informations they carry about the parameter /spl eta/. We then assume that two densities g/sub 1/(/spl middot/) and g/sub 2/(/spl middot/) are given, but we do not know whether g/sub 1/(/spl middot/)=f/sub 1/(/spl middot/) and g/sub 2/(/spl middot/)=f/sub 2/(/spl middot/) or if the opposite holds, nor do we know /spl eta/. Thus the learning problem consists of both estimating the optimum partition of the observation space and assigning the classifications to the decision regions. Here, we show that labeled samples are necessary to construct a classification rule and that they are exponentially more valuable than unlabeled samples.
Vittorio Castelli, Thomas M. Cover
IEEE Trans. Inf. Theory1
1995 On the exponential value of labeled samples
abstract
Consider the problem of classifying a sample X0 into one of two classes, using a training set Q. Let Q be composed of l labeled samples {(X1, θ1), …, (Xl, θl)} and u unlabeled samples {X′1, …, X′u}, where the labels θi are i.i.d. Bernoulli(η) random variables over the set {1, 2}, the observations {Xi}i=1l are distributed according to fθi(·) and the unlabeled observations {X′j}j=1u are independently distributed according to the mixture density fX′(·) = ηf1(·) + (1−η)f2(·). We assume that f1(·),f2(·) and η are all unknown. Let f1(·) and f2(·) belong to a known family F, and assume that the mixtures of elements of F are identifiable. Even when the number of unlabeled samples is infinite and the decision regions can therefore be identified, one still needs labeled samples to label the decision regions with the correct classification. Letting R(l, u) denote the optimal probability of error for l labeled and u unlabeled samples, and assuming that the pairwise mixtures of F are identifiable, we obtain the obvious statements R(0, u) = R(0, ∞) = 12, R(1, 0) ⪕ 2ηη, R(∞, u) = R∗, and then prove R(1, ∞) = 2R∗(1−R∗), where R∗ is the Bayes probability of error, and R(l, ∞) = R∗ + exp{ −αl + o(l)}, where the exponent α is given by −log(2√ηη∫ √f1(x)f2(x) dx). Thus the first labeled sample reduces the risk from 12 to 2R∗(1−R∗) and subsequent labeled samples in the training set reduce the probability of error exponentially fast to the Bayes risk.
Vittorio Castelli, Thomas M. Cover
Pattern Recognit. Lett.1