Vo Ngoc Anh

dblp:a/VoNgocAnh · DBLP profile ↗
← Back
15ranked-venue papers
12as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 11 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Information retrieval · 97% Query processing and optimization · 3%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 50% High-performance computing · 50%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › genomic data compression
lossy compression
0.112012
Transformations for the compression of FASTQ quality scores of next-generation sequencing data · Bioinform. 2012
Bioinformatics and computational biology › bioinformatics infrastructure
next-generation sequencing data compression
0.112012
Transformations for the compression of FASTQ quality scores of next-generation sequencing data · Bioinform. 2012
Information retrieval
ranking
0.122005
Simplified similarity scoring using term ranks · SIGIR 2005
Impact transformation: effective and efficient web retrieval · SIGIR 2002
Information retrieval › indexing › inverted index
impact-ordered indexes
0.112006
Pruned query evaluation using pre-computed impacts · SIGIR 2006
Information retrieval › indexing
index compression
0.112006
Improved Word-Aligned Binary Compression for Text Indexing · IEEE Trans. Knowl. Data Eng. 2006
Information retrieval › indexing
inverted index
0.112006
Improved Word-Aligned Binary Compression for Text Indexing · IEEE Trans. Knowl. Data Eng. 2006
Information retrieval › query processing
dynamic pruning
0.112005
Simplified similarity scoring using term ranks · SIGIR 2005
Storage systems
data compression
0.012012
Transformations for the compression of FASTQ quality scores of next-generation sequencing data · Bioinform. 2012
High-performance computing
lossy compression
0.012012
Transformations for the compression of FASTQ quality scores of next-generation sequencing data · Bioinform. 2012
Information retrieval › query processing
early termination
0.012001
Vector-Space Ranking with Effective Early Termination · SIGIR 2001
Information retrieval › indexing
inverted file
0.012001
Vector-Space Ranking with Effective Early Termination · SIGIR 2001
Information retrieval › evaluation
retrieval effectiveness
0.012001
Vector-Space Ranking with Effective Early Termination · SIGIR 2001
Information retrieval › indexing › index compression
inverted index compression
0.011998
Compressed Inverted Files with Reduced Decoding Overheads · SIGIR 1998
Information retrieval
query processing
0.011998
Compressed Inverted Files with Reduced Decoding Overheads · SIGIR 1998
Query processing and optimization › ranking query
ranked query evaluation
0.011998
Compressed Inverted Files with Reduced Decoding Overheads · SIGIR 1998
Information retrieval
evaluation
0.012006
Pruned query evaluation using pre-computed impacts · SIGIR 2006
Information retrieval › retrieval models › probabilistic retrieval model
BM25
0.012005
Simplified similarity scoring using term ranks · SIGIR 2005
Information retrieval
retrieval models
0.012005
Simplified similarity scoring using term ranks · SIGIR 2005
Information retrieval › web search
web information retrieval
0.012002
Impact transformation: effective and efficient web retrieval · SIGIR 2002

Methods — techniques the papers use, named apart from their topics

lossy transformation · 0.3lossless transformation · 0.3impact-sorted indexing · 0.1carry method · 0.1accumulator management · 0.1integer arithmetic · 0.1document-centric scoring · 0.1vector-space similarity · 0.0thresholding · 0.0quantization · 0.0
YearPublicationVenuePosition
2012 Transformations for the compression of FASTQ quality scores of next-generation sequencing data
abstract
MOTIVATION: The growth of next-generation sequencing means that more effective and efficient archiving methods are needed to store the generated data for public dissemination and in anticipation of more mature analytical methods later. This article examines methods for compressing the quality score component of the data to partly address this problem. RESULTS: We compare several compression policies for quality scores, in terms of both compression effectiveness and overall efficiency. The policies employ lossy and lossless transformations with one of several coding schemes. Experiments show that both lossy and lossless transformations are useful, and that simple coding methods, which consume less computing resources, are highly competitive, especially when random access to reads is needed. AVAILABILITY AND IMPLEMENTATION: Our C++ implementation, released under the Lesser General Public License, is available for download at http://www.cb.k.u-tokyo.ac.jp/asailab/members/rwan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Raymond Wan 0001, Vo Ngoc Anh, Kiyoshi Asai
Bioinform.2
2010 Local Modeling for WebGraph Compression
abstract
We describe a simple hierarchical scheme for Webgraph compression, which supports efficient in-memory and from-disk decoding of page neighborhoods, for neighborhoods defined for both incoming and outgoing links. The scheme is highly competitive in terms of both compression effectiveness and decoding speed.
Vo Ngoc Anh, Alistair Moffat
DCC1
2010 Index compression using 64-bit words
abstract
Abstract Modern computers typically make use of 64‐bit words as the fundamental unit of data access. However the decade‐long migration from 32‐bit architectures has not been reflected in compression technology, because of a widespread assumption that effective compression techniques operate in terms of bits or bytes, rather than words. Here we demonstrate that the use of 64‐bit access units, especially in connection with word‐bounded codes, does indeed provide the opportunity for improving the compression performance. In particular, we extend several 32‐bit word‐bounded coding schemes to 64‐bit operation and explore their uses in information retrieval applications. Our results show that the Simple‐8b approach, a 64‐bit word‐bounded code, is an excellent self‐skipping code, and has a clear advantage over its competitors in supporting fast query evaluation when the data being compressed represents the inverted index for a large text collection. The advantages of the new code also accrue on 32‐bit architectures, and for all of Boolean, ranked, and phrase queries; which means that it can be used in any situation. Copyright © 2010 John Wiley & Sons, Ltd.
Vo Ngoc Anh, Alistair Moffat
Softw. Pract. Exp.1
2008 Term Impacts as Normalized Term Frequencies for BM25 Similarity Scoring
Vo Ngoc Anh, Raymond Wan 0001, Alistair Moffat
SPIRE1
2006 Pruning strategies for mixed-mode querying
abstract
Web information retrieval systems face a range of unique challenges, not the least of which is the sheer scale of the data that must be handled. Also specific to web retrieval is that queries may be a mix of Boolean and ranked features, and documents may have static score components that must also be factored into the ranking process. In this paper we consider a range of query semantics used in web retrieval systems, and show that impact-sorted indexes provide support for dynamic pruning mechanisms and in doing so allow fast document-at-a-time resolution of typical mixed-mode queries, even on relatively large volumes of data. Our techniques also extend to more complex query semantics, including the use of phrase, proximity, and structural constraints.
Vo Ngoc Anh, Alistair Moffat
CIKM1
2006 Pruned query evaluation using pre-computed impacts
abstract
Exhaustive evaluation of ranked queries can be expensive, particularly when only a small subset of the overall ranking is required, or when queries contain common terms. This concern gives rise to techniques for dynamic query pruning, that is, methods for eliminating redundant parts of the usual exhaustive evaluation, yet still generating a demonstrably "good enough" set of answers to the query. In this work we propose new pruning methods that make use of impact-sorted indexes. Compared to exhaustive evaluation, the new methods reduce the amount of computation performed, reduce the amount of memory required for accumulators, reduce the amount of data transferred from disk, and at the same time allow performance guarantees in terms of precision and mean average precision. These strong claims are backed by experiments using the TREC Terabyte collection and queries.
Vo Ngoc Anh, Alistair Moffat
SIGIR1
2006 Structured Index Organizations for High-Throughput Text Querying
Vo Ngoc Anh, Alistair Moffat
SPIRE1
2006 Binary codes for locally homogeneous sequences
Alistair Moffat, Vo Ngoc Anh
Inf. Process. Lett.2
2006 Improved Word-Aligned Binary Compression for Text Indexing
abstract
We present an improved compression mechanism for handling the compressed inverted indexes used in text retrieval systems, extending the word-aligned binary coding carry method. Experiments using two typical document collections show that the new method obtains superior compression to previous static codes, without penalty in terms of execution speed
Vo Ngoc Anh, Alistair Moffat
IEEE Trans. Knowl. Data Eng.1
2005 Binary Codes for Non-Uniform Sources
abstract
In many applications of compression, decoding speed is at least as important as compression effectiveness. For example, the large inverted indexes associated with text retrieval mechanisms are best stored compressed, but a working system must also process queries at high speed. Here we present two coding methods that make use of fixed binary representations. They have all of the consequent benefits in terms of decoding performance, but are also sensitive to localized variations in the source data, and in practice give excellent compression. The methods are validated by applying them to various test data, including the index of an 18 GB document collection.
Alistair Moffat, Vo Ngoc Anh
DCC2
2005 Simplified similarity scoring using term ranks
abstract
We propose a method for document ranking that combines a simple document-centric view of text, and fast evaluation strategies that have been developed in connection with the vector space model. The new method defines the importance of a term within a document qualitatively rather than quantitatively, and in doing so reduces the need for tuning parameters. In addition, the method supports very fast query processing, with most of the computation carried out on small integers, and dynamic pruning an effective option. Experiments on a wide range of TREC data show that the new method provides retrieval effectiveness as good as or better than the Okapi BM25 formulation, and variants of language models.
Vo Ngoc Anh, Alistair Moffat
SIGIR1
2005 Inverted Index Compression Using Word-Aligned Binary Codes
Vo Ngoc Anh, Alistair Moffat
Inf. Retr.1
2002 Impact transformation: effective and efficient web retrieval
abstract
We extend the applicability of impact transformation, which is a technique for adjusting the term weights assigned to documents so as to boost the effectiveness of retrieval when short queries are applied to large document collections. In conjunction with techniques called quantization and thresholding, impact transformation allows improved query execution rates compared to traditional vector-space similarity computations, as the number of arithmetic operations can be reduced. The transformation also facilitates a new dynamic query pruning heuristic. We give results based upon the trec web data that show the combination of these various techniques to yield highly competitive retrieval, in terms of both effectiveness and efficiency, for both short and long queries.
Vo Ngoc Anh, Alistair Moffat
SIGIR1
2001 Vector-Space Ranking with Effective Early Termination
abstract
Considerable research effort has been invested in improving the effectiveness of information retrieval systems. Techniques such as relevance feedback, thesaural expansion, and pivoting all provide better quality responses to queries when tested in standard evaluation frameworks. But such enhancements can add to the cost of evaluating queries. In this paper we consider the pragmatic issue of how to improve the cost-effectiveness of searching. We describe a new inverted file structure using quantized weights that provides superior retrieval effectiveness compared to conventional inverted file structures when early termination heuristics are employed. That is, we are able to reach similar effectiveness levels with less computational cost, and so provide a better cost/performance compromise than previous inverted file organisations.
Vo Ngoc Anh, Owen de Kretser, Alistair Moffat
SIGIR1
1998 Compressed Inverted Files with Reduced Decoding Overheads
abstract
Compressed inverted files are the most compact way of indexing large text databases, typically occupying around 10% of the space of the collection they index.The drawback of compression is the need to decompress the index lists during query processing.Here we describe an improved implementation of compressed inverted lists that eliminates almost all redundant decoding and allows extremely fast processing of conjunctive Boolean queries and ranked queries.We also describe a pruning method to reduce the number of candidate documents considered during the evaluation of ranked queries.Experimental results with a database of 510 Mb show that the new mechanism can reduce the CPU and elapsed time for Boolean queries of 4-10 terms to one tenth and one fifth respectively of the standard technique.For ranked queries, the new mechanism reduces both CPU and elapsed time to one third and memory usage to less than one tenth of the standard algorithm, with no degradation in retrieval effectiveness.
Vo Ngoc Anh, Alistair Moffat
SIGIR1