Ingemar J. Cox

dblp:c/IngemarJCox · also Ingemar Johansson Cox · DBLP profile ↗
← Back
106ranked-venue papers
35as first author
2since 2021 · last 2025
0000-0002-6662-417XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 16 first-author · 1 since 2021Artificial intelligence and machine learning · 38 · 21 first-authorDatabases, data management, data science and information retrieval · 33 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Security and privacy · 6 · 4 first-authorSystems, architecture and hardware · 5 · 3 first-authorComputer networks · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
14 papers
Information retrieval · 77% Data mining · 17% Web and social media mining · 6%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Medical and health informatics · 94% Computational social science and digital humanities · 6%
Network and information security
10 papers
Digital forensics and information hiding · 46% Privacy and data protection · 42% Network security · 10%
Artificial intelligence
18 papers
Generative modeling · 29% Deep learning architectures and training · 29% 3D vision · 12%
Computer graphics and multimedia
9 papers
Image and video processing · 55% Image and video coding · 42% Multimedia systems and quality of experience · 1%

Topics — the 30 heaviest of 109, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics › public health › public health informatics
disease surveillance
1.032019
Transfer Learning for Unsupervised Influenza-like Illness Models from Online Search Data · WWW 2019
Multi-Task Learning Improves Disease Models from Web Search · WWW 2018
Enhancing Feature Selection Using Word Embeddings: The Case of Flu Surveillance · WWW 2017
Medical and health informatics › public health › public health informatics
web search-based disease surveillance
0.722019
Transfer Learning for Unsupervised Influenza-like Illness Models from Online Search Data · WWW 2019
Multi-Task Learning Improves Disease Models from Web Search · WWW 2018
Machine learning › Deep learning architectures and training › loss function design
perceptual loss
0.412020
A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model · NeurIPS 2020
Machine learning › Generative modeling
variational autoencoder
0.412020
A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model · NeurIPS 2020
Image and video processing › perceptual modeling › visual perception modeling
human visual system model
0.412020
A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model · NeurIPS 2020
Image and video coding › image quality assessment
perceptual image quality
0.412020
A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model · NeurIPS 2020
Information retrieval › query formulation
query selection
0.422017
Seasonal Web Search Query Selection for Influenza-Like Illness (ILI) Estimation · SIGIR 2017
An uncertainty-aware query selection model for evaluation of IR systems · SIGIR 2012
Privacy and data protection › data aggregation
privacy-preserving data aggregation
0.412019
Privacy-Preserving Crowd-Sourcing of Web Searches with Private Data Donor · WWW 2019
Privacy and data protection
privacy-preserving data analysis
0.412019
Privacy-Preserving Crowd-Sourcing of Web Searches with Private Data Donor · WWW 2019
Information retrieval
retrieval models
0.332014
Estimating global statistics for unstructured P2P search in the presence of adversarial peers · SIGIR 2014
Risky business: modeling and exploiting uncertainty in information retrieval · SIGIR 2009
The Bayesian image retrieval system, PicHunter: theory, implementation, and psychophysical experiments · IEEE Trans. Image Process. 2000
Medical and health informatics › public health › public health informatics
influenza surveillance
0.312017
Enhancing Feature Selection Using Word Embeddings: The Case of Flu Surveillance · WWW 2017
Data mining › dimensionality reduction
feature selection
0.312017
Enhancing Feature Selection Using Word Embeddings: The Case of Flu Surveillance · WWW 2017
Information retrieval
web search
0.312017
Seasonal Web Search Query Selection for Influenza-Like Illness (ILI) Estimation · SIGIR 2017
Information retrieval › retrieval evaluation
interleaving
0.212016
An Improved Multileaving Algorithm for Online Ranker Evaluation · SIGIR 2016
Information retrieval › retrieval evaluation
ranking evaluation
0.212016
An Improved Multileaving Algorithm for Online Ranker Evaluation · SIGIR 2016
Information retrieval
evaluation
0.222012
An uncertainty-aware query selection model for evaluation of IR systems · SIGIR 2012
Topic (query) selection for IR evaluation · SIGIR 2009
Digital forensics and information hiding
watermarking
0.272007
Using Perceptual Models to Improve Fidelity and Provide Resistance to Valumetric Scaling for Quantization Index Modulation Watermarking · IEEE Trans. Inf. Forensics Secur. 2007
Applying informed coding and embedding to design a robust high-capacity watermark · IEEE Trans. Image Process. 2004
Rotation, scale, and translation resilient watermarking for images · IEEE Trans. Image Process. 2001
Medical and health informatics › public health › public health informatics
public health surveillance
0.212015
Learning About Health and Medicine from Internet Data · WSDM 2015
Web and social media mining
social media analysis
0.222017
WSDM 2017 Workshop on Mining Online Health Reports: MOHRS 2017 · WSDM 2017
Learning About Health and Medicine from Internet Data · WSDM 2015
Information retrieval › evaluation › test collection
test collection construction
0.112012
An uncertainty-aware query selection model for evaluation of IR systems · SIGIR 2012
Digital forensics and information hiding
digital forensics
0.112012
Normalized Energy Density-Based Forensic Detection of Resampled Images · IEEE Trans. Multim. 2012
Digital forensics and information hiding › digital forensics › multimedia forensics › image forensics
image forgery detection
0.112012
Normalized Energy Density-Based Forensic Detection of Resampled Images · IEEE Trans. Multim. 2012
Digital forensics and information hiding › digital forensics › multimedia forensics › image forensics
resampling detection
0.112012
Normalized Energy Density-Based Forensic Detection of Resampled Images · IEEE Trans. Multim. 2012
Information retrieval
search engines
0.112019
Privacy-Preserving Crowd-Sourcing of Web Searches with Private Data Donor · WWW 2019
Information retrieval › user behavior
search logs
0.112019
Privacy-Preserving Crowd-Sourcing of Web Searches with Private Data Donor · WWW 2019
Machine learning › Learning paradigms
multi-task learning
0.112018
Multi-Task Learning Improves Disease Models from Web Search · WWW 2018
Information retrieval › evaluation › query performance prediction
query difficulty
0.112009
Topic (query) selection for IR evaluation · SIGIR 2009
Information retrieval › evaluation › test collection
topic selection
0.112009
Topic (query) selection for IR evaluation · SIGIR 2009
Digital forensics and information hiding › watermarking
quantization index modulation
0.112007
Using Perceptual Models to Improve Fidelity and Provide Resistance to Valumetric Scaling for Quantization Index Modulation Watermarking · IEEE Trans. Inf. Forensics Secur. 2007
Information retrieval › search engines
search engine query analysis
0.112015
Learning About Health and Medicine from Internet Data · WSDM 2015

Methods — techniques the papers use, named apart from their topics

watson perceptual model · 0.9fourier transform · 0.9contrast masking · 0.9temporal similarity · 0.8semantic similarity · 0.8regularized regression · 0.8cryptographic protocol for privacy-preserving data aggregation · 0.8multi-task gaussian process · 0.7multi-task elastic net · 0.7word embeddings · 0.6seasonal modeling · 0.6residual correlation · 0.6regression · 0.6simulation · 0.6dirichlet smoothing · 0.6BM25 · 0.6online experimentation · 0.4support vector machine · 0.3
YearPublicationVenuePosition
2025 GENIE: Socially Unbiased Generative Text-to-Image Editing
abstract
Generative diffusion models often exhibit societal biases in sensitive personal attributes such as age, gender, and race. In this work, we describe GENIE – a method to reduce such biases in a variety of classifier-free diffusion models used for image editing. Our method implicitly incorporates debiasing terms together with the user’s explicit edit instruction to reduce bias. This automatic method relieves the user from needing to modify edit instructions in order to avoid bias. Further, no additional training is needed. Experimental results are provided based on modifications to four diffusion models, namely InstructPix2Pix, Stable Diffusion 1.5, Stable Diffusion 2.1, and Stable Diffusion XL. We show that, on average, bias is reduced by 31% in gender, 15% in age, 39% in race.
Julia Kaiwen Lau, Raphael C.-W. Phan, Sailaja Rajanala, Ingemar J. Cox, Arghya Pal
ICASSP4
2023 Neural network models for influenza forecasting with associated uncertainty using Web search activity trends
abstract
Influenza affects millions of people every year. It causes a considerable amount of medical visits and hospitalisations as well as hundreds of thousands of deaths. Forecasting influenza prevalence with good accuracy can significantly help public health agencies to timely react to seasonal or novel strain epidemics. Although significant progress has been made, influenza forecasting remains a challenging modelling task. In this paper, we propose a methodological framework that improves over the state-of-the-art forecasting accuracy of influenza-like illness (ILI) rates in the United States. We achieve this by using Web search activity time series in conjunction with historical ILI rates as observations for training neural network (NN) architectures. The proposed models incorporate Bayesian layers to produce associated uncertainty intervals to their forecast estimates, positioning themselves as legitimate complementary solutions to more conventional approaches. The best performing NN, referred to as the iterative recurrent neural network (IRNN) architecture, reduces mean absolute error by 10.3% and improves skill by 17.1% on average in nowcasting and forecasting tasks across 4 consecutive flu seasons.
Michael Morris, Peter Hayes, Ingemar J. Cox, Vasileios Lampos
PLoS Comput. Biol.3
2020 A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model
abstract
To train Variational Autoencoders (VAEs) to generate realistic imagery requires a loss function that reflects human perception of image similarity. We propose such a loss function based on Watson's perceptual model, which computes a weighted distance in frequency space and accounts for luminance and contrast masking. We extend the model to color images, increase its robustness to translation by using the Fourier Transform, remove artifacts due to splitting the image into blocks, and make it differentiable. In experiments, VAEs trained with the new loss function generated realistic, high-quality image samples. Compared to using the Euclidean distance and the Structural Similarity Index, the images were less blurry; compared to deep neural network based losses, the new approach required less computational resources and generated images with less artifacts.
Steffen Czolbe, Oswin Krause, Ingemar J. Cox, Christian Igel
NeurIPS3
2019 Privacy-Preserving Crowd-Sourcing of Web Searches with Private Data Donor
abstract
Search engines play an important role on the Web, helping users find relevant resources and answers to their questions. At the same time, search logs can also be of great utility to researchers. For instance, a number of recent research efforts have relied on them to build prediction and inference models, for applications ranging from economics and marketing to public health surveillance. However, companies rarely release search logs, also due to the related privacy issues that ensue, as they are inherently hard to anonymize. As a result, it is very difficult for researchers to have access to search data, and even if they do, they are fully dependent on the company providing them. Aiming to overcome these issues, this paper presents Private Data Donor (PDD), a decentralized and private-by-design platform providing crowd-sourced Web searches to researchers. We build on a cryptographic protocol for privacy-preserving data aggregation, and address a few practical challenges to add reliability into the system with regards to users disconnecting or stopping using the platform. We discuss how PDD can be used to build a flu monitoring model, and evaluate the impact of the privacy-preserving layer on the quality of the results. Finally, we present the implementation of our platform, as a browser extension and a server, and report on a pilot deployment with real users.
Vincent Primault, Vasileios Lampos, Ingemar J. Cox, Emiliano De Cristofaro
WWW3
2019 Transfer Learning for Unsupervised Influenza-like Illness Models from Online Search Data
abstract
A considerable body of research has demonstrated that online search data can be used to complement current syndromic surveillance systems. The vast majority of previous work proposes solutions that are based on supervised learning paradigms, in which historical disease rates are required for training a model. However, for many geographical regions this information is either sparse or not available due to a poor health infrastructure. It is these regions that have the most to benefit from inferring population health statistics from online user search activity. To address this issue, we propose a statistical framework in which we first learn a supervised model for a region with adequate historical disease rates, and then transfer it to a target region, where no syndromic surveillance data exists. This transfer learning solution consists of three steps: (i) learn a regularized regression model for a source country, (ii) map the source queries to target ones using semantic and temporal similarity metrics, and (iii) re-adjust the weights of the target queries. It is evaluated on the task of estimating influenza-like illness (ILI) rates. We learn a source model for the United States, and subsequently transfer it to three other countries, namely France, Spain and Australia. Overall, the transferred (unsupervised) models achieve strong performance in terms of Pearson correlation with the ground truth (> .92 on average), and their mean absolute error does not deviate greatly from a fully supervised baseline.
Bin Zou 0006, Vasileios Lampos, Ingemar J. Cox
WWW3
2018 Multi-Task Learning Improves Disease Models from Web Search
abstract
We investigate the utility of multi-task learning to disease surveillance using Web search data. Our motivation is two-fold. Firstly, we assess whether concurrently training models for various geographies - inside a country or across different countries - can improve accuracy. We also test the ability of such models to assist health systems that are producing sporadic disease surveillance reports that reduce the quantity of available training data. We explore both linear and nonlinear models, specifically a multi-task expansion of elastic net and a multi-task Gaussian Process, and compare them to their respective single task formulations. We use influenza-like illness as a case study and conduct experiments on the United States (US) as well as England, where both health and Google search data were obtained. Our empirical results indicate that multi-task learning improves regional as well as national models for the US. The percentage of improvement on mean absolute error increases up to 14.8% as the historical training data is reduced from 5 to 1 year(s), illustrating that accurate models can be obtained, even by training on relatively short time intervals. Furthermore, in simulated scenarios, where only a few health reports (training data) are available, we show that multi-task learning helps to maintain a stable performance across all the affected locations. Finally, we present results from a cross-country experiment, where data from the US improves the estimates for England. As the historical training data for England is reduced, the benefits of multi-task learning increase, reducing mean absolute error by up to 40%.
Bin Zou 0006, Vasileios Lampos, Ingemar J. Cox
WWW3
2017 Seasonal Web Search Query Selection for Influenza-Like Illness (ILI) Estimation
abstract
Influenza-like illness (ILI) estimation from web search data is an important web analytics task. The basic idea is to use the frequencies of queries in web search logs that are correlated with past ILI activity as features when estimating current ILI activity. It has been noted that since influenza is seasonal, this approach can lead to spurious correlations with features/queries that also exhibit seasonality, but have no relationship with ILI. Spurious correlations can, in turn, degrade performance. To address this issue, we propose modeling the seasonal variation in ILI activity and selecting queries that are correlated with the residual of the seasonal model and the observed ILI signal. Experimental results show that re-ranking queries obtained by Google Correlate based on their correlation with the residual strongly favours ILI-related queries.
Niels Dalum Hansen, Kåre Mølbak, Ingemar J. Cox, Christina Lioma
SIGIR3
2017 WSDM 2017 Workshop on Mining Online Health Reports: MOHRS 2017
abstract
The workshop on Mining Online Health Reports (MOHRS) draws upon the rapidly developing field of Computational Health, focusing on textual content that has been generated through the various facets of Web activity. Online user-generated information mining, especially from social media platforms and search engines, has been in the forefront of many research efforts, especially in the fields of Information Retrieval and Natural Language Processing. The incorporation of such data and techniques in a number of health-oriented applications has provided strong evidence about the potential benefits, which include better population coverage, timeliness and the operational ability in places with less established health infrastructure. The workshop aims to create a platform where relevant state-of-the-art research is presented, but at the same time discussions among researchers with cross-disciplinary backgrounds can take place. It will focus on the characterisation of data sources, the essential methods for mining this textual information, as well as potential real-world applications and the arising ethical issues. MOHRS '17 will feature 3 keynote talks and 4 accepted paper presentations, together with a panel discussion session.
Nigel Collier, Nut Limsopatham, Aron Culotta, Mike Conway, Ingemar J. Cox, Vasileios Lampos
WSDM5
2017 Enhancing Feature Selection Using Word Embeddings: The Case of Flu Surveillance
abstract
Health surveillance systems based on online user-generated content often rely on the identification of textual markers that are related to a target disease. Given the high volume of available data, these systems benefit from an automatic feature selection process. This is accomplished either by applying statistical learning techniques, which do not consider the semantic relationship between the selected features and the inference task, or by developing labour-intensive text classifiers. In this paper, we use neural word embeddings, trained on social media content from Twitter, to determine, in an unsupervised manner, how strongly textual features are semantically linked to an underlying health concept. We then refine conventional feature selection methods by a priori operating on textual variables that are sufficiently close to a target concept. Our experiments focus on the supervised learning problem of estimating influenza-like illness rates from Google search queries. A "flu infection" concept is formulated and used to reduce spurious and potentially confounding features that were selected by previously applied approaches. In this way, we also address forms of scepticism regarding the appropriateness of the feature space, alleviating potential cases of overfitting. Ultimately, the proposed hybrid feature selection method creates a more reliable model that, according to our empirical analysis, improves the inference performance (Mean Absolute Error) of linear and nonlinear regressors by 12% and 28.7%, respectively.
Vasileios Lampos, Bin Zou 0006, Ingemar J. Cox
WWW3
2016 Multi-Dueling Bandits and Their Application to Online Ranker Evaluation
abstract
Online ranker evaluation focuses on the challenge of efficiently determining, from implicit user feedback, which ranker out of a finite set of rankers is the best. It can be modeled by dueling bandits, a mathematical model for online learning under limited feedback from pairwise comparisons. Comparisons of pairs of rankers is performed by interleaving their result sets and examining which documents users click on. The dueling bandits model addresses the key issue of which pair of rankers to compare at each iteration.
Brian Brost, Yevgeny Seldin, Ingemar J. Cox, Christina Lioma
CIKM3
2016 Inferring the Socioeconomic Status of Social Media Users Based on Behaviour and Language
Vasileios Lampos, Nikolaos Aletras, Jens K. Geyti, Bin Zou 0006, Ingemar J. Cox
ECIR5
2016 An Improved Multileaving Algorithm for Online Ranker Evaluation
abstract
Online ranker evaluation is a key challenge in information retrieval. An important task in the online evaluation of rankers is using implicit user feedback for inferring preferences between rankers. Interleaving methods have been found to be efficient and sensitive, i.e. they can quickly detect even small differences in quality. It has recently been shown that multileaving methods exhibit similar sensitivity but can be more efficient than interleaving methods. This paper presents empirical results demonstrating that existing multileaving methods either do not scale well with the number of rankers, or, more problematically, can produce results which substantially differ from evaluation measures like NDCG. The latter problem is caused by the fact that they do not correctly account for the similarities that can occur between rankers being multileaved. We propose a new multileaving method for handling this problem and demonstrate that it substantially outperforms existing methods, in some cases reducing errors by as much as 50%.
Brian Brost, Ingemar J. Cox, Yevgeny Seldin, Christina Lioma
SIGIR2
2015 Learning About Health and Medicine from Internet Data
abstract
Surveys show that around 70% of US Internet users consult the Internet when they require medical information. People seek this information using both traditional search engines and via social media. The information created using the search process offers an unprecedented opportunity for applications to monitor and improve the quality of life of people with a variety of medical conditions. In recent years, research in this area has addressed public-health questions such as the effect of media on development of anorexia, developed tools for measuring influenza rates and assessing drug safety, and examined the effects of health information on individual wellbeing. This tutorial will show how Internet data can facilitate medical research, providing an overview of the state-of-the-art in this area. During the tutorial we will discuss the information which can be gleaned from a variety of Internet data sources, including social media, search engines, and specialized medical websites. We will provide an overview of analysis methods used in recent literature, and show how results can be evaluated using publicly-available health information and online experimentation. Finally, we will discuss ethical and privacy issues and possible technological solutions. This tutorial is intended for researchers of user generated content who are interested in applying their knowledge to improve health and medicine.
Elad Yom-Tov, Ingemar J. Cox, Vasileios Lampos
WSDM2
2015 Assessing the impact of a health intervention via user-generated Internet content
abstract
Assessing the effect of a health-oriented intervention by traditional epidemiological methods is commonly based only on population segments that use healthcare services. Here we introduce a complementary framework for evaluating the impact of a targeted intervention, such as a vaccination campaign against an infectious disease, through a statistical analysis of user-generated content submitted on web platforms. Using supervised learning, we derive a nonlinear regression model for estimating the prevalence of a health event in a population from Internet data. This model is applied to identify control location groups that correlate historically with the areas, where a specific intervention campaign has taken place. We then determine the impact of the intervention by inferring a projection of the disease rates that could have emerged in the absence of a campaign. Our case study focuses on the influenza vaccination program that was launched in England during the 2013/14 season, and our observations consist of millions of geo-located search queries to the Bing search engine and posts on Twitter. The impact estimates derived from the application of the proposed statistical framework support conventional assessments of the campaign.
Vasileios Lampos, Elad Yom-Tov, Richard Pebody, Ingemar J. Cox
Data Min. Knowl. Discov.4
2015 Multi-Keyword Multi-Click Advertisement Option Contracts for Sponsored Search
abstract
In sponsored search, advertisement (abbreviated ad) slots are usually sold by a search engine to an advertiser through an auction mechanism in which advertisers bid on keywords. In theory, auction mechanisms have many desirable economic properties. However, keyword auctions have a number of limitations including: the uncertainty in payment prices for advertisers; the volatility in the search engine’s revenue; and the weak loyalty between advertiser and search engine. In this article, we propose a special ad option that alleviates these problems. In our proposal, an advertiser can purchase an option from a search engine in advance by paying an upfront fee, known as the option price. The advertiser then has the right, but no obligation, to purchase among the prespecified set of keywords at the fixed cost-per-clicks (CPCs) for a specified number of clicks in a specified period of time. The proposed option is closely related to a special exotic option in finance that contains multiple underlying assets (multi-keyword) and is also multi-exercisable (multi-click). This novel structure has many benefits: advertisers can have reduced uncertainty in advertising; the search engine can improve the advertisers’ loyalty as well as obtain a stable and increased expected revenue over time. Since the proposed ad option can be implemented in conjunction with the existing keyword auctions, the option price and corresponding fixed CPCs must be set such that there is no arbitrage between the two markets. Option pricing methods are discussed and our experimental results validate the development. Compared to keyword auctions, a search engine can have an increased expected revenue by selling an ad option.
Bowei Chen 0001, Jun Wang 0012, Ingemar J. Cox, Mohan Kankanhalli
ACM Trans. Intell. Syst. Technol.3
2014 Estimating global statistics for unstructured P2P search in the presence of adversarial peers
abstract
A common problem in unstructured peer-to-peer (P2P) information retrieval is the need to compute global statistics of the full collection, when only a small subset of the collection is visible to a peer. Without accurate estimates of these statistics, the effectiveness of modern retrieval models can be reduced. We show that for the case of a probably approximately correct P2P architecture, and using either the BM25 retrieval model or a language model with Dirichlet smoothing, very close approximations of the required global statistics can be estimated with very little overhead and a small extension to the protocol. However, through theoretical modeling and simulations we demonstrate this technique also greatly increases the ability for adversarial peers to manipulate search results. We show an adversary controlling fewer than 10% of peers can censor or increase the rank of documents, or disrupt overall search results. As a defense, we propose a simple modification to the extension, and show global statistics estimation is viable even when up to 40% of peers are adversarial.
Sami Richardson, Ingemar J. Cox
SIGIR2
2013 Retrieval of trending keywords in a peer-to-peer micro-blogging OSN
abstract
We investigate the problem of identifying trending information in a peer-to-peer micro-blogging online social network. In a distributed decentralized environment, the participating nodes do not have access to global statistics such as the frequencies of the keywords and the information creation rate. We propose a two step solution. First, nodes make a local estimate of the frequency of keywords in the network based on their local information. At each iteration a subset of nodes collect this information from a small subset of random nodes in the network and aggregate the results. The most frequently occurring keywords are identified. In the second step, a node requests another small random subset of nodes to identify when, in the recent past, the more frequently occurring keywords were seen in micro-blogs. Once again this information is aggregated the fraction of time within a consecutive period that keywords were encountered is calculated. If this fraction, referred to as the trending fraction, is close to 1, then the keyword is predicted to be trending. A simulation on a network of 10,000 nodes shows that the solution is capable of detecting multiple trending keywords with a moderate increase in bandwidth.
H. Asthana, Ingemar J. Cox
CIKM2
2013 Ranked Accuracy and Unstructured Distributed Search
Sami Richardson, Ingemar J. Cox
ECIR2
2013 Trending topics in a peer-to-peer micro-blogging social network
abstract
We investigate the problem of identifying trending information in a peer-to-peer micro-blogging social network. Trending topics are valuable since they often reflect a news-worthy event and invite users to search on the topic to see what others are posting about the event. Whilst identifying trending topics in a centralized system is relatively straightforward, in a peer-to-peer environment the participating nodes do not have access to global statistics such as the micro-blog post creation rate and the frequencies of the keywords.
H. Asthana, Ingemar J. Cox
P2P2
2013 On unstructured distributed search over BitTorrent
abstract
Current BitTorrent discovery methods rely on either centralised systems or structured peer-to-peer (P2P) networks. These methods present security weaknesses that can be exploited in order to censor or remove information from the network. To alleviate this threat, we propose incorporating an unstructured peer-to-peer information discovery mechanism that can be used in the event that the centralised or structured P2P mechanisms are compromised. Unstructured P2P information discovery has fewer security weaknesses. However, in this case, the performance of the search is nondeterministic since it is not practical to perform an exhaustive search. The search performance then strongly depends on the distribution of documents in the network. To determine the practicality of unstructured P2P search over BitTorrent, we first conducted a 64 day study of BitTorrent activities, looking at the distribution of 1.6 million torrents on 5.4 million peers. We found that the distribution of torrents follows a power law which is not amenable to unstructured search. To address this, we introduce a simple modification to BitTorrent which enables each peer to index a random subset of tracking data, i.e. the torrent ID and list of participating nodes. A successful search is then one that finds a peer with tracking data, rather than a peer directly participating in the torrent. The distribution of this tracking data is shown to be capable of supporting an accurate unstructured search for torrents.We assess the overheads introduced by our extension and conclude that we would require small amounts of bandwidth, that are easily provided by current home broadband capabilities. We also simulate our extension to verify our model and to explore our extension's capabilities in different situations. We demonstrate that our extension can satisfy PAC search queries for torrents, under network churn and complex node behaviours.
William Mayor, Ingemar J. Cox
P2P2
2013 Increasing ranked accuracy for unstructured distributed search with dynamic replication
abstract
In unstructured peer-to-peer networks it is necessary to query every node to guarantee finding a particular document. This is usually not feasible, so search is probabilistic. In previous work the metric of rank-accuracy was proposed to measure the performance of a query. It compares the set of documents retrieved from a search of a random subset of nodes to the set that would have been retrieved from an exhaustive search of all nodes, assigning a higher weight to higher ranked documents. It was shown that rank-accuracy can be significantly increased by replicating documents across nodes in proportion to their weighted retrieval rate, a function of query frequency and rank. However, this work assumed that the query distribution is known and unchanging, both of which are seldom true for practical systems. In this paper we make no such assumptions. We investigate how the popularity- and rank-aware distribution of documents can be approximated by replicating documents as queries are made. Using a simulated network of 10,000 nodes and 10,000,000 queries drawn from a Yahoo! web search engine log, we show that such a scheme can achieve over a 450% increase in overall rank-accuracy when compared to a uniform random distribution of documents. We also show this increase can be improved up to about 650% when the dynamically replicated documents are pruned down to just the terms that appear in queries.
Sami Richardson, Ingemar J. Cox
P2P2
2012 On Aggregating Labels from Multiple Crowd Workers to Infer Relevance of Documents
Mehdi Hosseini 0001, Ingemar J. Cox, Natasa Milic-Frayling, Gabriella Kazai, Vishwa Vinay
ECIR2
2012 An uncertainty-aware query selection model for evaluation of IR systems
abstract
We propose a mathematical framework for query selection as a mechanism for reducing the cost of constructing information retrieval test collections. In particular, our mathematical formulation explicitly models the uncertainty in the retrieval effectiveness metrics that is introduced by the absence of relevance judgments. Since the optimization problem is computationally intractable, we devise an adaptive query selection algorithm, referred to as Adaptive, that provides an approximate solution. Adaptive selects queries iteratively and assumes that no relevance judgments are available for the query under consideration. Once a query is selected, the associated relevance assessments are acquired and then used to aid the selection of subsequent queries. We demonstrate the effectiveness of the algorithm on two TREC test collections as well as a test collection of an online search engine with 1000 queries. Our experimental results show that the queries chosen by Adaptive produce reliable performance ranking of systems. The ranking is better correlated with the actual systems ranking than the rankings produced by queries that were selected using the considered baseline methods.
Mehdi Hosseini 0001, Ingemar J. Cox, Natasa Milic-Frayling, Milad Shokouhi, Emine Yilmaz
SIGIR2
2012 Normalized Energy Density-Based Forensic Detection of Resampled Images
abstract
We propose a new method to detect resampled imagery. The method is based on examining the normalized energy density present within windows of varying size in the second derivative of the image in the frequency domain, and exploiting this characteristic to derive a 19-D feature vector that is used to train a SVM classifier. Experimental results are reported on 7500 raw images from the BOSS database. Comparison with prior work reveals that the proposed algorithm performs similarly for resampling rates greater than 1, and is superior to prior work for resampling rates less than 1. Experiments are performed for both bilinear and bicubic interpolations, and qualitatively similar results are observed for each. Results are also provided for the detection of resampled imagery with noise corruption and JPEG compression. As expected, some degradation in performance is observed as the noise increases or the JPEG quality factor declines.
Xiaoying Feng, Ingemar J. Cox, Gwenaël J. Doërr
IEEE Trans. Multim.2
2011 Prioritizing relevance judgments to improve the construction of IR test collections
abstract
We consider the problem of optimally allocating a fixed budget to construct a test collection with associated relevance judgements, such that it can (i) accurately evaluate the relative performance of the participating systems, and (ii) generalize to new, previously unseen systems. We propose a two stage approach. For a given set of queries, we adopt the traditional pooling method and use a portion of the budget to evaluate a set of documents retrieved by the participating systems. Next, we analyze the relevance judgments to prioritize the queries and remaining pooled documents for further relevance assessments. The query prioritization is formulated as a convex optimization problem, thereby permitting efficient solution and providing a flexible framework to incorporate various constraints. Query-document pairs with the highest priority scores are evaluated using the remaining budget. We evaluate our resource optimization approach on the TREC 2004 Robust track collection. We demonstrate that our optimization techniques are cost efficient and yield a significant improvement in the reusability of the test collections.
Mehdi Hosseini 0001, Ingemar J. Cox, Natasa Milic-Frayling, Trevor J. Sweeting, Vishwa Vinay
CIKM2
2011 An energy-based method for the forensic detection of Re-sampled images
abstract
We propose a new method to detect re-sampled imagery. The method is based on examining the normalized energy density present within windows of varying size in the second derivative of the frequency domain, and exploiting this characteristic to derive a 19-dimensional feature vector that is used to train a SVM classifier. Experimental results are reported on 7,500 raw images from the BOSS database. Comparison with prior work reveals that the proposed algorithm performs similarly for re-sampling rates greater than 1, and is superior to prior work for re-sampling rates less than 1. Experiments are performed for both bilinear and bicubic interpolation, and qualitatively similar results are observed for each. Results are also provided for the detection of re-sampled imagery that subsequently undergoes JPEG compression. Results are quantitatively similar with some small degradation in performance as the quality factor is reduced.
Xiaoying Feng, Ingemar J. Cox, Gwenaël J. Doërr
ICME2
2011 A comparison of extended fingerprint hashing and locality sensitive hashing for binary audio fingerprints
abstract
Hash tables have been proposed for the indexing of high-dimensional binary vectors, specifically for the identification of media by fingerprints. In this paper we develop a new model to predict the performance of a hash-based method (Fingerprint Hashing) under varying levels of noise. We show that by the adjustment of two parameters, robustness to a higher level of noise is achieved. We extend Fingerprint Hashing to a multi-table range search (Extended Fingerprint Hashing) and show this approach also increases robustness to noise. We then show the relationship between Extended Fingerprint Hashing and Locality Sensitive Hashing and investigate design choices for dealing with higher noise levels. If index size must be held constant, the Extended Fingerprint Hash is a superior method. We also show that to achieve similar performance at a given level of noise a Locality Sensitive Hash requires nearly a six-fold increase in index size which is likely to be impractical for many applications.
Kimberly Moravec, Ingemar J. Cox
ICMR2
2010 Improving Query Correctness Using Centralized Probably Approximately Correct (PAC) Search
Ingemar J. Cox, Jianhan Zhu, Ruoxun Fu, Lars Kai Hansen
ECIR1
2009 Entropy-Based Static Index Pruning
Ingemar J. Cox
ECIR2
2009 Risk-Aware Information Retrieval
Jianhan Zhu, Jun Wang 0012, Michael J. Taylor 0001, Ingemar J. Cox
ECIR4
2009 Re-ranking Documents Based on Query-Independent Document Specificity
Ingemar J. Cox
FQAS2
2009 Data Hiding and the Statistics of Images
Ingemar J. Cox
IWDW1
2009 Risky business: modeling and exploiting uncertainty in information retrieval
abstract
Most retrieval models estimate the relevance of each document to a query and rank the documents accordingly. However, such an approach ignores the uncertainty associated with the estimates of relevancy. If a high estimate of relevancy also has a high uncertainty, then the document may be very relevant or not relevant at all. Another document may have a slightly lower estimate of relevancy but the corresponding uncertainty may be much less. In such a circumstance, should the retrieval engine risk ranking the first document highest, or should it choose a more conservative (safer) strategy that gives preference to the second document? There is no definitive answer to this question, as it depends on the risk preferences of the user and the information retrieval system. In this paper we present a general framework for modeling uncertainty and introduce an asymmetric loss function with a single parameter that can model the level of risk the system is willing to accept. By adjusting the risk preference parameter, our approach can effectively adapt to users' different retrieval strategies.
Jianhan Zhu, Jun Wang 0012, Ingemar J. Cox, Michael J. Taylor 0001
SIGIR3
2009 Topic (query) selection for IR evaluation
abstract
The need for evaluating large amounts of topics (queries) makes IR evaluation an uneasy task. In this paper, we study a topic selection problem for IR evaluation. The selection criterion is based on the overall difficulty of the chosen set, as well as the uncertainty of the final IR metric applied to the systems. Our preliminary experiments demonstrate that our approach helps to identify a set of topics that provides confident estimates of systems' performance while keeping the requirement of the query difficulty.
Jianhan Zhu, Jun Wang 0012, Vishwa Vinay, Ingemar J. Cox
SIGIR4
2008 Estimating retrieval effectiveness using rank distributions
abstract
In this paper, we consider the task of estimating query effectiveness, i.e., assessment of the retrieval system performance in absence of the user relevance judgments. In our approach we model the score associated with each document in the result set as a Gaussian random variable. The mean and the variance of each document score can then be used to estimate the probability that a document will be ranked above another one and thus calculate the expected rank of the document in the ranked list. We propose to measure the effectiveness of the system performance by comparing the predicted and actual ranks of the retrieved documents. In our experiments we consider two retrieval models and five document scoring methods and evaluate their impact on the proposed estimation measures. Our experiments with standardized data sets that include document relevance judgments and the task of predicting the relative query effectiveness show that the expected rank metric is robust to variations in document scoring and retrieval algorithms.
Vishwa Vinay, Natasa Milic-Frayling, Ingemar J. Cox
CIKM3
2008 Detection of +/-1 LSB steganography based on the amplitude of histogram local extrema
abstract
Recently Zhang et al described an algorithm for the detection of plusmn1 LSB steganography based on the statistics of the amplitudes of local extrema in the greylevel histogram. Experimental results demonstrated performance comparable or superior to other state-of-the-art algorithms. In this paper, we describe improvements to this algorithm to (i) reduce the noise associated with border effects in the histogram, and (ii) extend the analysis to amplitudes of local extrema in the 2D adjacency histogram. Experimental results on a composite database of 7125 images, averaged over a 20-fold cross validation, with classification based on Fisher linear discriminant analysis, demonstrate that the improved algorithm exhibits significantly better performance. The experimetal results are reported in the form of receiver operating characteristic (ROC) curves and summarized by computing the area under the ROC curve (AUC). The new algorithm, using 10 features derived from the ID and 2D histograms, has an AUC value of 0.77 compared to 0.57 for the original algorithm. It also significantly outperforms other state-of-the-art steganalysers.
Giacomo Cancelli, Gwenaël J. Doërr, Ingemar J. Cox, Mauro Barni
ICIP3
2008 A comparative study of +/- steganalyzers
abstract
We compare the performance of three steganalysis system for detection of plusmn1 steganography. We examine the relative performance of each system on three commonly used image databases. Experimental results clearly demonstrate that both absolute and relative performance of all three algorithms vary considerably across databases. This sensitivity suggests that considerably more work is needed to develop databases that are more representative of diverse imagery. In addition, we investigate how performance varies based on a variety of training and testing assumptions, specifically (i) that training and testing are performed for a fixed and known embedding rate, (ii) training is performed at one embedding rate, but testing is over a range of embedding rates, (iii) training and testing are performed over a range of embedding rates. As expected, experimental results show that performance under (ii) and (iii) is inferior to (i). The experimental results also suggest that test results for different embedding rates should not be consolidated into a single score, but rather reported separately. Otherwise, good performance at high embedding rates may mask poor performance at low embedding rates.
Giacomo Cancelli, Gwenaël J. Doërr, Mauro Barni, Ingemar J. Cox
MMSP4
2008 Ranked-Listed or Categorized Results in IR: 2 Is Better Than 1
Ingemar J. Cox, Mark Levene
NLDB2
2008 Watermarking, Steganography and Content Forensics
Ingemar J. Cox
SECRYPT1
2007 Improved Spread Transform Dither Modulation using a Perceptual Model: Robustness to Amplitude Scaling and JPEG Compression
abstract
Spread transform dither modulation (STDM) is a form of quantization index modulation (QIM) that is more robust to re-quantization. However, the robustness of STDM to JPEG compression is still very poor and it remains very sensitive to amplitude scaling. Here, we show how a perceptual model that scales linearly with amplitude scaling can be used to (i) provide robustness to amplitude scaling, (ii) reduce the perceptual distortion at the embedder and (iii) significantly improve the robustness to re-quantization.
Qiao Li 0008, Ingemar J. Cox
ICASSP (2)2
2007 JPEG Based Conditional Entropy Coding for Correlated Steganography
abstract
Correlated steganography considers the case in which the cover work is chosen to be correlated with the covert message that is to be hidden. The advantage of this is that, at least theoretically, the number of bits needed to encode the hidden message can be considerably reduced since it is based on the conditional entropy of the message given the cover. This may be much less than the entropy of the message itself. And if the number of bits needed to embed the hidden message is significantly reduced, then it is more likely that the steganographic algorithm will be secure, i.e. undetectable. In this paper, we describe an example of correlated steganography. Specifically, we are interested in embedding a covert image into a cover image. Comparative experiments indicate that selecting a cover Work that is correlated with the covert message can reduce the number of bits needed to represent the covert image below that needed by standard JPEG compression, provided the two images are sufficiently correlated.
Ingemar J. Cox
ICME2
2007 Steganalysis for LSB Matching in Images with High-frequency Noise
abstract
Considerable progress has been made in the detection of steganographic algorithms based on replacement of the least significant bit (LSB) plane. However, if LSB matching, also known as -1 embedding, is used, the detection rates are considerably reduced. In particular, since LSB embedding is modeled as an additive noise process, detection is especially poor for images that exhibit high-frequency noise - the high-frequency noise is often incorrectly thought to be indicative of a hidden message. To overcome this, we propose a targeted steganalysis algorithm that exploits the fact that after LSB matching, the local maxima of an images graylevel or color histogram decrease and the local minima increase. Consequently, the sum of the absolute differences between local extrema and their neighbors in the intensity histogram of stego images will be smaller than for cover images. Experimental results on two datasets, each of 2000 images, demonstrate that this method has superior results compared with other recently proposed algorithms when the images contain high-frequency noise, e.g. never-compressed imagery such as high-resolution scans of photographs and video. However, the method is inferior to the prior art when applied to decompressed imagery with little or no high-frequency noise.
Ingemar J. Cox, Gwenaël J. Doërr
MMSP2
2007 Using Perceptual Models to Improve Fidelity and Provide Resistance to Valumetric Scaling for Quantization Index Modulation Watermarking
abstract
Traditional quantization index modulation (QIM) methods are based on a fixed quantization step size, which may lead to poor fidelity in some areas of the content. A more serious limitation of the original QIM algorithm is its sensitivity to valumetric changes (e.g., changes in amplitude). In this paper, we first propose using Watson's perceptual model to adaptively select the quantization step size based on the calculated perceptual "slack". Experimental results on 1000 images indicate improvements in fidelity as well as improved robustness in high-noise regimes. Watson's perceptual model is then modified such that the slacks scale linearly with valumetric scaling, thereby providing a QIM algorithm that is theoretically invariant to valumetric scaling. In practice, scaling can still result in errors due to cropping and roundoff that are an indirect effect of scaling. Two new algorithms are proposed - the first based on traditional QIM and the second based on rational dither modulation. A comparison with other methods demonstrates improved performance over other recently proposed valumetric-invariant QIM algorithms, with only small degradations in fidelity
Qiao Li 0008, Ingemar J. Cox
IEEE Trans. Inf. Forensics Secur.2
2006 Measuring the Complexity of a Collection of Documents
Vishwa Vinay, Ingemar J. Cox, Natasa Milic-Frayling, Kenneth R. Wood
ECIR2
2006 Toward a Better Understanding of Dirty Paper Trellis Codes
abstract
Dirty paper trellis codes have been introduced as an alternative to lattice codes to implement watermarking as communications with side information. Their key feature is robustness against valuemetric scaling in comparison with lattice codes. Despite the strong academic recognition, parametrization issues remain unclear. For instance, the impact of the trellis configuration on performance is still not well understood. In this paper, experiments on synthetic signals are reported to investigate how the trellis configuration influences the bit error rate and the computational complexity.
Chin Kiong Wang, Gwenaël J. Doërr, Ingemar J. Cox
ICASSP (2)3
2006 Watermarking Is Not Cryptography
Ingemar J. Cox, Gwenaël J. Doërr, Teddy Furon
IWDW1
2006 Spread Transform Dither Modulation using a Perceptual Model
abstract
In previous work, we demonstrated how perceptual modeling can be applied to dither modulated quantization index modulation and rational dither modulation, to improve both robustness and fidelity. These algorithms were shown to be significantly more robust to valumetric scaling. However, they, and their predecessors, remain extremely sensitive to re-quantization which commonly occurs due to JPEG compression, numerical rounding and analog-to-digital conversion. It is well known that spread transform dither modulation (STDM) is more robust to re-quantization. In this paper we describe how to incorporate a perceptual model into this framework and present two algorithms based on Watson's perceptual model. Experimental results of robustness to JPEG compression are reported for 1000 images at embedding rates of 1/32 and 1/320. At the high embedding rate, the robustness of the two algorithms is the same as STDM but the perceptual distortion is reduced 23 to about 4, based on Watson's perceptual distance. At the lower embedding rate, we simultaneously observed superior robustness to STDM as well as improved fidelity. If the perceptual distance rather than the document-to-watermark ratio (DWR) is held fixed, then the two adaptive spread transform methods exhibit significant improvements in robustness to JPEG compression
Qiao Li 0008, Gwenaël J. Doërr, Ingemar J. Cox
MMSP3
2006 On ranking the effectiveness of searches
abstract
There is a growing interest in estimating the effectiveness of search. Two approaches are typically considered: examining the search queries and examining the retrieved document sets. In this paper, we take the latter approach. We use four measures to characterize the retrieved document sets and estimate the quality of search. These measures are (i) the clustering tendency as measured by the Cox-Lewis statistic, (ii) the sensitivity to document perturbation, (iii) the sensitivity to query perturbation and (iv) the local intrinsic dimensionality. We present experimental results for the task of ranking 200 queries according to the search effectiveness over the TREC (discs 4 and 5) dataset. Our ranking of queries is compared with the ranking based on the average precision using the Kendall t statistic. The best individual estimator is the sensitivity to document perturbation and yields Kendall t of 0.521. When combined with the clustering tendency based on the Cox-Lewis statistic and the query perturbation measure, it results in Kendall t of 0.562 which to our knowledge is the highest correlation with the average precision reported to date.
Vishwa Vinay, Ingemar J. Cox, Natasa Milic-Frayling, Kenneth R. Wood
SIGIR2
2006 The web structure of e-government - developing a methodology for quantitative evaluation
abstract
In this paper we describe preliminary work that examines whether statistical properties of the structure of websites can be an informative measure of their quality. We aim to develop a new method for evaluating e-government. E-government websites are evaluated regularly by consulting companies, international organizations and academic researchers using a variety of subjective measures. We aim to improve on these evaluations using a range of techniques from webmetric and social network analysis. To pilot our methodology, we examine the structure of government audit office sites in Canada, the USA, the UK, New Zealand and the Czech Republic.We report experimental values for a variety of characteristics, including the connected components, the average distance between nodes, the distribution of paths lengths, and the indegree and outdegree. These measures are expected to correlate with (i) the navigability of a website and (ii) with its "nodalityö which is a combination of hubness and authority. Comparison of websites based on these characteristics raised a number of issues, related to the proportion of non-hyperlinked content (e.g. pdf and doc files) within a site, and both the very significant differences in the size of the websites and their respective national populations. Methods to account for these issues are proposed and discussed.There appears to be some correlation between the values measured and the league tables reported in the literature. However, this multi dimensional analysis provides a richer source of evaluative techniques than previous work. Our analysis indicates that the US and Canada provide better navigability, much better than the UK; however, the UK site is shown to have the strongest "nodalityö on the Web.
Vaclav Petricek, Tobias Escher, Ingemar J. Cox, Helen Z. Margetts
WWW3
2006 Can constrained relevance feedback and display strategies help users retrieve items on mobile devices?
Vishwa Vinay, Ingemar J. Cox, Natasa Milic-Frayling, Kenneth R. Wood
Inf. Retr.2
2005 Evaluating Relevance Feedback Algorithms for Searching on Small Displays
Vishwa Vinay, Ingemar J. Cox, Natasa Milic-Frayling, Kenneth R. Wood
ECIR2
2005 Using Perceptual Models to Improve Fidelity and Provide Invariance to Valumetric Scaling for Quantization Index Modulation Watermarking
abstract
Quantization index modulation (QIM) is a computationally efficient method of watermarking with side information. This paper proposes two improvements to the original algorithm. First, the fixed quantization step size is replaced with an adaptive step size that is determined using Watson's perceptual model. Experimental results on a database of 1000 images illustrate significant improvements in both fidelity and robustness to additive white Gaussian noise. Second, modifying the Watson model such that it scales linearly with valumetric (amplitude) scaling, results in a QIM algorithm that is invariant to valumetric scaling. Experimental results compare this algorithm with both the original QIM and an adaptive QIM and demonstrate superior performance.
Qiao Li 0008, Ingemar J. Cox
ICASSP (2)2
2005 An efficient algorithm for informed embedding of dirty-paper trellis codes for watermarking
abstract
Dirty paper trellis codes are a form of watermarking with side information. These codes have the advantage of being invariant to valuemetric scaling of the cover work. However, the original proposal requires a computational expensive second stage, informed embedding, to embed the chosen code into the cover Work. In this paper, we present a computational efficient algorithm for informed embedding. This is accomplished by recognizing that all possible code words are uniformly distributed on the surface of a high n-dimensional sphere. Each codeword is then contained within an (n - 1)-dimensional region which defines an n-dimensional cone with the centre of the sphere. This approximates the detection region. This is equivalent to the detection region for normalized correlation detection, for which there are known analytic methods to embed a watermark in a cover work. We use a previously described technique for embedding with a constant robustness. However, rather than moving the cover work to the closest Euclidean point on the defined surface, we find the point on the surface which has the smallest perceptual distortion. Experimental results on 2000 images demonstrate a 600-fold computational improvement together with an improved quality of embedding.
Gwenaël J. Doërr, Ingemar J. Cox, Matthew L. Miller
ICIP (1)3
2005 A comparison of dimensionality reduction techniques for text retrieval
abstract
The growth of digital information increases the need to build better techniques for automatically storing, organizing and retrieving it. Much of this information is textual in nature and existing representation models struggle to deal with the high dimensionality of the resulting feature space. Techniques like latent semantic indexing address, to some degree, the problem of high dimensionality in information retrieval. However, promising alternatives, like random mapping (RM), have yet to be completely studied in this context. In this paper, we show that despite the attention RM has received in other applications, in the case of text retrieval it is outperformed not only by principal component analysis (PCA) and independent component analysis (ICA) but also by a simple noise reduction algorithm.
Vishwa Vinay, Ingemar J. Cox, Kenneth R. Wood, Natasa Milic-Frayling
ICMLA2
2005 Information Transmission and Steganography
Ingemar J. Cox, Ton Kalker, Georg Pakura, Mathias Scheel
IWDW1
2005 Rational dither modulation watermarking using a perceptual model
abstract
Quantization index modulation (QIM) is a computationally efficient method of informed watermarking. However, the original method is particularly sensitive to variations in the amplitude of the signal. Previously, we proposed using a modification of Watson's perceptual model to adaptively adjust the quantization index step size. This simultaneously improved both the robustness and fidelity of the watermarked image and, most importantly, provided invariance (to a large degree) to valumetric scaling. Contemporaneously, rational dither modulation was proposed as an alternative QIM with valumetric invariance. In this paper, we combine the two methods and compare the performance of the new algorithm with our previous results. Experimental results demonstrate that the new algorithm outperforms the previous algorithms over the entire range of valumetric scale factors, albeit at the expense of a small decrease in fidelity. However all algorithms have a superior performance and improved fidelity compared with QIM
Qiao Li 0008, Ingemar J. Cox
MMSP2
2004 Using perceptual distance to improve the selection of dirty paper trellis codes for watermarking
abstract
Previous watermarking research based on dirty paper trellis coding [G.L. Miller et al., 2002] proposed a method for informed coding by which the best code from the set of codewords representing the message was selected based on maximizing the linear correlation between the codewords and the original cover Work. However, this does not guarantee that the linear correlation is maximized in the watermarked cover Work. This is because the chosen codeword must be attenuated due to fidelity constraints. Since there is no clear relationship between linear correlation and fidelity, a codeword that is chosen to maximize linear correlation may be very difficult to embed if it is perceptually very different from the underlying cover Work. We show that this is in fact the case and suggest a solution to this problem that involves a cost function that is a linear combination of perceptual distance and linear correlation. Experimental results demonstrate 50% and 25% improvements in bit and message error rates respectively.
Chin Kiong Wang, Matthew L. Miller, Ingemar J. Cox
MMSP3
2004 Evaluating Relevance Feedback and Display Strategies for Searching on Small Displays
Vishwa Vinay, Ingemar J. Cox, Natasa Milic-Frayling, Kenneth R. Wood
SPIRE2
2004 Applying informed coding and embedding to design a robust high-capacity watermark
abstract
We describe a new watermarking system based on the principles of informed coding and informed embedding. This system is capable of embedding 1380 bits of information in images with dimensions 240 x 368 pixels. Experiments on 2000 images indicate the watermarks are robust to significant valumetric distortions, including additive noise, low-pass filtering, changes in contrast, and lossy compression. Our system encodes watermark messages with a modified trellis code in which a given message may be represented by a variety of different signals, with the embedded signal selected according to the cover image. The signal is embedded by an iterative method that seeks to ensure the message will not be confused with other messages, even after addition of noise. Fidelity is improved by the incorporation of perceptual shaping into the embedding process. We show that each of these three components improves performance substantially.
Matthew L. Miller, Gwenaël J. Doërr, Ingemar J. Cox
IEEE Trans. Image Process.3
2002 Dirty-paper trellis codes for watermarking
abstract
Informed coding is the practice of representing watermark messages with patterns that are dependent on the cover works. This requires the use of a dirty-paper code, in which each message is represented by a large number of alternative vectors. Most previous dirty-paper codes are based on lattice codes, in which each code vector, or pattern, is a point in a regular lattice. While such codes are very efficient to implement, they suffer from inherent weakness against valumetric scaling, such as changes in audio volume or image brightness. In the present paper, we present an alternative to lattice codes that is inherently robust to valumetric scaling. This code is based on a trellis that has been modified so that each bit value may be coded by traversing several alternative arcs. A Viterbi decoder is used in the detector to identify the path with the highest correlation to the input work. Since relative correlation values are unaffected by valumetric scaling, the same message will be detected no matter how the input has been scaled.
Matthew L. Miller, Gwenaël J. Doërr, Ingemar J. Cox
ICIP (2)3
2002 Informed Embedding for Multi-bit Watermarks
Matthew L. Miller, Gwenaël J. Doërr, Ingemar J. Cox
IWDW3
2001 Electronic watermarking: the first 50 years
abstract
Electronic watermarking can be traced back as far as 1954. The 1990s have seen considerable interest in digital watermarking, due in large part to concerns about illegal piracy of copyrighted content. In this paper, we consider the following questions: is the interest warranted? What are the commercial applications of the technology? What scientific progress has been made? What are the most exciting areas for research? And where might the first decade of the new century take us? In our opinion, the interest in watermarking is appropriate. However, we expect that copyright applications will be overshadowed by applications such as broadcast monitoring, authentication, and tracking content distributed within corporations. We further see a variety of applications emerging that add value to media, such as annotation and linking content to the Web. These latter applications may turn out to be the most compelling. Considerable progress has been made toward enabling these applications-perceptual modelling, security threats and countermeasures, and the development of a bag of tricks for efficient implementations. Further progress is needed in methods for handling geometric and temporal distortions. We expect other exciting developments to arise from research in informed watermarking.
Ingemar J. Cox, Matthew L. Miller
MMSP1
2001 Rotation, scale, and translation resilient watermarking for images
abstract
Many electronic watermarks for still images and video content are sensitive to geometric distortions. For example, simple rotation, scaling, and/or translation (RST) of an image can prevent blind detection of a public watermark. In this paper, we propose a watermarking algorithm that is robust to RST distortions. The watermark is embedded into a one-dimensional (1-D) signal obtained by taking the Fourier transform of the image, resampling the Fourier magnitudes into log-polar coordinates, and then summing a function of those magnitudes along the log-radius axis. Rotation of the image results in a cyclical shift of the extracted signal. Scaling of the image results in amplification of the extracted signal, and translation of the image has no effect on the extracted signal. We can therefore compensate for rotation with a simple search, and compensate for scaling by using the correlation coefficient as the detection measure. False positive results on a database of 10,000 images are reported. Robustness results on a database of 2000 images are described. It is shown that the watermark is robust to rotation, scale, and translation. In addition, we describe tests examining the watermarks resistance to cropping and JPEG compression.
Ching-Yung Lin, Min Wu 0001, Jeffrey A. Bloom, Ingemar J. Cox, Matthew L. Miller, Yui Man Lui
IEEE Trans. Image Process.4
2000 Informed Embedding: Exploiting Image and Detector Information During Watermark Insertion
abstract
Usually watermark embedding simply adds a globally or locally attenuated watermark pattern to the cover data (photograph, music, movie). The attenuation is required to maintain fidelity of the cover data to an observer while the watermark detector considers the cover data to be "noise". We refer to this as blind embedding. Cox, Miller and McKellips (see Proceedings of the IEEE, vol.87, no.7, p.1127-41, 1999) observed that the cover data is not noise, i.e. it is not random but completely known at the time of embedding. This knowledge, along with knowledge of the detection algorithm to be used, allows a new category of informed embedder to be realized. We describe a simple watermarking algorithm and then compare the performance of blind embedding with three types of informed embedding. Note that in all four cases, the watermark detector is unchanged, only the embedder is altered. Experimental results clearly reveal the improvement of informed over blind embedding.
Matthew L. Miller, Ingemar J. Cox, Jeffrey A. Bloom
ICIP2
2000 The Bayesian image retrieval system, PicHunter: theory, implementation, and psychophysical experiments
abstract
This paper presents the theory, design principles, implementation and performance results of PicHunter, a prototype content-based image retrieval (CBIR) system. In addition, this document presents the rationale, design and results of psychophysical experiments that were conducted to address some key issues that arose during PicHunter's development. The PicHunter project makes four primary contributions to research on CBIR. First, PicHunter represents a simple instance of a general Bayesian framework which we describe for using relevance feedback to direct a search. With an explicit model of what users would do, given the target image they want, PicHunter uses Bayes's rule to predict the target they want, given their actions. This is done via a probability distribution over possible image targets, rather than by refining a query. Second, an entropy-minimizing display algorithm is described that attempts to maximize the information obtained from a user at each iteration of the search. Third, PicHunter makes use of hidden annotation rather than a possibly inaccurate/inconsistent annotation structure that the user must learn and make queries in. Finally, PicHunter introduces two experimental paradigms to quantitatively evaluate the performance of the system, and psychophysical experiments are presented that support the theoretical claims.
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Thomas V. Papathomas, Peter N. Yianilos
IEEE Trans. Image Process.1
2000 Correction to "the Bayesian image retrieval system, pichunter: theory, implementation, and psychophysical experiments"
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Thomas V. Papathomas, Peter N. Yianilos
IEEE Trans. Image Process.1
1999 A rotation, scale and translation resilient public watermark
abstract
Summary form only given. Watermarking algorithms that are robust to the common geometric transformations of rotation, scale and translation (RST) have been reported for cases in which the original unwatermarked content is available at the detector so as to allow the transformations to be inverted. However, for public watermarks the problem is significantly more difficult since there is no original content to register with. Two classes of solution have been proposed. The first embeds a registration pattern into the content while the second seeks to apply detection methods that are invariant to these geometric transformations. This paper describes a public watermarking method which is invariant (or bares a simple relation) to the common geometric transforms of rotation, scale, and translation. It is based on the Fourier-Mellin transform which has previously been suggested. We extend this work, using a variation based on the Radon transform. The watermark is inserted into a projection of the image. The properties of this projection are such that RST transforms produce simple or no effects on the projection waveform. When a watermark is inserted into a projection, the signal must eventually be back projected to the original image dimensions. This is a one to many mapping that allows for considerable flexibility in the watermark insertion process. We highlight some theoretical and practical issues that affect the implementation of an RST invariant watermark. Finally, we describe preliminary experimental results.
Min Wu 0001, Matthew L. Miller, Jeffrey A. Bloom, Ingemar J. Cox
ICASSP4
1999 Introduction: Computer Vision Research at NECI
Ingemar J. Cox
Int. J. Comput. Vis.1
1999 Copy protection for DVD video
abstract
The prospect of consumer digital versatile disk (DVD) recorders highlights the challenge of protecting copyrighted video content from piracy. We describe the copy-protection system currently under consideration for DVD. The copy-protection system broadly tries to prevent illicit copies from being made from either the analog or digital I/O channels of DVD recorders. An analog copy-protection system is utilized to protect the NTSC/PAL output channel by preventing copies to VHS. The digital transmission of content is protected by a robust encryption protocol between two communicating devices. Watermarking is used to encode copy-control information retrievable from both digital and analog signals. Hence, such embedded signals avoid the need for metadata to be carried in either the digital or analog domains. Finally, the copy-protection system provides the capability for one-generation copying. We discuss some proposed solutions and some of the implementation issues that are being addressed.
Jefferey A. Bloom, Ingemar J. Cox, Ton Kalker, Jean-Paul Linnartz, Matthew L. Miller, C. Brendan S. Traw
Proc. IEEE2
1999 Watermarking as communications with side information
abstract
Several authors have drawn comparison between embedded signaling or watermarking and communications, especially spread-spectrum communications. We examine the similarities and differences between watermarking and traditional communications. This comparison suggests that watermarking most closely resembles communications with side information at the transmitter and or detector, a configuration originally described by Shannon (1958). This leads to several novel characteristics and insights regarding embedded signaling which are discussed in detail.
Ingemar J. Cox, Matthew L. Miller, Andrew L. McKellips
Proc. IEEE1
1998 An Optimized Interaction Strategy for Bayesian Relevance Feedback
abstract
A new algorithm and systematic evaluation is presented for searching a database via relevance feedback. It represents a new image display strategy for the PicHunter system. The algorithm takes feedback in the form of relative judgments ("item A is more relevant than item B") as opposed to the stronger assumption of categorical relevance judgments ("item A is relevant but item B is not"). It also exploits a learned probabilistic model of human behavior to make better use of the feedback it obtains. The algorithm can be viewed as an extension of indexing schemes like the k-d tree to a stochastic setting, hence the name "stochastic-comparison search." In simulations, the amount of feedback required for the new algorithm scales like log/sub 2/ |D|, where |D| is the size of the database, while a simple query-by-example approach scales like |D|/sup /spl alpha//, where /spl alpha/<1 depends on the structure of the database. This theoretical advantage is reflected by experiments with real users on a database of 1500 stock photographs.
Ingemar J. Cox, Matthew L. Miller, Tom Minka, Peter N. Yianilos
CVPR1
1998 A Maximum-Flow Formulation of the N-Camera Stereo Correspondence Problem
abstract
This paper describes a new algorithm for solving the N-camera stereo correspondence problem by transforming it into a maximum-flow problem. Once solved, the minimum-cut associated to the maximum-flow yields a disparity surface for the whole image at once. This global approach to stereo analysis provides a more accurate and coherent depth map than the traditional line-by-line stereo. Moreover, the optimality of the depth surface is guaranteed and can be shown to be a generalization of the dynamic programming approach that is widely used in standard stereo. Results show improved depth estimation as well as better handling of depth discontinuities. While the worst case running time is O(n/sup 2/d/sup 2/log(nd)), the observed average running time is O(n/sup 1.2/ d/sup 1.3/) for an image size of n pixels and depth resolution d.
Sébastien Roy 0001, Ingemar J. Cox
ICCV2
1998 Some general methods for tampering with watermarks
abstract
Watermarks allow embedded signals to be extracted from audio and video content for a variety of purposes. One application is for copyright control, where it is envisaged that digital video recorders will not permit the recording of content that is watermarked as "never copy". In such a scenario, it is important that the watermark survive both normal signal transformations and attempts to remove the watermark so that an illegal copy can be made. We discuss to what extent a watermark can be resistant to tampering and describe a variety of possible attacks.
Ingemar J. Cox, Jean-Paul Linnartz
IEEE J. Sel. Areas Commun.1
1997 Cylindrical rectification to minimize epipolar distortion
abstract
We propose anew rectification method for aligning epipolar lines of a pair of stereo images taken under any camera geometry. It effectively remaps both images onto the surface of a cylinder instead of a plane, which is used in common rectification methods. For a large set of camera motions, remapping to a plane has the drawback of creating rectified images that are potentially infinitely large and presents a loss of pixel information along epipolar lines. In contrast, cylindrical rectification guarantees that the rectified images are bounded for all possible camera motions and minimizes the loss of pixel information along epipolar line. The processes (e.g., stereo matching, etc.) subsequently applied to the rectified images are thus more accurate and general since they can accommodate any camera geometry.
Sébastien Roy 0001, Jean Meunier, Ingemar J. Cox
CVPR3
1997 A multiple-baseline stereo for precise human face acquisition
Shizuo Sakamoto, Ingemar J. Cox, Johji Tajima
Pattern Recognit. Lett.2
1997 Secure spread spectrum watermarking for multimedia
abstract
This paper presents a secure (tamper-resistant) algorithm for watermarking images, and a methodology for digital watermarking that may be generalized to audio, video, and multimedia data. We advocate that a watermark should be constructed as an independent and identically distributed (i.i.d.) Gaussian random vector that is imperceptibly inserted in a spread-spectrum-like fashion into the perceptually most significant spectral components of the data. We argue that insertion of a watermark under this regime makes the watermark robust to signal processing operations (such as lossy compression, filtering, digital-analog and analog-digital conversion, requantization, etc.), and common geometric transformations (such as cropping, scaling, translation, and rotation) provided that the original image is available and that it can be successfully registered against the transformed watermarked image. In these cases, the watermark detector unambiguously identifies the owner. Further, the use of Gaussian noise, ensures strong resilience to multiple-document, or collusional, attacks. Experimental results are provided to support these claims, along with an exposition of pending open problems.
Ingemar J. Cox, Joe Kilian, Frank Thomson Leighton, Talal Shamoon
IEEE Trans. Image Process.1
1996 Feature-Based Face Recognition Using Mixture-Distance
abstract
We consider the problem of feature-based face recognition in the setting where only a single example of each face is available for training. The mixture-distance technique we introduce achieves a recognition rate of 95% on a database of 685 people in which each face is represented by 30 measured distances. This is currently the best recorded recognition rate for a feature-based system applied to a database of this size. By comparison, nearest neighbor search using Euclidean distance yields 84%. In our work a novel distance function is constructed based on local second order statistics as estimated by modeling the training data as a mixture of normal densities. We report on the results from mixtures of several sizes. We demonstrate that a flat mixture of mixtures performs as well as the best model and therefore represents an effective solution to the model selection problem. A mixture perspective is also taken for individual Gaussians to choose between first order (variance) and second order (covariance) models. Here an approximation to flat combination is proposed and seen to perform well in practice. Our results demonstrate that even in the absence of multiple training examples for each class, it is sometimes possible to infer from a statistical model of training data, a significantly improved distance function for use in pattern recognition.
Ingemar J. Cox, Joumana Ghosn, Peter N. Yianilos
CVPR1
1996 Secure spread spectrum watermarking for images, audio and video
abstract
We describe a digital watermarking method for use in audio, image, video and multimedia data. We argue that a watermark must be placed in perceptually significant components of a signal if it is to be robust to common signal distortions and malicious attack. However, it is well known that modification of these components can lead to perceptual degradation of the signal. To avoid this, we propose to insert a watermark into the spectral components of the data using techniques analogous to spread spectrum communications, hiding a narrow band signal in a wideband channel that is the data. The watermark is difficult for an attacker to remove, even when several individuals conspire together with independently watermarked copies of the data. It is also robust to common signal and geometric distortions such as digital-to-analog and analog-to-digital conversion, resampling, quantization, dithering, compression, rotation, translation, cropping and scaling. The same digital watermarking algorithm can be applied to all three media under consideration with only minor modifications, making it especially appropriate for multimedia products. Retrieval of the watermark unambiguously identifies the owner, and the watermark can be constructed to make counterfeiting almost impossible. We present experimental results to support these claims.
Ingemar J. Cox, Joe Kilian, Frank Thomson Leighton, Talal Shamoon
ICIP (3)1
1996 PicHunter: Bayesian relevance feedback for image retrieval
abstract
This paper describes PicHunter, an image retrieval system that implements a novel approach to relevance feedback, such that the entire history of user selections contributes to the system's estimate of the user's goal image. To accomplish this, PicHunter uses Bayesian learning based on a probabilistic model of a user's behavior. The predictions of this model are combined with the selections made during a search to estimate the probability associated with each image. These probabilities are then used to select images for display. Details of our model of a user's behavior were tuned using an off-line leaning algorithm. For clarity, our studies were done with the simplest possible user interface but the algorithm can easily be incorporated into systems which support complex queries, including most previously proposed systems. However, even with this constraint and simple image features, PicHunter is able to locate randomly selected targets in a database of 4522 images after displaying an average of only 55 groups of 4 images which is over 10 times better than chance. We therefore expect that the performance of current image database retrieval systems can be improved by incorporation of the techniques described here.
Ingemar J. Cox, Matthew L. Miller, Stephen M. Omohundro, Peter N. Yianilos
ICPR1
1996 "Ratio regions": a technique for image segmentation
abstract
We develop a image segmentation algorithm in which the segmented region has both an exterior boundary cost and an interior benefit associated with it. Our segmentation method proceeds by minimizing the ratio between the exterior boundary cost and the enclosed interior benefit using a computationally efficient graph partitioning algorithm. Our interest is motivated by very efficient algorithms for finding the globally optimum solution, and a desire to investigate how weak smoothness constraints may be globally imposed without disallowing very high local curvature. We analyze the performance of the approach, indicating both strengths and weaknesses, and discuss its connections with prior image partitioning algorithms. The relationship with snakes is discussed in detail and it is shown how to efficiently compute an approximation to common snakes under the additional constraint that it enclose a given point. When user interaction is available, there is a clear advantage to minimizing user interaction for purposes of improved speed and ease of use and for robustness. "Ratio regions" can accommodate several levels of user interaction and it is empirically shown that very coarse initializations can be tolerated. User interaction not only guides the algorithm to perceptually salient regions but can also be exploited to significantly reduce the computational cost.
Ingemar J. Cox, Satish Rao
ICPR1
1996 Motion without structure
abstract
We propose a new paradigm, motion without structure, for determining the ego-motion between two frames. It is best suited for cases where reliable feature point correspondence is difficult, or for cases where the expected camera motion is large. The problem is posed as a five-dimensional search over the space of possible motions during which the structural information present in the two views is neither implicitly or explicitly used or estimated. To accomplish this search, a cost function is devised that measures the relative likelihood of each hypothesized motion. This cost function is invariant to the structure present in the scene. An analysis of the global scene statistics present in an image, together with the geometry of epipolar misalignment, suggests a measure based on the sum of squared differences between pixels in the first image and their corresponding epipolar line segments in the second image. The measure relies on a simple statistical characteristic of neighboring image intensity levels. Specifically, that the variance of intensity differences between two arbitrary points in an image is a monotonically increasing symmetrical function of the distance between the two points. This assumption is almost always true, though the size of the neighborhood over which the monotonic dependency holds varies from image to image. This range determines the maximum permissible motion between two frames, which can be quite large. Experiments with both outdoor scenes and an indoor calibrated sequence achieve very good accuracy (less then 1 pixel image displacement error) and robustness to noise.
Sébastien Roy 0001, Ingemar J. Cox
ICPR2
1996 A Maximum Likelihood Stereo Algorithm
Ingemar J. Cox, Sunita L. Hingorani, Satish Rao, Bruce M. Maggs
Comput. Vis. Image Underst.1
1996 An Efficient Implementation of Reid's Multiple Hypothesis Tracking Algorithm and Its Evaluation for the Purpose of Visual Tracking
abstract
An efficient implementation of Reid's multiple hypothesis tracking (MHT) algorithm is presented in which the k-best hypotheses are determined in polynomial time using an algorithm due to Murly (1968). The MHT algorithm is then applied to several motion sequences. The MHT capabilities of track initiation, termination, and continuation are demonstrated together with the latter's capability to provide low level support of temporary occlusion of tracks. Between 50 and 150 corner features are simultaneously tracked in the image plane over a sequence of up to 51 frames. Each corner is tracked using a simple linear Kalman filter and any data association uncertainty is resolved by the MHT. Kalman filter parameter estimation is discussed, and experimental results show that the algorithm is robust to errors in the motion model. An investigation of the performance of the algorithm as a function of look-ahead (tree depth) indicates that high accuracy can be obtained for tree depths as shallow as three. Experimental results suggest that a real-time MHT solution to the motion correspondence problem is possible for certain classes of scenes.
Ingemar J. Cox, Sunita L. Hingorani
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 Direct Estimation of Rotation from Two Frames via Epipolar Search
Sébastien Roy 0001, Ingemar J. Cox
CAIP2
1995 Dynamic histogram warping of image pairs for constant image brightness
abstract
The constant image brightness (CIB) assumption assumes that the intensities of corresponding points in two images are equal. This assumption is central to much of computer vision. However, surprisingly little work has been performed to support this assumption, despite the fact the many of algorithms are very sensitive to deviations from CIB. An examination of the images contained in the SRI JISCT stereo database revealed that the constant image brightness assumption is indeed often false. Moreover, the simple additive/multiplicative models of the form I/sub L/=/spl beta/I/sub R/+/spl alpha/ do not adequately represent the observed deviations. A comprehensive physical model of the observed deviations is difficult to develop. However, many potential sources of deviations can be represented by a nonlinear monotonically increasing relationship between intensities. Under these conditions, we believe that an expansion/contraction matching of the intensity histograms represents the best method to both measure the degree of validity of the CIB assumption and correct for it. Dynamic histogram warping (DHW) is closely related to histogram specification. It is shown that histogram specification introduces artifacts that do not occur with dynamic histogram warping. Experimental results show that image histograms are closely matched after DHW, especially when both histograms are modified simultaneously. DHW is also capable of removing simple constant additive and multiplicative biases without derivative operations, thereby avoiding amplification of high frequency noise. DHW can improve the estimates from stereo and optical flow estimators.
Ingemar J. Cox, Sébastien Roy 0001, Sunita L. Hingorani
ICIP1
1995 Underwater Sonar Data Fusion Using Efficient Multiple Hypothesis Algorithm
abstract
This paper describes a geometric approach to underwater environmental modeling using sonar. We classify and localize geometric features of man-made objects by combining the boundary constraints of sonar returns obtained from multiple vantage points. The approach builds on our previous use of Reid's (1979) multiple hypothesis tracking (MHT) algorithm in order to resolve data association and motion correspondence ambiguities thereby to construct a model of the observed environment (Cox and Leonard, 1994). In particular we describe a new, computationally efficient implementation of the MHT algorithm originally reported in (Cox and Miller, 1995) and validate target models previously developed for air sonar. The technique fuses data by modeling the physics of underwater sonar and its interaction with different object features. We illustrate the approach in two dimensions with real acoustic data taken using a high-frequency (1.25 MHz) pencil-beam profiling sonar, manually positioned along trajectories which circumnavigate prismatic objects.
John J. Leonard, Bradley A. Moran, Ingemar J. Cox, Matthew L. Miller
ICRA3
1994 A maximum likelihood N-camera stereo algorithm
abstract
This paper extends results of a maximum likelihood two-frame stereo algorithm to the case of N cameras. The N-camera stereo algorithm determines the "best" set of correspondences between a given pair of cameras, referred to as the principal cameras. Knowledge of the relative positions of the cameras allows the 3D point hypothesized by an assumed correspondence of two features in the principal pair to be projected onto the image plane of the remaining N-2 cameras. These N-2 points are then used to verify proposed matches. Not only does the algorithm explicitly model occlusion between features of the principal pair, but the possibility of occlusions in the N-2 additional views is also modelled. The benefits and importance of this are experimentally verified. Like other multi-frame stereo algorithms, the computational and memory costs of this approach increase linearly with each additional view. Experimental results are shown for two outdoor scenes. It is clearly demonstrated that the number of correspondence errors is significantly reduced as the number of views/cameras is increased.>
Ingemar J. Cox
CVPR1
1994 Recursive tracking of formants in speech signals
abstract
We report on an approach to recursively track parameters of a cascade formant model. The work follows from that of Rigoll (1986) who showed how an extended Kalman filter (EKF) may be used for recursive estimation of formants. The success of this approach depends on our ability to tune the model noise variances properly. The approach also fails when there is a mismatch between the complexity of the data and that of the model (i.e. wrong number of formants). We show how a multiple model (MM) approach may be used to overcome these problems. We run several models in parallel and use the innovation probabilities of the EKF to recursively evaluate the likelihoods of each of the models. Experimental results demonstrate the feasibility of the approach; accurate switching between models and good tracking of the formants is achieved.>
Mahesan Niranjan, Ingemar J. Cox, Sunita L. Hingorani
ICASSP (2)2
1994 Identification of Events from 3D Volumes of Seismic Data
abstract
Describes a method for extracting 3D events from a 3D volume of seismic data for the purposes of geophysical interpretation. An event is the recorded waveform caused by a seismic scatterer. It can be modeled as a hyperboloid with local perturbations. Extraction is difficult because of the local perturbations and, particularly, because events cross one another. The original data is in the form of 2D time sections which are vertical slices through the data volume. Once the event has been extracted, it is represented by fitting two types of surfaces. The first surface is a hyperboloid. The parameters that define the hyperboloid are used to determine geophysical parameters such as the position of the scatterer and the Earth's average propagation velocity. The second is a smooth approximating surface calculated using a regularising term. This surface is useful for visualising regions where the event pulls away from pure hyperbolic form.>
Peter H. Tu, Andrew Zisserman, Ian A. Mason, Ingemar J. Cox
ICIP (3)4
1994 An efficient implementation and evaluation of Reid's multiple hypothesis tracking algorithm for visual tracking
abstract
An efficient implementation of Reid's multiple hypothesis tracking (MHT) algorithm is presented in which the the k-best hypotheses are determined in polynomial time using an algorithm due to Murty (1968). The MHT algorithm is then applied to several motion sequences. The MHT capabilities of track initiation, termination and continuation are demonstrated. Continuation allows the MHT to function despite temporary occlusion of tracks. Between 50 and 150 corner features are simultaneously tracked in the image plane over a sequence of up to 60 frames. Each corner is tracked using a simple linear Kalman filter and any data association uncertainty is resolved by the MHT. Kalman filter parameter estimation is discussed and experimental results show that the algorithm is robust to errors in the motion model.
Ingemar J. Cox, Sunita L. Hingorani
ICPR (1)1
1994 Extraction of events from 3D volumes of seismic data
abstract
This paper describes a method for extracting 3D events from a 3D volume of seismic data for the purposes of geophysical interpretation. An event is the recoded waveform caused by a seismic scatterer. It can be modeled as hyperboloids with local perturbations. Extraction is difficult because of the local perturbations and, particularly, because events cross one another. The original data is in the form of 2D time sections which are vertical slices through the data volume. Constraints based on the physics of the seismic experiment are used to extract 2D event curves from the time sections and then combine them into 3D events. Two types of surfaces are fitted to the 3D event so that meaningful interpretations can be made.
Peter H. Tu, Andrew Zisserman, Ian A. Mason, Ingemar J. Cox
ICPR (3)4
1994 Modeling a Dynamic Environment Using a Bayesian Multiple Hypothesis Approach
Ingemar J. Cox, John J. Leonard
Artif. Intell.1
1993 A review of statistical data association techniques for motion correspondence
Ingemar J. Cox
Int. J. Comput. Vis.1
1993 A Bayesian multiple-hypothesis approach to edge grouping and contour segmentation
Ingemar J. Cox, James M. Rehg, Sunita L. Hingorani
Int. J. Comput. Vis.1
1992 Stereo Without Disparity Gradient Smoothing: A Bayesian Sensor Fusion Solution
Ingemar J. Cox, Sunita L. Hingorani, Bruce M. Maggs, Satish Rao
BMVC1
1992 A Bayesian Multiple Hypothesis Approach to Contour Grouping
Ingemar J. Cox, James M. Rehg, Sunita L. Hingorani
ECCV1
1992 An Analysis of Camera Noise
abstract
The class of cameras that are based on ionization sensors, which includes the most common charge-coupled device (CCD) and vidicon cameras, is examined. Camera signals are shown to be corrupted by direction-dependent stationary electronic noise sources and fluctuations due to the statistical nature of the sensing process. The authors develop and test a model of the inherent noises in cameras. These results are confirmed by measurement, and they suggest a locally stationary model of noise for adaptive signal processing.>
Robert A. Boie, Ingemar J. Cox
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 Blanche-an experiment in guidance and navigation of an autonomous robot vehicle
abstract
The principal components and capabilities of Blanche, an autonomous robot vehicle, are described. Blanche is designed for use in structured office or factory environments rather than unstructured natural environments, and it is assumed that an offline path planner provides the vehicle with a series of collision-free maneuvers, consisting of line and arc segments, to move the vehicle to a desired position. These segments are sent to a low-level trajectory generator and closed-loop motion control. The controller assumes accurate knowledge of the vehicle's position. Blanche's position estimation system consists of a priori map of its environment and a robust matching algorithm. The matching algorithm also estimates the precision of the corresponding match/correction that is then optimally (in a maximum-likelihood sense) combined with the current odometric position to provide an improved estimate of the vehicle's position. The system does not use passive or active beacons. Experimental results are reported.>
Ingemar J. Cox
IEEE Trans. Robotics Autom.1
1990 Line recognition
abstract
The use of edge detection and localizing filters for the recognition of lines (narrow contrasting strips) in visual scenes is shown to produce systematic errors while being suboptimum with respect to noise. The authors derive optimum filters for the detection and localization of lines based on matched filtering in the line normal direction and a Wiener filter in the direction tangential to the line. The matched filters for the detection and localization of line normals in white noise have the form of the system response and its first derivative, respectively. The orthogonal least-restrictive Wiener filter is also closely approximated by the form of the system response in cased where Gaussian response and white noise dominate. Close approximations to the optimum filters for line recognition under these common conditions are readily realized. Line filters were integrated with the Boie-Cox edge detector. Results for images with mixed lines and edges show that the integrated Boie-Cox system is free of the systematic errors common to other edge recognition systems.>
Ingemar J. Cox, Robert A. Boie, Deborah A. Wallach
ICPR (1)1
1990 Predicting and Estimating the Accuracy of a Subpixel Registration Algorithm
abstract
It is shown that an efficient practical registration algorithm previously described by H.G. Barrow et al. (1977) can provide high-accuracy registration. Experiments with a quantized video image of a solid triangle yielded registration that was accurate to 3% of the interpixel spacing (i.e. accurate to 0.5 mils) in the x and y directions and 0.015 degrees in rotation. It may also be important to predict the accuracy in advance, to see whether specifications can be met and to estimate accuracy during registration, in order to control quality. The authors provide practical formulas for both purposes for two kinds of image point data: edge detection (ED) data and direct measurement (DM) data. In two experiments using ED data, the predicted, estimated, and observed accuracies are all in agreement. The prediction theory developed suggests five precautions to avoid loss of registration accuracy. Perhaps most important is Precaution c, the necessity of the ED case not to have a large fraction of the total segment length of the model aligned with the horizontal or vertical directions of the pixel grid. When the model consists largely of horizontal and vertical segments, a good way to observe this precaution is to tilt the pixel grid a few degrees away from perfect alignment, e.g. by tilting the video camera. A third experiment verifies that violating Precaution c can seriously degrade accuracy.>
Ingemar J. Cox, Joseph B. Kruskal, Deborah A. Wallach
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 On The Congruence Of Noisy Images To Line Segment Models
abstract
An algorithm is described for matching two-dimensional images that does not depend on the extraction of features, a step which is often difficult or impractical. It combines computational efficiency with robustness against noisy data. It has been tested with two kinds of data. The algorithm matches a two-dimensional image, in the form of a set of points, to a line segment model. It estimates a positive congruence (i.e., a translation plus a rotation) that moves the image onto the model. The algorithm is most appropriate for matching when the required congruence is small, and it seems very well suited to motion-tracking. Experimental results have been obtained for two kinds of images: (1) points obtained from an optical range finder, and (2) edge points extracted from an intensity image. Only the former results are presented.
Ingemar J. Cox, Joseph B. Kruskal
ICCV1
1988 C++ language support for guaranteed initialization, safe termination and error recovery in robotics
abstract
Software issues related to the reliability of robot systems are considered. It is shown how more reliable robot systems can be built using data abstraction and object-oriented programming, as supported within C++, a general-purpose programming language. It is also shown how the constructor mechanism associated with C++ classes can be used to guarantee initialization and self-test of each subsystem of the robot. A complementary destructor mechanism can be used to guarantee safe termination of subsystems under most conditions. These mechanisms are completely transparent to both the user and his application program. Exception handling is discussed and it is shown how object-oriented programming facilities can be used to provide transparent recovery from subsystem failures during program execution, given some hardware redundancy. Most of the examples have been demonstrated and tested using C++ running on an autonomous robot vehicle.>
Ingemar J. Cox
ICRA1
1988 Blanche: an autonomous robot vehicle for structured environments
abstract
Blanche is an experimental vehicle designed to operate autonomously within a structured office or factory environment. Blanche has been designed with several goals in mind. First and foremost, Blanche is a testbed with which to experiment with such things as robot programming languages, sensor integration/data fusion techniques, i.e. the management of sparse, conflicting, and/or uncertain information, and with techniques for error detection and recovery. Second, Blanche is designed to be low-cost and its dependence on only two sensors, an optical rangefinder and odometry, reflects this. A description is given of the structure of the vehicle and its associated sensors. Experimental results characterizing the accuracy and repeatability of the cart and sensors are presented. A brief description of the obstacle-free trajectory generation and low-level control algorithms are also given.>
Ingemar J. Cox
ICRA1
1988 Local path control for an autonomous vehicle
abstract
A control system for an autonomous robot cart designed to operate in well-structured environments such as offices and factories is described. The onboard navigation system comprises a reference-state generator, an error-feedback controller, and cart-location sensing using odometry. There is a convenient separation between the path guidance and control logic. Under normal operating conditions, the controller ensures that the errors between the measured and reference states are small. These errors only exceed set limits if the cart is malfunctioning. Major hardware failures can be detected in this way and failsafe procedures invoked. Results on the control system performance derived from a computer simulation of the cart and its operating environment, and from an experimental cart, indicate that the system can provide reliable, accurate, and safe operation.>
Winston L. Nelson, Ingemar J. Cox
ICRA2
1987 Concurrent C and robotics
abstract
Many current robot systems exhibit a significant degree of concurrency, doing many activities in parallel. Future sensor-based robots are expected to exhibit even more concurrency. Programs to control such robots are characterized by the need to wait for external events and/or handle interrupts, deal with concurrent activities, synchronize actions with external events and communicate with other robots/processes. In this paper, we focus on the advantages of concurrent programming for robotics and suggest that a general purpose language with the right facilities is a good vehicle for robot programming. In this context we will discuss Concurrent C, an upward-compatible extension of the C language that provides high-level concurrent programming facilities. We give a brief description of Concurrent C followed by a description of how Concurrent C programs communicate with robots and devices. We then show, by means of examples, all of which were implemented, how Concurrent C simplifies the writing of robot programs. Of specific interest are the process interaction and related interrupt handling facilities.
Ingemar J. Cox, Narain H. Gehani
ICRA1
1983 Digital image processing of confocal images
Ingemar J. Cox, C. J. R. Sheppard
Image Vis. Comput.1