Mark J. Carman

dblp:38/3874 · also Mark James Carman · DBLP profile ↗
← Back
55ranked-venue papers
8as first author
3since 2021 · last 2022
0000-0001-6575-9737ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 32 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Information retrieval · 62% Data mining · 35% Web and social media mining · 2%
Artificial intelligence
5 papers
Information extraction and text analysis · 42% Probabilistic and Bayesian machine learning · 25% Language models and text generation · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 67% Medical and health informatics · 33%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › sentiment analysis
sarcasm detection
0.522017
Sarcasm Suite: A Browser-Based Engine for Sarcasm Detection and Generation · AAAI 2017
Are Word Embedding-based Features Useful for Sarcasm Detection? · EMNLP 2016
Collaborative and social computing
crowdsourcing
0.412020
A technical survey on statistical modelling and design methods for crowdsourcing quality control · Artif. Intell. 2020
Medical and health informatics
epidemic modeling
0.312018
SIR-Hawkes: Linking Epidemic Models and Hawkes Processes to Model Diffusions in Finite Populations · WWW 2018
Computational social science and digital humanities
hawkes process
0.312018
SIR-Hawkes: Linking Epidemic Models and Hawkes Processes to Model Diffusions in Finite Populations · WWW 2018
Computational social science and digital humanities › social network analysis
information diffusion
0.312018
SIR-Hawkes: Linking Epidemic Models and Hawkes Processes to Model Diffusions in Finite Populations · WWW 2018
Natural language and speech › Language models and text generation › text generation › figurative language generation
sarcasm generation
0.312017
Sarcasm Suite: A Browser-Based Engine for Sarcasm Detection and Generation · AAAI 2017
Information retrieval › document retrieval
opinion retrieval
0.322012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Proximity-based opinion retrieval · SIGIR 2010
Data mining
anomaly detection
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining
clustering
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining › clustering
distance-based clustering
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Information retrieval › ranking
learning to rank
0.212016
Comparing Pointwise and Listwise Objective Functions for Random-Forest-Based Learning-to-Rank · ACM Trans. Inf. Syst. 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model
logistic regression
0.212014
Naive-Bayes Inspired Effective Pre-Conditioner for Speeding-Up Logistic Regression · ICDM 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network › bayesian network classifiers
naive bayes
0.212013
Alleviating naive Bayes attribute independence assumption by attribute weighting · J. Mach. Learn. Res. 2013
Information retrieval
ranking
0.112012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Information retrieval
retrieval models
0.112012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Information retrieval
similarity search
0.112012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Information retrieval › retrieval models
language model
0.112011
Improving social bookmark search using personalised latent variable language models · WSDM 2011
Information retrieval
personalized search
0.112011
Improving social bookmark search using personalised latent variable language models · WSDM 2011
Information retrieval › document retrieval › opinion retrieval
blog opinion retrieval
0.112010
Proximity-based opinion retrieval · SIGIR 2010
Information retrieval › web search › web information retrieval › social media retrieval
blog distillation
0.112009
Blog distillation using random walks · SIGIR 2009
Information retrieval › web search › web information retrieval › social media retrieval
blog retrieval
0.112009
Blog distillation using random walks · SIGIR 2009
Information retrieval
query log analysis
0.112009
A statistical comparison of tag and query logs · SIGIR 2009
Information retrieval
distributed information retrieval
0.112008
Towards personalized distributed information retrieval · SIGIR 2008
Information retrieval › ranking › learning to rank
NDCG optimization
0.112016
Comparing Pointwise and Listwise Objective Functions for Random-Forest-Based Learning-to-Rank · ACM Trans. Inf. Syst. 2016
Information retrieval
evaluation
0.122012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Towards personalized distributed information retrieval · SIGIR 2008
Web and social media mining
social tagging
0.122011
Improving social bookmark search using personalised latent variable language models · WSDM 2011
A statistical comparison of tag and query logs · SIGIR 2009
Services computing and microservices
web services
0.112005
Learning Source Descriptions for Web Services · AAAI 2005
Information retrieval › evaluation
retrieval effectiveness
0.012012
Aggregation Methods for Proximity-Based Opinion Retrieval · ACM Trans. Inf. Syst. 2012
Data mining › text mining › text classification
tag prediction
0.012009
A statistical comparison of tag and query logs · SIGIR 2009

Methods — techniques the papers use, named apart from their topics

statistical modelling · 0.4stochastic modeling · 0.3hawkes process · 0.3SIR model · 0.3incongruity modeling · 0.3chatbot · 0.3word2vec · 0.2word embeddings · 0.2random forest · 0.2greedy optimization · 0.2glove · 0.2data-dependent dissimilarity measure · 0.2coordinate-wise optimization · 0.2LSA · 0.2stochastic gradient descent · 0.2quasi-newton · 0.2naive bayes · 0.2gradient descent · 0.2
YearPublicationVenuePosition
2022 Investigating Deep Learning Based Breast Cancer Subtyping Using Pan-Cancer and Multi-Omic Data
abstract
Breast Cancer comprises multiple subtypes implicated in prognosis. Existing stratification methods rely on the expression quantification of small gene sets. Next Generation Sequencing promises large amounts of omic data in the next years. In this scenario, we explore the potential of machine learning and, particularly, deep learning for breast cancer subtyping. Due to the paucity of publicly available data, we leverage on pan-cancer and non-cancer data to design semi-supervised settings. We make use of multi-omic data, including microRNA expressions and copy number alterations, and we provide an in-depth investigation of several supervised and semi-supervised architectures. Obtained accuracy results show simpler models to perform at least as well as the deep semi-supervised approaches on our task over gene expression data. When multi-omic data types are combined together, performance of deep models shows little (if any) improvement in accuracy, indicating the need for further analysis on larger datasets of multi-omic data as and when they become available. From a biological perspective, our linear model mostly confirms known gene-subtype annotations. Conversely, deep approaches model non-linear relationships, which is reflected in a more varied and still unexplored set of representative omic features that may prove useful for breast cancer subtyping.
Francisco Cristovao, Silvia Cascianelli, Arif Canakoglu, Mark J. Carman, Luca Nanni, Pietro Pinoli, Marco Masseroli
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Perception Visualization: Seeing Through The Eyes Of a DNN
Loris Giulivi, Mark J. Carman, Giacomo Boracchi
BMVC2
2021 CDF Transform-and-Shift: An effective way to deal with datasets of inhomogeneous cluster densities
Ye Zhu 0002, Kai Ming Ting, Mark J. Carman, Maia Angelova
Pattern Recognit.3
2020 A technical survey on statistical modelling and design methods for crowdsourcing quality control
Mark J. Carman, Ye Zhu 0002, Yong Xiang 0001
Artif. Intell.2
2020 Self-labeling methods for unsupervised transfer ranking
Mark Sanderson, Mark J. Carman, Falk Scholer
Inf. Sci.3
2019 OCR On-the-Go: Robust End-to-End Systems for Reading License Plates & Street Signs
abstract
We work on the problem of recognizing license plates and street signs automatically in challenging conditions such as chaotic traffic. We leverage state-of-the-art text spotters to generate a large amount of noisy labeled training data. The data is filtered using a pattern derived from domain knowledge. We augment training and testing data with interpolated boxes and annotations that makes our training and testing robust. We further use synthetic data during training to increase the coverage of the training data. We train two different models for recognition. Our baseline is a conventional Convolution Neural Network (CNN) encoder followed by a Recurrent Neural Network (RNN) decoder. As our first contribution, we bypass the detection phase by augmenting the baseline with an Attention mechanism in the RNN decoder. Next, we build in the capability of training the model end-to-end on scenes containing license plates by incorporating inception based CNN encoder that makes the model robust to multiple scales. We achieve improvements of 3.75% at the sequence level, over the baseline model. We present the first results of using multi-headed attention models on text recognition in images and illustrate the advantages of using multiple-heads over a single head. We observe gains as large as 7.18% by incorporating multi-headed attention. We also experiment with multi-headed attention models on French Street Name Signs dataset (FSNS) and a new Indian Street dataset that we release for experiments. We observe that such models with multiple attention masks perform better than the model with single-headed attention on three different datasets with varying complexities. Our models outperform state-of-the-art methods on FSNS and IIIT-ILST Devanagari datasets by 1.1% and 8.19% respectively.
Rohit Saluja, Ayush Maheshwari, Ganesh Ramakrishnan, Parag Chaudhuri, Mark J. Carman
ICDAR5
2019 Sub-Word Embeddings for OCR Corrections in Highly Fusional Indic Languages
abstract
Texts in Indic Languages contain a large proportion of out-of-vocabulary (OOV) words due to frequent fusion using conjoining rules (of which there are around 4000 in Sanskrit). OCR errors further accentuate this complexity for the error correction systems. Variations of sub-word units such as n-grams, possibly encapsulating the context, can be extracted from the OCR text as well as the language text individually. Some of the sub-word units that are derived from the texts in such languages highly correlate to the word conjoining rules. Signals such as frequency values (on a corpus) associated with such sub-word units have been used previously with log-linear classifiers for detecting errors in Indic OCR texts. We explore two different encodings to capture such signals and augment the input to Long Short Term Memory (LSTM) based OCR correction models, that have proven useful in the past for jointly learning the language as well as OCR-specific confusions. The first type of encoding makes direct use of sub-word unit frequency values, derived from the training data. The formulation results in faster convergence and better accuracy values of the error correction model on four different languages with varying complexities. The second type of encoding makes use of trainable sub-word embeddings. We introduce a new procedure for training fastText embeddings on the sub-word units and further observe a large gain in F-Scores, as well as word-level accuracy values.
Rohit Saluja, Mayur Punjabi, Mark J. Carman, Ganesh Ramakrishnan, Parag Chaudhuri
ICDAR3
2019 Lowest probability mass neighbour algorithms: relaxing the metric constraint in distance-based neighbourhood algorithms
Kai Ming Ting, Ye Zhu 0002, Mark J. Carman, Yue Zhu 0001, Takashi Washio, Zhi-Hua Zhou
Mach. Learn.3
2018 Distinguishing Question Subjectivity from Difficulty for Improved Crowdsourcing
abstract
The questions in a crowdsourcing task typically exhibit varying degrees of difficulty and subjectivity. Their joint effects give rise to the variation in responses to the same question by different crowd-workers. This variation is low when the question is easy to answer and objective, and high when it is difficult and subjective. Unfortunately, current quality control methods for crowdsourcing consider only the question difficulty to account for the variation. As a result, these methods cannot distinguish workers' ,personal preferences for different correct answers of a partially subjective question from their ability to avoid objectively incorrect answers for that question. To address this issue, we present a probabilistic model which (i) explicitly encodes question difficulty as a model parameter and (ii) implicitly encodes question subjectivity via latent preference factors for crowd-workers. We show that question subjectivity induces grouping of crowd-workers, revealed through clustering of their latent preferences. Moreover, we develop a quantitative measure for the question subjectivity. Experiments show that our model (1) improves both the question true answer prediction and the unseen worker response prediction, and (2) can potentially provide rankings of questions coherent with human assessment in terms of difficulty and subjectivity.
Mark J. Carman, Ye Zhu 0002, Wray L. Buntine
ACML2
2018 Sarcasm Target Identification: Dataset and An Introductory Approach
Aditya Joshi 0001, Pranav Goel 0001, Pushpak Bhattacharyya, Mark J. Carman
LREC4
2018 Leveraging Label Category Relationships in Multi-class Crowdsourcing
Lan Du 0002, Ye Zhu 0002, Mark J. Carman
PAKDD (2)4
2018 SIR-Hawkes: Linking Epidemic Models and Hawkes Processes to Model Diffusions in Finite Populations
abstract
Among the statistical tools for online information diffusion modeling, both epidemic models and Hawkes point processes are popular choices. The former originate from epidemiology, and consider information as a viral contagion which spreads into a population of online users. The latter have roots in geophysics and finance, view individual actions as discrete events in continuous time, and modulate the rate of events according to the self-exciting nature of event sequences. Here, we establish a novel connection between these two frameworks. Namely, the rate of events in an extended Hawkes model is identical to the rate of new infections in the Susceptible-Infected-Recovered (SIR) model after marginalizing out recovery events -- which are unobserved in a Hawkes process. This result paves the way to apply tools developed for SIR to Hawkes, and vice versa. It also leads to HawkesN, a generalization of the Hawkes model which accounts for a finite population size. Finally, we derive the distribution of cascade sizes for HawkesN, inspired by methods in stochastic SIR. Such distributions provide nuanced explanations to the general unpredictability of popularity: the distribution for diffusion cascade sizes tends to have two modes, one corresponding to large cascade sizes and another one around zero.
Marian-Andrei Rizoiu, Swapnil Mishra, Quyu Kong, Mark J. Carman, Lexing Xie
WWW4
2018 Grouping points by shared subspaces for effective subspace clustering
Ye Zhu 0002, Kai Ming Ting, Mark J. Carman
Pattern Recognit.3
2017 Sarcasm Suite: A Browser-Based Engine for Sarcasm Detection and Generation
abstract
Sarcasm Suite is a browser-based engine that deploys five of our past papers in sarcasm detection and generation. The sarcasm detection modules use four kinds of incongruity: sentiment incongruity, semantic incongruity, historical context incongruity and conversational context incongruity. The sarcasm generation module is a chatbot that responds sarcastically to user input. With a visually appealing interface that indicates predictions using `faces' of our co-authors from our past papers, Sarcasm Suite is our first demonstration of our work in computational sarcasm.
Aditya Joshi 0001, Diptesh Kanojia, Pushpak Bhattacharyya, Mark J. Carman
AAAI4
2017 Using Knowledge Graphs to Explain Entity Co-occurrence in Twitter
abstract
Modern Knowledge Graphs such as DBPedia contain significant information regarding Named Entities and the logical relationships which exist between them. Twitter on the other hand, contains important information on the popularity and frequency with which these entities are mentioned and discussed in combination with one another. In this paper we investigate whether these two sources of information can be used to complement and explain one another. In particular, we would like to know whether the logical relationships (a.k.a. semantic paths) which exist between pairs of known entities can help to explain the frequency with which those entities co-occur with one another in Twitter. To do this we train a ranking function over semantic paths between pairs of entities. The aim of the ranker is to identify the path that most likely explains why a particular pair of entities have appeared together in a particular tweet. We train the ranking model using a number of lexical, graph-embedding and popularity-based features over semantic paths containing a single intermediate entity and demonstrate the efficacy of the model for determining why pairs of entities occur together in tweets.
Yiwei Wang 0001, Mark J. Carman, Yuan-Fang Li
CIKM2
2017 Efficient Benchmarking of NLP APIs using Multi-armed Bandits
abstract
Comparing NLP systems to select the best one for a task of interest, such as named entity recognition, is critical for practitioners and researchers.A rigorous approach involves setting up a hypothesis testing scenario using the performance of the systems on query documents.However, often the hypothesis testing approach needs to send a large number of document queries to the systems, which can be problematic.In this paper, we present an effective alternative based on the multi-armed bandit (MAB).We propose a hierarchical generative model to represent the uncertainty in the performance measures of the competing systems, to be used by Thompson Sampling to solve the resulting MAB.Experimental results on both synthetic and real data show that our approach requires significantly fewer queries compared to the standard benchmarking technique to identify the best system according to Fmeasure.
Gholamreza Haffari, Tuan Dung Tran, Mark J. Carman
EACL (1)3
2017 Leveraging Side Information to Improve Label Quality Control in Crowd-Sourcing
abstract
We investigate the possibility of leveraging side information for improving quality control over crowd-sourced data. We extend the GLAD model, which governs the probability of correct labeling through a logistic function in which worker expertise counteracts item difficulty, by systematically encod- ing different types of side information, including worker in- formation drawn from demographics and personality traits, item information drawn from item genres and content, and contextual information drawn from worker responses and la- beling sessions. Modeling side information allows for better estimation of worker expertise and item difficulty in sparse data situations and accounts for worker biases, leading to bet- ter prediction of posterior true label probabilities. We demon- strate the efficacy of the proposed framework with overall improvements in both the true label prediction and the un- seen worker response prediction based on different combina- tions of the various types of side information across three new crowd-sourcing datasets. In addition, we show the framework exhibits potential of identifying salient side information fea- tures for predicting the correctness of responses without the need of knowing any true label information.
Mark J. Carman, Dongwoo Kim 0002, Lexing Xie
HCOMP2
2017 Error Detection and Corrections in Indic OCR Using LSTMs
abstract
Conventional approaches to spell checking suggest spelling corrections using proximity-based matches to a known vocabulary. For highly inflectional Indian languages, any off-the-shelf vocabulary is significantly incomplete, since a large fraction of words in Indic documents are generated using word conjoining rules. Therefore, a tremendous manual effort is needed in spell-correcting words in Indic OCR documents. Moreover, in a spell checking system, a vocabulary may suggest multiple alternatives to the incorrect word. The ranking of these corrective suggestions is improved using language models. Owing to corpus resource scarcity, however, Indian languages lack reliable language models. Thus, learning the character (or n-gram) confusions or error patterns of the OCR system can be helpful in correcting the Out of Vocabulary (OOV) words in OCR documents. We adopt a Long Short-Term Memory (LSTM) based character level language model with a fixed delay for discriminative language modeling in the context of OCR errors for jointly addressing the problems of error detection and correction in Indic OCR. For words that need not be corrected in the OCR output, our model simply abstains from suggesting any changes. We present extensive results to validate the performance of our model on four Indian languages with different inflectional complexities. We achieve F-Scores above 92.4% and decreases in Word Error Rates (WER) of at least 26.7% across the four languages.
Rohit Saluja, Devaraj Adiga, Parag Chaudhuri, Ganesh Ramakrishnan, Mark J. Carman
ICDAR5
2017 Multi-domain evaluation framework for named entity recognition tools
Zahraa Said Abdallah, Mark J. Carman, Gholamreza Haffari
Comput. Speech Lang.2
2017 Efficient parameter learning of Bayesian network classifiers
abstract
Recent advances have demonstrated substantial benefits from learning with both generative and discriminative parameters. On the one hand, generative approaches address the estimation of the parameters of the joint distribution— $$\mathrm{P}(y,\mathbf{x})$$ , which for most network types is very computationally efficient (a notable exception to this are Markov networks) and on the other hand, discriminative approaches address the estimation of the parameters of the posterior distribution—and, are more effective for classification, since they fit $$\mathrm{P}(y|\mathbf{x})$$ directly. However, discriminative approaches are less computationally efficient as the normalization factor in the conditional log-likelihood precludes the derivation of closed-form estimation of parameters. This paper introduces a new discriminative parameter learning method for Bayesian network classifiers that combines in an elegant fashion parameters learned using both generative and discriminative methods. The proposed method is discriminative in nature, but uses estimates of generative probabilities to speed-up the optimization process. A second contribution is to propose a simple framework to characterize the parameter learning task for Bayesian network classifiers. We conduct an extensive set of experiments on 72 standard datasets and demonstrate that our proposed discriminative parameterization provides an efficient alternative to other state-of-the-art parameterizations.
Nayyar Abbas Zaidi, Geoffrey I. Webb, Mark J. Carman, François Petitjean, Wray L. Buntine, Mike Hynes, Hans De Sterck
Mach. Learn.3
2016 Beyond Clustering: Sub-DAG Discovery for Categorising Documents
abstract
We study the problem of generating DAG-structured category hierarchies over a given set of documents associated with "importance" scores. Example application includes automatically generating Wikipedia disambiguation pages for a set of articles having click counts associated with them. Unlike previous works, which focus on clustering the set of documents using the category hierarchy as features, we directly pose the problem as that of finding a DAG structured generative mode that has maximum likelihood of generating the observed "importance" scores for each document where documents are modeled as the leaf nodes in the DAG structure. Desirable properties of the categories in the inferred DAG-structured hierarchy include document coverage and category relevance, each of which, we show, is naturally modeled by our generative model. We propose two different algorithms for estimating the model parameters. One by modeling the DAG as a Bayesian Network and estimating its parameters via Gibbs Sampling; and the other by estimating the path probabilities using the Expectation Maximization algorithm. We empirically evaluate our method on the problem of automatically generating Wikipedia disambiguation pages using human generated clusterings as the ground truth. We find that our framework improves upon the baselines according to the F1 score and Entropy that are used as standard metrics to evaluate the hierarchical clustering.
Ramakrishna Bairi, Mark J. Carman, Ganesh Ramakrishnan
CIKM2
2016 On the Effectiveness of Query Weighting for Adapting Rank Learners to New Unlabelled Collections
abstract
Query-level instance weighting is a technique for unsupervised transfer ranking, which aims to train a ranker on a source collection so that it also performs effectively on a target collection, even if no judgement information exists for the latter. Past work has shown that this approach can be used to significantly improve effectiveness; in this work, the approach is re-examined on a wide set of publicly available L2R test collections with more advanced learning to rank algorithms. Different query-level weighting strategies are examined against two transfer ranking frameworks: AdaRank and a new weighted LambdaMART algorithm. Our experimental results show that the effectiveness of different weighting strategies, including those shown in past work, vary under different transferring environments. In particular, (i) Kullback-Leibler based density-ratio estimation tends to outperform a classification-based approach and (ii) aggregating document-level weights into query-level weights is likely superior to direct estimation using a query-level representation. The Nemenyi statistical test, applied across multiple datasets, indicates that most weighting transfer learning methods do not significantly outperform baselines, although there is potential for the further development of such techniques.
Mark Sanderson, Mark J. Carman, Falk Scholer
CIKM3
2016 Harnessing Sequence Labeling for Sarcasm Detection in Dialogue from TV Series 'Friends'
abstract
This paper is a novel study that views sarcasm detection in dialogue as a sequence labeling task, where a dialogue is made up of a sequence of utterances.We create a manuallylabeled dataset of dialogue from TV series 'Friends' annotated with sarcasm.Our goal is to predict sarcasm in each utterance, using sequential nature of a scene.We show performance gain using sequence labeling as compared to classification-based approaches.Our experiments are based on three sets of features, one is derived from information in our dataset, the other two are from past works.Two sequence labeling algorithms (SVM-HMM and SEARN) outperform three classification algorithms (SVM, Naive Bayes) for all these feature sets, with an increase in F-score of around 4%.Our observations highlight the viability of sequence labeling techniques for sarcasm detection of dialogue.
Aditya Joshi 0001, Vaibhav Tripathi, Pushpak Bhattacharyya, Mark J. Carman
CoNLL4
2016 Are Word Embedding-based Features Useful for Sarcasm Detection?
abstract
This paper makes a simple increment to state-of-the-art in sarcasm detection research. Existing approaches are unable to capture subtle forms of context incongruity which lies at the heart of sarcasm. We explore if prior work can be enhanced using semantic similarity/discordance between word embeddings. We augment word embedding-based features to four feature sets reported in the past. We also experiment with four types of word embeddings. We observe an improvement in sarcasm detection, irrespective of the word embedding used or the original feature set to which our features are augmented. For example, this augmentation results in an improvement in F-score of around 4\% for three out of these four feature sets, and a minor degradation in case of the fourth, when Word2Vec embeddings are used. Finally, a comparison of the four embeddings shows that Word2Vec and dependency weight-based features outperform LSA and GloVe, in terms of their benefit to sarcasm detection.
Aditya Joshi 0001, Vaibhav Tripathi, Kevin Patel, Pushpak Bhattacharyya, Mark J. Carman
EMNLP5
2016 Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure
abstract
This paper introduces the first generic version of data dependent dissimilarity and shows that it provides a better closest match than distance measures for three existing algorithms in clustering, anomaly detection and multi-label classification. For each algorithm, we show that by simply replacing the distance measure with the data dependent dissimilarity measure, it overcomes a key weakness of the otherwise unchanged algorithm.
Kai Ming Ting, Ye Zhu 0002, Mark J. Carman, Yue Zhu 0001, Zhi-Hua Zhou
KDD3
2016 That'll Do Fine!: A Coarse Lexical Resource for English-Hindi MT, Using Polylingual Topic Models
Diptesh Kanojia, Aditya Joshi 0001, Pushpak Bhattacharyya, Mark J. Carman
LREC4
2016 Tinder Me Softly - How Safe Are You Really on Tinder?
Mark J. Carman, Kim-Kwang Raymond Choo
SecureComm1
2016 ALRn: accelerated higher-order logistic regression
abstract
This paper introduces Accelerated Logistic Regression : a hybrid generative-discriminative approach to training Logistic Regression with high-order features. We present two main results: (1) that our combined generative-discriminative approach significantly improves the efficiency of Logistic Regression and (2) that incorporating higher order features (i.e. features that are the Cartesian products of the original features) reduces the bias of Logistic Regression, which in turn significantly reduces its error on large datasets. We assess the efficacy of Accelerated Logistic Regression by conducting an extensive set of experiments on 75 standard datasets. We demonstrate its competitiveness, particularly on large datasets, by comparing against state-of-the-art classifiers including Random Forest and Averaged n -Dependence Estimators.
Nayyar Abbas Zaidi, Geoffrey I. Webb, Mark J. Carman, François Petitjean, Jesús Cerquides
Mach. Learn.3
2016 Density-ratio based clustering for discovering clusters with varying densities
Ye Zhu 0002, Kai Ming Ting, Mark J. Carman
Pattern Recognit.3
2016 Comparing Pointwise and Listwise Objective Functions for Random-Forest-Based Learning-to-Rank
abstract
Current random-forest (RF)-based learning-to-rank (LtR) algorithms use a classification or regression framework to solve the ranking problem in a pointwise manner. The success of this simple yet effective approach coupled with the inherent parallelizability of the learning algorithm makes it a strong candidate for widespread adoption. In this article, we aim to better understand the effectiveness of RF-based rank-learning algorithms with a focus on the comparison between pointwise and listwise approaches. We introduce what we believe to be the first listwise version of an RF-based LtR algorithm. The algorithm directly optimizes an information retrieval metric of choice (in our case, NDCG) in a greedy manner. Direct optimization of the listwise objective functions is computationally prohibitive for most learning algorithms, but possible in RF since each tree maximizes the objective in a coordinate-wise fashion. Computational complexity of the listwise approach is higher than the pointwise counterpart; hence for larger datasets, we design a hybrid algorithm that combines a listwise objective in the early stages of tree construction and a pointwise objective in the latter stages. We also study the effect of the discount function of NDCG on the listwise algorithm. Experimental results on several publicly available LtR datasets reveal that the listwise/hybrid algorithm outperforms the pointwise approach on the majority (but not all) of the datasets. We then investigate several aspects of the two algorithms to better understand the inevitable performance tradeoffs. The aspects include examining an RF-based unsupervised LtR algorithm and comparing individual tree strength. Finally, we compare the the investigated RF-based algorithms with several other LtR algorithms.
Muhammad Ibrahim 0006, Mark J. Carman
ACM Trans. Inf. Syst.2
2014 Naive-Bayes Inspired Effective Pre-Conditioner for Speeding-Up Logistic Regression
abstract
We propose an alternative parameterization of Logistic Regression (LR) for the categorical data, multi-class setting. LR optimizes the conditional log-likelihood over the training data and is based on an iterative optimization procedure to tune this objective function. The optimization procedure employed may be sensitive to scale and hence an effective pre-conditioning method is recommended. Many problems in machine learning involve arbitrary scales or categorical data (where simple standardization of features is not applicable). The problem can be alleviated by using optimization routines that are invariant to scale such as (second-order) Newton methods. However, computing and inverting the Hessian is a costly procedure and not feasible for big data. Thus one must often rely on first-order methods such as gradient descent (GD), stochastic gradient descent (SGD) or approximate second-order such as quasi-Newton (QN) routines, which are not invariant to scale. This paper proposes a simple yet effective pre-conditioner for speeding-up LR based on naive Bayes conditional probability estimates. The idea is to scale each attribute by the log of the conditional probability of that attribute given the class. This formulation substantially speeds-up LR's convergence. It also provides a weighted naive Bayes formulation which yields an effective framework for hybrid generative-discriminative classification.
Nayyar Abbas Zaidi, Mark J. Carman, Jesús Cerquides, Geoffrey I. Webb
ICDM2
2013 Building user profiles from topic models for personalised search
abstract
Personalisation is an important area in the field of IR that attempts to adapt ranking algorithms so that the results returned are tuned towards the searcher's interests. In this work we use query logs to build personalised ranking models in which user profiles are constructed based on the representation of clicked documents over a topic space. Instead of employing a human-generated ontology, we use novel latent topic models to determine these topics. Our experiments show that by subtly introducing user profiles as part of the ranking algorithm, rather than by re-ranking an existing list, we can provide personalised ranked lists of documents which improve significantly over a non-personalised baseline. Further examination shows that the performance of the personalised system is particularly good in cases where prior knowledge of the search query is limited.
Morgan Harvey, Fabio Crestani, Mark J. Carman
CIKM3
2013 Alleviating naive Bayes attribute independence assumption by attribute weighting
Nayyar Abbas Zaidi, Jesús Cerquides, Mark J. Carman, Geoffrey I. Webb
J. Mach. Learn. Res.3
2012 Comparing Tweets and Tags for URLs
Morgan Harvey, Mark J. Carman, David Elsweiler
ECIR2
2012 Employing document dependency in blog search
abstract
Abstract The goal in blog search is to rank blogs according to their recurrent relevance to the topic of the query. State‐of‐the‐art approaches view it as an expert search or resource selection problem. We investigate the effect of content‐based similarity between posts on the performance of the retrieval system. We test two different approaches for smoothing (regularizing) relevance scores of posts based on their dependencies. In the first approach, we smooth term distributions describing posts by performing a random walk over a document‐term graph in which similar posts are highly connected. In the second, we directly smooth scores for posts using a regularization framework that aims to minimize the discrepancy between scores for similar documents. We then extend these approaches to consider the time interval between the posts in smoothing the scores. The idea is that if two posts are temporally close, then they are good sources for smoothing each other's relevance scores. We compare these methods with the state‐of‐the‐art approaches in blog search that employ Language Modeling‐based resource selection algorithms and fusion‐based methods for aggregating post relevance scores. We show performance gains over the baseline techniques which do not take advantage of the relation between posts for smoothing relevance estimates.
Mostafa Keikha, Fabio Crestani, Mark J. Carman
J. Assoc. Inf. Sci. Technol.3
2012 Aggregation Methods for Proximity-Based Opinion Retrieval
abstract
The enormous amount of user-generated data available on the Web provides a great opportunity to understand, analyze, and exploit people’s opinions on different topics. Traditional Information Retrieval methods consider the relevance of documents to a topic but are unable to differentiate between subjective and objective documents. Opinion retrieval is a retrieval task in which not only the relevance of a document to the topic is important but also the amount of opinion expressed in the document about the topic. In this article, we address the blog post opinion retrieval task and propose methods that rank blog posts according to their relevance and opinionatedness toward a topic. We propose estimating the opinion density at each position in a document using a general opinion lexicon and kernel density functions. We propose and investigate different models for aggregating the opinion density at query terms positions to estimate the opinion score of every document. We then combine the opinion score with the relevance score based on a probabilistic justification. Experimental results on the BLOG06 dataset show that the proposed method provides significant improvement over the standard TREC baselines. The proposed models also achieve much higher performance compared to all state of the art methods.
Shima Gerani, Mark J. Carman, Fabio Crestani
ACM Trans. Inf. Syst.2
2011 Bayesian latent variable models for collaborative item rating prediction
abstract
Collaborative filtering systems based on ratings make it easier for users to find content of interest on the Web and as such they constitute an area of much research. In this paper we first present a Bayesian latent variable model for rating prediction that models ratings over each user's latent interests and also each item's latent topics. We describe a Gibbs sampling procedure that can be used to estimate its parameters and show by experiment that it is competitive with the gradient descent SVD methods commonly used in state-of-the-art systems. We then proceed to make an important and novel extension to this model, enhancing it with user-dependent and item-dependant biases to significantly improve rating estimation. We show by experiment on a large set of real ratings data that these models are able to outperform 3 common baselines, including a very competitive and modern SVD-based model. Furthermore we illustrate other advantages of our approach beyond simply its ability to provide more accurate ratings and show that it is able to perform better on the common and important case where the user profile is short.
Morgan Harvey, Mark J. Carman, Ian Ruthven, Fabio Crestani
CIKM2
2011 Personal Blog Retrieval Using Opinion Features
Shima Gerani, Mostafa Keikha, Mark J. Carman, Fabio Crestani
ECIR3
2011 Investigating the Statistical Properties of User-Generated Documents
Giacomo Inches, Mark J. Carman, Fabio Crestani
FQAS2
2011 Improving social bookmark search using personalised latent variable language models
abstract
Social tagging systems have recently become very popular as a method of categorising information online and have been used to annotate a wide range of different resources. In such systems users are free to choose whatever keywords or "tags" they wish to annotate each resource, resulting in a highly personalised, unrestricted vocabulary. While this freedom of choice has several notable advantages, it does come at the cost of making searching of these systems more difficult as the vocabulary problem introduced is more pronounced than in a normal information retrieval setting.
Morgan Harvey, Ian Ruthven, Mark J. Carman
WSDM3
2011 A multi-collection latent topic model for federated search
Mark Baillie, Mark J. Carman, Fabio Crestani
Inf. Retr.2
2010 Towards query log based personalization using topic models
abstract
We investigate the utility of topic models for the task of personalizing search results based on information present in a large query log. We define generative models that take both the user and the clicked document into account when estimating the probability of query terms. These models can then be used to rank documents by their likelihood given a particular query and user pair.
Mark J. Carman, Fabio Crestani, Morgan Harvey, Mark Baillie
CIKM1
2010 Ranking social bookmarks using topic models
abstract
Ranking of resources in social tagging systems is a difficult problem due to the inherent sparsity of the data and the vocabulary problems introduced by having a completely unrestricted lexicon. In this paper we propose to use hidden topic models as a principled way of reducing the dimensionality of this data to provide more accurate resource rankings with higher recall. We first describe Latent Dirichlet Allocation (LDA) and then show how it can be used to rank resources in a social bookmarking system. We test the LDA tagging model and compare it with 3 non-topic model baselines on a large data sample obtained from the Delicious social bookmarking site. Our evaluations show that our LDA-based method significantly outperforms all of the baselines.
Morgan Harvey, Ian Ruthven, Mark J. Carman
CIKM3
2010 Tripartite Hidden Topic Models for Personalised Tag Suggestion
Morgan Harvey, Mark Baillie, Ian Ruthven, Mark J. Carman
ECIR4
2010 Statistics of Online User-Generated Short Documents
Giacomo Inches, Mark J. Carman, Fabio Crestani
ECIR2
2010 Proximity-based opinion retrieval
abstract
Blog post opinion retrieval aims at finding blog posts that are relevant and opinionated about a user's query. In this paper we propose a simple probabilistic model for assigning relevant opinion scores to documents. The key problem is how to capture opinion expressions in the document, that are related to the query topic. Current solutions enrich general opinion lexicons by finding query-specific opinion lexicons using pseudo-relevance feedback on external corpora or the collection itself. In this paper we use a general opinion lexicon and propose using proximity information in order to capture opinion term relatedness to the query. We propose a proximity-based opinion propagation method to calculate the opinion density at each point in a document. The opinion density at the position of a query term in the document can then be considered as the probability of opinion about the query term at that position. The effect of different kernels for capturing the proximity is also discussed. Experimental results on the BLOG06 dataset show that the proposed method provides significant improvement over standard TREC baselines and achieves a 2.5% increase in MAP over the best performing run in the TREC 2008 blog track.
Shima Gerani, Mark J. Carman, Fabio Crestani
SIGIR2
2009 A Topic-Based Measure of Resource Description Quality for Distributed Information Retrieval
Mark Baillie, Mark J. Carman, Fabio Crestani
ECIR2
2009 Investigating Learning Approaches for Blog Post Opinion Retrieval
Shima Gerani, Mark J. Carman, Fabio Crestani
ECIR2
2009 A statistical comparison of tag and query logs
abstract
We investigate tag and query logs to see if the terms people use to annotate websites are similar to the ones they use to query for them. Over a set of URLs, we compare the distribution of tags used to annotate each URL with the distribution of query terms for clicks on the same URL. Understanding the relationship between the distributions is important to determine how useful tag data may be for improving search results and conversely, query data for improving tag prediction. In our study, we compare both term frequency distributions using vocabulary overlap and relative entropy. We also test statistically whether the term counts come from the same underlying distribution. Our results indicate that the vocabulary used for tagging and searching for content are similar but not identical. We further investigate the content of the websites to see which of the two distributions (tag or query) is most similar to the content of the annotated/searched URL. Finally, we analyze the similarity for different categories of URLs in our sample to see if the similarity between distributions is dependent on the topic of the website or the popularity of the URL.
Mark J. Carman, Mark Baillie, Robert Gwadera, Fabio Crestani
SIGIR1
2009 Blog distillation using random walks
abstract
This paper addresses the blog distillation problem. That is, given a user query find the blogs most related to the query topic. We model the blogosphere as a single graph that includes extra information besides the content of the posts. By performing a random walk on this graph we extract most relevant blogs for each query. Our experiments on the TREC'07 data set show 15% improvement in MAP and 8% improvement in [email protected] over the Language Modeling baseline.
Mostafa Keikha, Mark J. Carman, Fabio Crestani
SIGIR2
2008 Towards personalized distributed information retrieval
abstract
Our aim is to investigate if and how the performance of Distributed Information Retrieval (DIR) systems can be improved through personalization. Toward this aim we are building a testbed of document collections and corresponding personalized relevance judgments. In this paper we discuss our intended approach for personalizing the three different phases of the DIR process. We also describe the test collection we are building and discuss our methodology for evaluating personalized DIR using relevance information taken from social bookmarking data.
Mark J. Carman, Fabio Crestani
SIGIR1
2007 Learning Semantic Descriptions of Web Information Sources
Mark J. Carman, Craig A. Knoblock
IJCAI1
2007 Learning Semantic Definitions of Online Information Sources
abstract
The Internet contains a very large number of information sources providing many types of data from weather forecasts to travel deals and financial information. These sources can be accessed via Web-forms, Web Services, RSS feeds and so on. In order to make automated use of these sources, we need to model them semantically, but writing semantic descriptions for Web Services is both tedious and error prone. In this paper we investigate the problem of automatically generating such models. We introduce a framework for learning Datalog definitions of Web sources. In order to learn these definitions, our system actively invokes the sources and compares the data they produce with that of known sources of information. It then performs an inductive logic search through the space of plausible source definitions in order to learn the best possible semantic model for each new source. In this paper we perform an empirical evaluation of the system using real-world Web sources. The evaluation demonstrates the effectiveness of the approach, showing that we can automatically learn complex models for real sources in reasonable time. We also compare our system with a complex schema matching system, showing that our approach can handle the kinds of problems tackled by the latter.
Mark J. Carman, Craig A. Knoblock
J. Artif. Intell. Res.1
2005 Learning Source Descriptions for Web Services
Mark J. Carman
AAAI1
2002 Towards an Economy-Based Optimisation of File Access and Replication on a Data Grid
abstract
We are working on a system for the optimised access and replication of data on a Data Grid. Our approach is based on the use of an economic model that includes the actors and the resources in the Grid. Optimisation is obtained via interaction of the actors in the model, whose goals are maximising the profits and minimising the costs of data resource management. In the system, local optimisation results in global optimisation through emergent marketplace behaviour. In this paper we give an overview of our model and present part of the complex economic reasoning required to support this desired marketplace interaction model.
Mark J. Carman, Floriano Zini, Luciano Serafini, Kurt Stockinger
CCGRID1