EDBT 2026 Demo / reviewers in the wild / expert
Michal Lukasik
dblp:72/11338
· DBLP profile ↗
22ranked-venue papers
10as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Language models and text generation · 28% Learning theory · 20% Deep learning architectures and training · 13% | |
| Databases, data mining, and information retrieval
6 papers |
Web and social media mining · 32% Data mining · 29% Information retrieval · 20% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 30 heaviest of 44, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model fine-tuning |
1.6 | 2 | 2025 | Better autoregressive regression with LLMs via regression-aware fine-tuning · ICLR 2025 Two-stage LLM Fine-tuning with Less Specialization and More Generalization · ICLR 2024 |
Machine learning › Learning theory › ranking
AUC optimization |
0.9 | 1 | 2025 | Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation · ICML 2025 |
Machine learning › Learning theory › statistical pattern recognition
bayes optimal classifier |
0.9 | 1 | 2025 | Better autoregressive regression with LLMs via regression-aware fine-tuning · ICLR 2025 |
Machine learning › Learning theory › ranking
bipartite ranking |
0.9 | 1 | 2025 | Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation · ICML 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › crowdsourced annotation
crowdsourced label aggregation |
0.9 | 1 | 2025 | Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge · ACL (1) 2025 |
Machine learning › Deep learning architectures and training › regularization
label smoothing |
0.9 | 2 | 2020 | Does label smoothing mitigate label noise? · ICML 2020 Semantic Label Smoothing for Sequence to Sequence Problems · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge · ACL (1) 2025 |
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff |
0.8 | 1 | 2024 | On Bias-Variance Alignment in Deep Models · ICLR 2024 |
Machine learning › Trustworthy machine learning
calibration |
0.8 | 1 | 2024 | On Bias-Variance Alignment in Deep Models · ICLR 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 1 | 2024 | Two-stage LLM Fine-tuning with Less Specialization and More Generalization · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
fine-tuning generalization |
0.8 | 1 | 2024 | Two-stage LLM Fine-tuning with Less Specialization and More Generalization · ICLR 2024 |
Machine learning › Deep learning architectures and training
neural collapse |
0.8 | 1 | 2024 | On Bias-Variance Alignment in Deep Models · ICLR 2024 |
Natural language and speech › Language models and text generation
prompt tuning |
0.8 | 1 | 2024 | Two-stage LLM Fine-tuning with Less Specialization and More Generalization · ICLR 2024 |
Natural language and speech › Language models and text generation › prompt tuning
soft prompt tuning |
0.8 | 1 | 2024 | Two-stage LLM Fine-tuning with Less Specialization and More Generalization · ICLR 2024 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | ResMem: Learn what you can and memorize the rest · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
residual learning |
0.7 | 1 | 2023 | ResMem: Learn what you can and memorize the rest · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis › text segmentation
discourse segmentation |
0.4 | 1 | 2020 | Text Segmentation by Cross Segment Attention · EMNLP (1) 2020 |
Machine learning › Graph learning
graph neural network |
0.4 | 1 | 2020 | Scaling Graph Neural Networks with Approximate PageRank · KDD 2020 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.4 | 1 | 2020 | Does label smoothing mitigate label noise? · ICML 2020 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
loss correction |
0.4 | 1 | 2020 | Does label smoothing mitigate label noise? · ICML 2020 |
Machine learning › Graph learning › graph neural network
scalable graph neural network |
0.4 | 1 | 2020 | Scaling Graph Neural Networks with Approximate PageRank · KDD 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.4 | 1 | 2020 | Semantic Label Smoothing for Sequence to Sequence Problems · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.4 | 1 | 2020 | Text Segmentation by Cross Segment Attention · EMNLP (1) 2020 |
Computer vision › Segmentation and scene understanding › semantic segmentation
transformer-based segmentation |
0.4 | 1 | 2020 | Text Segmentation by Cross Segment Attention · EMNLP (1) 2020 |
Graph algorithms and graph theory › centrality
pagerank |
0.4 | 1 | 2020 | Scaling Graph Neural Networks with Approximate PageRank · KDD 2020 |
Graph algorithms and graph theory › centrality
pagerank approximation |
0.4 | 1 | 2020 | Scaling Graph Neural Networks with Approximate PageRank · KDD 2020 |
Data mining › text mining
text classification |
0.4 | 1 | 2019 | Gaussian Processes for Rumour Stance Classification in Social Media · ACM Trans. Inf. Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
regression-aware fine-tuning · 1.7cross-entropy loss · 1.7loss aggregation · 0.9chain-of-thought · 0.9bayes-optimal analysis · 0.9autoregressive sampling · 0.9message passing · 0.9approximate pagerank · 0.9prompt tuning · 0.8in-context learning · 0.8ensemble analysis · 0.8linear regression · 0.7multi-task learning · 0.6supervised learning · 0.5inhomogeneous poisson process · 0.4gaussian process classifier · 0.4deep learning · 0.3cross-entropy loss modification · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-JudgeabstractThe LLM-as-a-judge paradigm uses large language models (LLMs) for automated text evaluation, which assigns a score to the text based on some scoring rubrics.Existing methods for LLM-as-a-judge use cross-entropy (CE) loss for fine-tuning, which neglects the numeric nature of score prediction.Recent work addresses numerical prediction limitations of LLM fine-tuning through regression-aware finetuning, which, however, does not consider chain-of-thought (CoT) reasoning for score prediction.In this paper, we introduce TRACT (Two-stage Regression-Aware fine-tuning with CoT), a method combining CoT reasoning with regression-aware training.The training objective of TRACT combines the CE loss for learning the CoT reasoning and the regression-aware loss for the score prediction.TRACT consists of two stages: first, a seed LLM is fine-tuned to generate CoTs; next, we retrain the seed LLM using the CoTs generated by the LLM trained in stage 1.Experiments across four LLM-as-a-judge datasets and two LLMs show that TRACT significantly outperforms existing methods.Extensive ablation studies validate the importance of each component in TRACT. 1 Cheng-Han Chiang, Hung-yi Lee, Michal Lukasik |
ACL (1) | 3 |
| 2025 | Better autoregressive regression with LLMs via regression-aware fine-tuningabstractDecoder-based large language models (LLMs) have proven highly versatile, with remarkable successes even on problems ostensibly removed from traditional language generation. One such example is solving regression problems, where the targets are real numbers rather than textual tokens. A common approach to use LLMs on such problems is to perform fine-tuning based on the cross-entropy loss, and use autoregressive sampling at inference time. Another approach relies on fine-tuning a separate predictive head with a suitable loss such as squared error. While each approach has had success, there has been limited study on principled ways of using decoder LLMs for regression. In this work, we compare different prior works under a unified view, and introduce regression-aware fine-tuning(RAFT), a novel approach based on the Bayes-optimal decision rule. We demonstrate how RAFT improves over established baselines on several benchmarks and model families. Michal Lukasik, Harikrishna Narasimhan, Yin-Wen Chang, Aditya Krishna Menon, Felix Yu, Sanjiv Kumar |
ICLR | 1 |
| 2025 | Bipartite Ranking From Multiple Labels: On Loss Versus Label AggregationabstractBipartite ranking is a fundamental supervised learning problem, with the goal of learning a ranking over instances with maximal area under the ROC curve (AUC) against a single binary target label. However, one may often observe multiple binary target labels, e.g., from distinct human annotators. How can one synthesize such labels into a single coherent ranking? In this work, we formally analyze two approaches to this problem—loss aggregation and label aggregation—by characterizing their Bayes-optimal solutions. We show that while both approaches can yield Pareto-optimal solutions, loss aggregation can exhibit label dictatorship: one can inadvertently (and undesirably) favor one label over others. This suggests that label aggregation can be preferable to loss aggregation, which we empirically verify. Michal Lukasik, Lin Chen 0003, Harikrishna Narasimhan, Aditya Krishna Menon, Wittawat Jitkrittum, Felix X. Yu, Sashank J. Reddi, Mohammad Hossein Bateni 0001, Sanjiv Kumar |
ICML | 1 |
| 2024 | On Bias-Variance Alignment in Deep ModelsabstractClassical wisdom in machine learning holds that the generalization error can be decomposed into bias and variance, and these two terms exhibit a \emph{trade-off}. However, in this paper, we show that for an ensemble of deep learning based classification models, bias and variance are \emph{aligned} at a sample level, where squared bias is approximately \emph{equal} to variance for correctly classified sample points. We present empirical evidence confirming this phenomenon in a variety of deep learning models and datasets. Moreover, we study this phenomenon from two theoretical perspectives: calibration and neural collapse. We first show theoretically that under the assumption that the models are well calibrated, we can observe the bias-variance alignment. Second, starting from the picture provided by the neural collapse theory, we show an approximate correlation between bias and variance. Michal Lukasik, Wittawat Jitkrittum, Chong You, Sanjiv Kumar |
ICLR | 2 |
| 2024 | Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationabstractPretrained large language models (LLMs) are general purpose problem solvers applicable to a diverse set of tasks with prompts. They can be further improved towards a specific task by fine-tuning on a specialized dataset. However, fine-tuning usually makes the model narrowly specialized on this dataset with reduced general in-context learning performances, which is undesirable whenever the fine-tuned model needs to handle additional tasks where no fine-tuning data is available.
In this work, we first demonstrate that fine-tuning on a single task indeed decreases LLMs' general in-context learning performance. We discover one important cause of such forgetting, format specialization, where the model overfits to the format of the fine-tuned task.We further show that format specialization happens at the very beginning of fine-tuning. To solve this problem, we propose Prompt Tuning with MOdel Tuning (ProMoT), a simple yet effective two-stage fine-tuning framework that reduces format specialization and improves generalization.ProMoT offloads task-specific format learning into additional and removable parameters by first doing prompt tuning and then fine-tuning the model itself with this soft prompt attached.
With experiments on several fine-tuning tasks and 8 in-context evaluation tasks, we show that ProMoT achieves comparable performance on fine-tuned tasks to standard fine-tuning, but with much less loss of in-context learning performances across a board range of out-of-domain evaluation tasks. More importantly, ProMoT can even enhance generalization on in-context learning tasks that are semantically related to the fine-tuned task, e.g. ProMoT on En-Fr translation significantly improves performance on other language pairs, and ProMoT on NLI improves performance on summarization.
Experiments also show that ProMoT can improve the generalization performance of multi-task training. Si Si, Daliang Li, Michal Lukasik, Felix X. Yu, Cho-Jui Hsieh, Inderjit S. Dhillon, Sanjiv Kumar |
ICLR | 4 |
| 2023 | ResMem: Learn what you can and memorize the restabstractThe impressive generalization performance of modern neural networks is attributed in part to their ability to implicitly memorize complex training patterns.
Inspired by this, we explore a novel mechanism to improve model generalization via explicit memorization.
Specifically, we propose the residual-memorization (ResMem) algorithm, a new method that augments an existing prediction model (e.g., a neural network) by fitting the model's residuals with a nearest-neighbor based regressor.
The final prediction is then the sum of the original model and the fitted residual regressor.
By construction, ResMem can explicitly memorize the training labels.
We start by formulating a stylized linear regression problem and rigorously show that ResMem results in a more favorable test risk over a base linear neural network.
Then, we empirically show that ResMem consistently improves the test set generalization of the original prediction model across standard vision and natural language processing benchmarks. Zitong Yang, Michal Lukasik, Vaishnavh Nagarajan, Ankit Singh Rawat, Manzil Zaheer, Aditya Krishna Menon, Sanjiv Kumar |
NeurIPS | 2 |
| 2023 | Robust distillation for worst-class performance: on the interplay between teacher and student objectivesabstractKnowledge distillation is a popular technique that has been shown to produce remarkable gains in average accuracy. However, recent work has shown that these gains are not uniform across subgroups in the data, and can often come at the cost of accuracy on rare subgroups and classes. Robust optimization is a common remedy to improve worst-class accuracy in standard learning settings, but in distillation it is unknown whether it is best to apply robust objectives when training the teacher, the student, or both. This work studies the interplay between robust objectives for the teacher and student. Empirically, we show that that jointly modifying the teacher and student objectives can lead to better worst-class student performance and even Pareto improvement in the trade-off between worst-class and overall performance. Theoretically, we show that the per-class calibration of teacher scores is key when training a robust student. Both the theory and experiments support the surprising finding that applying a robust teacher training objective does not always yield a more robust student. Serena Lutong Wang, Harikrishna Narasimhan, Sara Hooker, Michal Lukasik, Aditya Krishna Menon |
UAI | 5 |
| 2022 | Hawkes Process Classification through Discriminative Modeling of TextabstractSocial media such as Twitter has provided a platform for users to gather and share information and stay updated with the news. However, restriction on the length, informal grammar and vocabulary of the posts pose challenges to perform classification from textual content alone. We propose models based on the Hawkes process (HP) which can naturally incorporate additional cues such as the temporal features and past labels of the posts, along with the textual features for improving short text classification. In particular, we propose a discriminative approach to model text in HP, where the text features parameterize the base intensity and the triggering kernel of the intensity function. This allows textual content to determine influence from past posts and consequently determine the intensity function and class label. Another major contribution is to model the kernel as a neural network function of both time and text, permitting more complex influence functions for Hawkes process. This will maintain the interpretability of Hawkes process models along with the improved function learning capability of the neural networks. The proposed HP models can easily consider pretrained word embeddings to represent text for classification. Experiments on the rumour stance classification problems in social media demonstrate the effectiveness of the proposed HP models. Rohan Tondulkar, Manisha Dubey, P. K. Srijith, Michal Lukasik |
IJCNN | 4 |
| 2020 | Text Segmentation by Cross Segment AttentionabstractDocument and discourse segmentation are two fundamental NLP tasks pertaining to breaking up text into constituents, which are commonly used to help downstream tasks such as information retrieval or text summarization.In this work, we propose three transformer-based architectures and provide comprehensive comparisons with previously proposed approaches on three standard datasets.We establish a new state-of-the-art, reducing in particular the error rates by a large margin in all cases.We further analyze model sizes and find that we can build models with many fewer parameters while keeping good performance, thus facilitating real-world applications.Early life and marriage: Franklin Delano Roosevelt was born on January 30, 1882, in the Hudson Valley town of Hyde Park, New York, to businessman James Roosevelt I and his second wife, Sara Ann Delano.(...) Aides began to refer to her at the time as "the president's girlfriend", and gossip linking the two romantically appeared in the newspapers.(...) Legacy: Roosevelt is widely considered to be one of the most important figures in the history of the United States, as well as one of the most influential figures of the 20th century.(...) Roosevelt has also appeared on several U.S. Postage stamps. Michal Lukasik, Boris Dadachev, Kishore Papineni, Gonçalo Simões |
EMNLP (1) | 1 |
| 2020 | Semantic Label Smoothing for Sequence to Sequence ProblemsabstractMichal Lukasik, Himanshu Jain, Aditya Menon, Seungyeon Kim, Srinadh Bhojanapalli, Felix Yu, Sanjiv Kumar. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Michal Lukasik, Himanshu Jain, Aditya Krishna Menon, Seungyeon Kim 0001, Srinadh Bhojanapalli, Felix X. Yu, Sanjiv Kumar |
EMNLP (1) | 1 |
| 2020 | Does label smoothing mitigate label noise?abstractLabel smoothing is commonly used in training deep learning models, wherein one-hot training labels are mixed with uniform label vectors. Empirically, smoothing has been shown to improve both predictive performance and model calibration. In this paper, we study whether label smoothing is also effective as a means of coping with label noise. While label smoothing apparently amplifies this problem — being equivalent to injecting symmetric noise to the labels — we show how it relates to a general family of loss-correction techniques from the label noise literature. Building on this connection, we show that label smoothing is competitive with loss-correction under label noise. Further, we show that when distilling models from noisy data, label smoothing of the teacher is beneficial; this is in contrast to recent findings for noise-free problems, and sheds further light on settings where label smoothing is beneficial. Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv Kumar |
ICML | 1 |
| 2020 | Scaling Graph Neural Networks with Approximate PageRankabstractGraph neural networks (GNNs) have emerged as a powerful approach for solving many network mining tasks. However, learning on large graphs remains a challenge -- many recently proposed scalable GNN approaches rely on an expensive message-passing procedure to propagate information through the graph. We present the PPRGo model which utilizes an efficient approximation of information diffusion in GNNs resulting in significant speed gains while maintaining state-of-the-art prediction performance. In addition to being faster, PPRGo is inherently scalable, and can be trivially parallelized for large datasets like those found in industry settings. Aleksandar Bojchevski, Johannes Gasteiger, Bryan Perozzi, Amol Kapoor, Martin Blais, Benedek Rozemberczki, Michal Lukasik, Stephan Günnemann |
KDD | 7 |
| 2019 | Gaussian Processes for Rumour Stance Classification in Social MediaabstractSocial media tend to be rife with rumours while new reports are released piecemeal during breaking news. Interestingly, one can mine multiple reactions expressed by social media users in those situations, exploring their stance towards rumours, ultimately enabling the flagging of highly disputed rumours as being potentially false. In this work, we set out to develop an automated, supervised classifier that uses multi-task learning to classify the stance expressed in each individual tweet in a conversation around a rumour as either supporting, denying or questioning the rumour. Using a Gaussian Process classifier, and exploring its effectiveness on two datasets with very different characteristics and varying distributions of stances, we show that our approach consistently outperforms competitive baseline classifiers. Our classifier is especially effective in estimating the distribution of different types of stance associated with a given rumour, which we set forth as a desired characteristic for a rumour-tracking system that will show both ordinary users of Twitter and professional news practitioners how others orient to the disputed veracity of a rumour, with the final aim of establishing its actual truth value. Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Arkaitz Zubiaga, Maria Liakata, Rob Procter |
ACM Trans. Inf. Syst. | 1 |
| 2018 | Content Explorer: Recommending Novel Entities for a Document WriterabstractBackground research is an essential part of document writing.Search engines are great for retrieving information once we know what to look for.However, the bigger challenge is often identifying topics for further research.Automated tools could help significantly in this discovery process and increase the productivity of the writer.In this paper, we formulate the problem of recommending topics to a writer.We consider this as a supervised learning problem and run a user study to validate this approach.We propose an evaluation metric and perform an empirical comparison of state-of-the-art models for extreme multi-label classification on a large data set.We demonstrate how a simple modification of the crossentropy loss function leads to improved results of the deep learning models. Michal Lukasik, Richard Zens |
EMNLP | 1 |
| 2018 | A Bayesian Point Process Model for User Return Time Prediction in Recommendation SystemsabstractIn order to sustain the user-base for a web service, it is important to know the return time of a user to the service. We propose a Bayesian point process, log Gaussian Cox process (LGCP), to model and predict return time of users. It allows encoding the prior domain knowledge and non-parametric estimation of latent intensity functions capturing user behaviour. We capture the similarities among the users in their return time by using a multi-task learning approach. We show the effectiveness of the proposed approaches on predicting the return time of users to last.fm music service. Sherin Thomas, P. K. Srijith, Michal Lukasik |
UMAP | 3 |
| 2018 | Discourse-aware rumour stance classification in social media using sequential classifiers
Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Isabelle Augenstein |
Inf. Process. Manag. | 5 |
| 2017 | Longitudinal Modeling of Social Media with Hawkes Process Based on Users and NetworksabstractOnline social media provide a platform for rapid network propagation of information at an unprecedented scale. In this paper, we study the evolution of information cascades in Twitter using a point process model of user activity. Twitter is rich with heterogenous information on users and network structure. We develop several Hawkes process models considering various properties of Twitter including conversational structure, users' connections and general features of users including the textual information, and show how they are helpful in modeling the social network activity. Evaluation on Twitter data sets shows that incorporating richer properties improves the performance in predicting future activity of users and memes. P. K. Srijith, Michal Lukasik, Kalina Bontcheva, Trevor Cohn |
ASONAM | 2 |
| 2016 | Convolution Kernels for Discriminative Learning from Streaming TextabstractTime series modeling is an important problem with many applications in different domains. Here we consider discriminative learning from time series, where we seek to predict an output response variable based on time series input. We develop a method based on convolution kernels to model discriminative learning over streams of text. Our method outperforms competitive baselines in three synthetic and two real datasets, rumour frequency modeling and popularity prediction tasks. Michal Lukasik, Trevor Cohn |
AAAI | 1 |
| 2016 | Stance Classification in Rumours as a Sequential Task Exploiting the Tree Structure of Social Media ConversationsabstractRumour stance classification, the task that determines if each tweet in a collection discussing a rumour is supporting, denying, questioning or simply commenting on the rumour, has been attracting substantial interest. Here we introduce a novel approach that makes use of the sequence of transitions observed in tree-structured conversation threads in Twitter. The conversation threads are formed by harvesting users’ replies to one another, which results in a nested tree-like structure. Previous work addressing the stance classification task has treated each tweet as a separate unit. Here we analyse tweets by virtue of their position in a sequence and test two sequential classifiers, Linear-Chain CRF and Tree CRF, each of which makes different assumptions about the conversational structure. We experiment with eight Twitter datasets, collected during breaking news, and show that exploiting the sequential structure of Twitter conversations achieves significant improvements over the non-sequential methods. Our work is the first to model Twitter conversations as a tree structure in this manner, introducing a novel way of tackling NLP tasks on Twitter conversations. Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik |
COLING | 5 |
| 2015 | Classifying Tweet Level Judgements of Rumours in Social MediaabstractSocial media is a rich source of rumours and corresponding community reactions.Rumours reflect different characteristics, some shared and some individual.We formulate the problem of classifying tweet level judgements of rumours as a supervised learning task.Both supervised and unsupervised domain adaptation are considered, in which tweets from a rumour are classified on the basis of other annotated rumours.We demonstrate how multi-task learning helps achieve good results on rumours from the 2011 England riots. Michal Lukasik, Trevor Cohn, Kalina Bontcheva |
EMNLP | 1 |
| 2015 | Modeling Tweet Arrival Times using Log-Gaussian Cox ProcessesabstractResearch on modeling time series text corpora has typically focused on predicting what text will come next, but less well studied is predicting when the next text event will occur.In this paper we address the latter case, framed as modeling continuous inter-arrival times under a log-Gaussian Cox process, a form of inhomogeneous Poisson process which captures the varying rate at which the tweets arrive over time.In an application to rumour modeling of tweets surrounding the 2014 Ferguson riots, we show how interarrival times between tweets can be accurately predicted, and that incorporating textual features further improves predictions. Michal Lukasik, P. K. Srijith, Trevor Cohn, Kalina Bontcheva |
EMNLP | 1 |
| 2012 | Evaluation of Features for Author Name Disambiguation Using Linear Support Vector MachinesabstractAuthor name disambiguation allows to distinguish between two or more authors sharing the same name. In a previous paper, we have proposed a name disambiguation framework in which for each author name in each article we build a context consisting of classification codes, bibliographic references, co-authors, etc. Then, by pair wise comparison of contexts, we have been grouping contributions likely referring to the same people. In this paper we examine which elements of the context are most effective in author name disambiguation. We employ linear Support Vector Machines (SVM) to find the most influential features. Piotr Jan Dendek, Lukasz Bolikowski, Michal Lukasik |
Document Analysis Systems | 3 |