VLDB 2026 Research / reviewers in the wild / expert
Advaith Siddharthan
dblp:95/3737
· DBLP profile ↗
33ranked-venue papers
12as first author
7since 2021 · last 2026
0000-0003-0796-8826ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 11 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sensing Nature: A School-ground Haptic Map and Tactile Garden for Texture ExplorationabstractThis paper presents HapticMap and Tactile Garden, an art installation with a demo contribution for IDC 2026 that invites participants to explore school-ground textures through vibrotactile interaction, projected drawings, and tactile surfaces. HapticMap features a touchscreen map of a real school site with a vibrotactile stylus and close-up images of bark, leaves, water and other natural materials. Tactile Garden extends this into a table-based installation built from children’s hand-drawn field sketches, projected onto translucent panels that deliver haptic-audio feedback on touch. The demo and table installation is grounded in a three-session study with 11 secondary school students in Edinburgh, in which outdoor sketching, guided haptic exploration and a return to school grounds was used to foreground touch as a mode of attention. Instead of attempting overly realistic representation of the textures, the system utilises a gap between digital and physical touch as a prompt for comparison, description and discussion. We present this installation as a contribution to IDC discussions around sensory design, place-based learning and ways in which interactive technologies can enable children and adults to notice everyday environments differently. Lisa J. Bowers, Nirwan Sharma, Jonathan Hancock, Poppy Lakeman Fraser, Julie Newman, Andrew Manches, Laura Colucci-Gray, Advaith Siddharthan |
IDC | 9 |
| 2026 | The role of touch in children's engagement with nature: implications for design: Role of touch in learningabstractEmerging technologies such as haptics offer increasing possibilities to support children's learning by exploiting their sense of touch. However, this potential is hindered by an empirical and conceptual gap in understanding the unique affordances of touch as distinct from broader physical interaction. We address this gap through in-depth video analysis of 82 children's (6–11-year-old) exploration and description of natural objects. Analysis revealed much variation in the role of touch across three themes identified from previous literature: children's propensity to touch, touch interaction, and touch communication. Subthemes developed from the data illustrate the importance of disentangling nuanced mechanisms of touch including the novel identification of gestures simulating touch which suggest the embodied internalization of touch experiences. The paper concludes with six intermediate knowledge themes to inform both pedagogical and haptic designs that might tap this undervalued sensory mode to support children's engagement and learning about their natural environment. Andrew Manches, Jonathan Hancock, Laura Colucci-Gray, Advaith Siddharthan |
IDC | 4 |
| 2024 | CSS: Contrastive Semantic Similarities for Uncertainty Quantification of LLMsabstractDespite the impressive capability of large language models (LLMs), knowing when to trust their generations remains an open challenge. The recent literature on uncertainty quantification of natural language generation (NLG) utilizes a conventional natural language inference (NLI) classifier to measure the semantic dispersion of LLMs responses. These studies employ logits of NLI classifier for semantic clustering to estimate uncertainty. However, logits represent the probability of the predicted class and barely contain feature information for potential clustering. Alternatively, CLIP (Contrastive Language{–}Image Pre-training) performs impressively in extracting image-text pair features and measuring their similarity. To extend its usability, we propose Contrastive Semantic Similarity, the CLIP-based feature extraction module to obtain similarity features for measuring uncertainty for text pairs. We apply this method to selective NLG, which detects and rejects unreliable generations for better trustworthiness of LLMs. We conduct extensive experiments with three LLMs on several benchmark question-answering datasets with comprehensive evaluation metrics. Results show that our proposed method performs better in estimating reliable responses of LLMs than comparable baselines. Shuang Ao, Stefan Rueger, Advaith Siddharthan |
UAI | 3 |
| 2023 | Two Sides of Miscalibration: Identifying Over and Under-Confidence Prediction for Network CalibrationabstractProper confidence calibration of deep neural networks is essential for reliable predictions in safety-critical tasks. Miscalibration can lead to model over-confidence and/or under-confidence; i.e., the model’s confidence in its prediction can be greater or less than the model’s accuracy. Recent studies have highlighted the over-confidence issue by introducing calibration techniques and demonstrated success on various tasks. However, miscalibration through under-confidence has not yet to receive much attention. In this paper, we address the necessity of paying attention to the under-confidence issue. We first introduce a novel metric, a miscalibration score, to identify the overall and class-wise calibration status, including being over or under-confident. Our proposed metric reveals the pitfalls of existing calibration techniques, where they often overly calibrate the model and worsen under-confident predictions. Then we utilize the class-wise miscalibration score as a proxy to design a calibration technique that can tackle both over and under-confidence. We report extensive experiments that show our proposed methods substantially outperforming existing calibration techniques. We also validate our proposed calibration technique on an automatic failure detection task with a risk-coverage curve, reporting that our methods improve failure detection as well as trustworthiness of the model. The code are available at \url{https://github.com/AoShuang92/miscalibration_TS}. Shuang Ao, Stefan Rueger, Advaith Siddharthan |
UAI | 3 |
| 2022 | Confidence-aware calibration and scoring functions for curriculum learningabstractDespite the great success of state-of-the-art deep neural networks, several studies have reported models to be over-confident in predictions, indicating miscalibration.Label Smoothing has been proposed as a solution to the over-confidence problem and works by softening hard targets during training, typically by distributing part of the probability mass from a 'one-hot' label uniformly to all other labels.However, neither model nor human confidence in a label are likely to be uniformly distributed in this manner, with some labels more likely to be confused than others.In this paper we integrate notions of model confidence and human confidence with label smoothing, respectively Model Confidence LS and Human Confidence LS, to achieve better model calibration and generalization.To enhance model generalization, we show how our model and human confidence scores can be successfully applied to curriculum learning, a training strategy inspired by learning of 'easier to harder' tasks.A higher model or human confidence score indicates a more recognisable and therefore easier sample, and can therefore be used as a scoring function to rank samples in curriculum learning.We evaluate our proposed methods with four state-of-the-art architectures for image and text classification task, using datasets with multi-rater label annotations by humans.We report that integrating model or human confidence information in label smoothing and curriculum learning improves both model performance and model calibration.The code are available at https://github.com/AoShuang92/ConfidenceCalibration CL. Shuang Ao, Stefan Rueger, Advaith Siddharthan |
ICMV | 3 |
| 2022 | Consensus Building in On-Line Citizen ScienceabstractA number of initiatives invite members of the public to perform online classification tasks such as identifying objects in images. These tasks are crucial to numerous large-scale Citizen Science projects in different disciplines, with volunteers using their knowledge and online support tools to, for example, identify species of wildlife or classify galaxies by their shapes. However, for complex classification tasks, such as this case study on identifying species of bumblebee, reaching an agreement between volunteers - or even between experts~-~may require consensus-building processes. Collaboration and teamwork approaches to problem solving and decision-making have been widely documented to improve both task performance and user learning in the real world. Most of these processes and projects are mediated online through feedback delivered in an asynchronous manner, and this article thus addresses a central research question: How do participants involved in species identification tasks respond to different forms of feedback provided in online collaboration, designed to support peer-learning and improve task performance? We tested four different approaches to feedback within a collaboration task, where participants reviewed their previously annotated data based on information curated from their peers on a long running online citizen science initiative. The selected interfaces have a strong foundation in social science and psychology literature and can be applied to citizen science practices as well as other online communities. Results showed that while all four approaches increased accuracy, there were differences based on the types of consensus that existed before collaboration. Such differences highlight the usefulness of different forms of feedback during collaboration for increasing data accuracy of identification and furthering users' expertise on identification tasks. We found that anonymised and goal-directed free text comments posted on social learning interfaces were most effective in improving data accuracy as well as creating opportunities for peer-learning, particularly where the species identification task was more difficult. This study has significant implications for extending the practice of citizen science across formal and informal learning environments and reaching out to a variety of users. Nirwan Sharma, Laura Colucci-Gray, René van der Wal, Advaith Siddharthan |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2021 | Summarising Historical Text in Modern LanguagesabstractWe introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language.This is a fundamentally important routine to historians and digital humanities researchers but has never been automated.We compile a high-quality gold-standard text summarisation dataset, which consists of historical German and Chinese news from hundreds of years ago summarised in modern German or Chinese.Based on cross-lingual transfer learning techniques, we propose a summarisation model that can be trained even with no cross-lingual (historical to modern) parallel data, and further benchmark it against state-of-the-art algorithms.We report automatic and human evaluations that distinguish the historic to modern language summarisation task from standard cross-lingual summarisation (i.e., modern to modern language), highlight the distinctness and value of our dataset, and demonstrate that our transfer learning approach outperforms standard cross-lingual benchmarks on this task.DE №34 Story Jhre Königl.Majest.befinden sich noch vnweit Thorn / ... / dahero zur Erledigung Hoffnung gemacht werden will.(Their Royal Majesties are still not far from Torn, ... , therefore completion of the hope is desired.)Summary Der Krieg zwischen Polen und Schweden dauert an.Von einem Friedensvertrag ist noch nicht der Rede.(The war between Poland and Sweden continues.There is still no talk on the peace treaty.)ZH №7 Story 有脚夫小民,三四千名集众围绕马监丞衙门,...,冒火突入,捧出敕印。 (Three to four thousand porters gathered around Majiancheng Yamen (a government office), ..., rushed into fire and salvaged the authority's seal.)Summary 小本生意免税条约未能落实,小商贩被严重剥削,以致百姓聚众闹事并火烧衙门,造成多人伤亡。王炀 抢救出公章。 (The tax-exemption act for small businesses was not well implemented and small traders were terribly exploited, leading to riot and arson attack on Yamen with many casualties.Yang Wang salvaged the authority's seal.) Xutan Peng, Chenghua Lin 0002, Advaith Siddharthan |
EACL | 4 |
| 2018 | Generating Summaries of Sets of Consumer Products: Learning from ExperimentsabstractWe explored the task of creating a textual summary describing a large set of objects characterised by a small number of features using an e-commerce dataset.When a set of consumer products is large and varied, it can be difficult for a consumer to understand how the products in the set differ; consequently, it can be challenging to choose the most suitable product from the set.To assist consumers, we generated high-level summaries of product sets.Two generation algorithms are presented, discussed, and evaluated with human users.Our evaluation results suggest a positive contribution to consumers' understanding of the domain. Kittipitch Kuptavanich, Ehud Reiter, Kees van Deemter, Advaith Siddharthan |
INLG | 4 |
| 2018 | Incorporating Constraints into Matrix Factorization for Clothes Package RecommendationabstractRecommender systems have been widely applied in the literature to suggest individual items to users. In this paper, we consider the harder problem of package recommendation, where items are recommended together as a package. We focus on the clothing domain, where a package recommendation involves a combination of a 'top' (e.g. a shirt) and a 'bottom' (e.g. a pair of trousers). The novelty in this work is that we combined matrix factorisation methods for collaborative filtering with hand-crafted and learnt fashion constraints on combining item features such as colour, formality and patterns. Finally, to better understand where the algorithms are underperforming, we conducted focus groups, which lead to deeper insights into how to use constraints to improve package recommendation in this domain. Agung Toto Wibowo, Advaith Siddharthan, Judith Masthoff, Chenghua Lin 0002 |
UMAP | 2 |
| 2018 | SaferDrive: An NLG-based behaviour change support system for driversabstractAbstract Despite the long history of Natural Language Generation (NLG) research, the potential for influencing real world behaviour through automatically generated texts has not received much attention. In this paper, we presentSaferDrive, a behaviour change support system that uses NLG and telematic data in order to create weekly textual feedback for automobile drivers, which is delivered through a smartphone application. Usage-based car insurances use sensors to track driver behaviour. Although the data collected by such insurances could provide detailed feedback about the driving style, they are typically withheld from the driver and used only to calculate insurance premiums.SaferDriveinstead provides detailed textual feedback about the driving style, with the intent to help drivers improve their driving habits. We evaluate the system with real drivers and report that the textual feedback generated by our system does have a positive influence on driving habits, especially with regard to speeding. Daniel Braun 0003, Ehud Reiter, Advaith Siddharthan |
Nat. Lang. Eng. | 3 |
| 2017 | Automatically Labelling Sentiment-Bearing Topics with Descriptive Sentence Labels
Mohamad Hardyman Barawi, Chenghua Lin 0002, Advaith Siddharthan |
NLDB | 3 |
| 2016 | Summarising News Stories for ChildrenabstractThis paper proposes a system to automatically summarise news articles in a manner suitable for children by deriving and combining statistical ratings for how important, positively oriented and easy to read each sentence is.Our results demonstrate that this approach succeeds in generating summaries that are suitable for children, and that there is further scope for combining this extractive approach with abstractive methods used in text simplification. Iain Macdonald, Advaith Siddharthan |
INLG | 2 |
| 2016 | Crowdsourcing Without a Crowd: Reliable Online Species Identification Using Bayesian Models to Minimize Crowd SizeabstractWe present an incremental Bayesian model that resolves key issues of crowd size and data quality for consensus labeling. We evaluate our method using data collected from a real-world citizen science program, B ee W atch , which invites members of the public in the United Kingdom to classify (label) photographs of bumblebees as one of 22 possible species. The biological recording domain poses two key and hitherto unaddressed challenges for consensus models of crowdsourcing: (1) the large number of potential species makes classification difficult, and (2) this is compounded by limited crowd availability, stemming from both the inherent difficulty of the task and the lack of relevant skills among the general public. We demonstrate that consensus labels can be reliably found in such circumstances with very small crowd sizes of around three to five users (i.e., through group sourcing). Our incremental Bayesian model, which minimizes crowd size by re-evaluating the quality of the consensus label following each species identification solicited from the crowd, is competitive with a Bayesian approach that uses a larger but fixed crowd size and outperforms majority voting. These results have important ecological applicability: biological recording programs such as B ee W atch can sustain themselves when resources such as taxonomic experts to confirm identifications by photo submitters are scarce (as is typically the case), and feedback can be provided to submitters in a timely fashion. More generally, our model provides benefits to any crowdsourced consensus labeling task where there is a cost (financial or otherwise) associated with soliciting a label. Advaith Siddharthan, Christopher Lambin, Anne-Marie Robinson, Nirwan Sharma, Richard Comont, Elaine O'Mahony, Chris Mellish, René van der Wal |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2014 | Lexico-syntactic text simplification and compression with typed dependencies
Angrosh Mandya, Tadashi Nomoto, Advaith Siddharthan |
COLING | 3 |
| 2014 | Hybrid text simplification using synchronous dependency grammars with hand-written and automatically harvested rulesabstractWe present an approach to text simplification based on synchronous dependency grammars.The higher level of abstraction afforded by dependency representations allows for a linguistically sound treatment of complex constructs requiring reordering and morphological change, such as conversion of passive voice to active.We present a synchronous grammar formalism in which it is easy to write rules by hand and also acquire them automatically from dependency parses of aligned English and Simple English sentences.The grammar formalism is optimised for monolingual translation in that it reuses ordering information from the source sentence where appropriate.We demonstrate the superiority of our approach over a leading contemporary system based on quasi-synchronous tree substitution grammars, both in terms of expressivity and performance. Advaith Siddharthan, Angrosh Mandya |
EACL | 1 |
| 2014 | Text simplification using synchronous dependency grammars: Generalising automatically harvested rulesabstractWe present an approach to text simplifi-cation based on synchronous dependency grammars. Our main contributions in this work are (a) a study of how automatically derived lexical simplification rules can be generalised to enable their application in new contexts without introducing errors, and (b) an evaluation of our hybrid sys-tem that combines a large set of automat-ically acquired rules with a small set of hand-crafted rules for common syntactic simplification. Our evaluation shows sig-nificant improvements over the state of the art, with scores comparable to human sim-plifications. 1 Angrosh Mandya, Advaith Siddharthan |
INLG | 2 |
| 2012 | Natural Language Generation for Nature Conservation: Automating Feedback to Help Volunteers Identify Bumblebee Species
Steven Blake, Advaith Siddharthan, Nirwan Sharma, Anne-Marie Robinson, Elaine O'Mahony, Ben Darvill, Chris Mellish, René van der Wal |
COLING | 2 |
| 2012 | Blogging birds: Generating narratives about reintroduced species to promote public engagement
Advaith Siddharthan, Matthew Green 0002, Kees van Deemter, Chris Mellish, René van der Wal |
INLG | 1 |
| 2011 | Information Status Distinctions and Referring Expressions: An Empirical Study of References to People in News SummariesabstractAlthough there has been much theoretical work on using various information status distinctions to explain the form of references in written text, there have been few studies that attempt to automatically learn these distinctions for generating references in the context of computer-regenerated text. In this article, we present a model for generating references to people in news summaries that incorporates insights from both theory and a corpus analysis of human written summaries. In particular, our model captures how two properties of a person referred to in the summary—familiarity to the reader and global salience in the news story—affect the content and form of the initial reference to that person in a summary. We demonstrate that these two distinctions can be learned from a typical input for multi-document summarization and that they can be used to make regeneration decisions that improve the quality of extractive summaries. Advaith Siddharthan, Ani Nenkova, Kathy McKeown |
Comput. Linguistics | 1 |
| 2010 | Complex Lexico-syntactic Reformulation of Sentences Using Typed Dependency Representations
Advaith Siddharthan |
INLG | 1 |
| 2010 | Corpora for the Conceptualisation and Zoning of Scientific Papers
Maria Liakata, Simone Teufel, Advaith Siddharthan, Colin R. Batchelor |
LREC | 3 |
| 2010 | Reformulating Discourse Connectives for Non-Expert Readers
Advaith Siddharthan, Napoleon Katsos |
HLT-NAACL | 1 |
| 2010 | Interlingual annotation of parallel text corpora: a new framework for annotation and evaluationabstractAbstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines. Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan |
Nat. Lang. Eng. | 12 |
| 2009 | Towards Domain-Independent Argumentative Zoning: Evidence from Chemistry and Computational Linguistics
Simone Teufel, Advaith Siddharthan, Colin R. Batchelor |
EMNLP | 2 |
| 2008 | Language Resources and Chemical Informatics
C. J. Rupp, Ann A. Copestake, Peter T. Corbett, Peter Murray-Rust, Advaith Siddharthan, Simone Teufel, Benjamin Waldron |
LREC | 5 |
| 2007 | Whose Idea Was This, and Why Does it Matter? Attributing Scientific Work to Citations
Advaith Siddharthan, Simone Teufel |
HLT-NAACL | 1 |
| 2006 | Automatic classification of citation function
Simone Teufel, Advaith Siddharthan, Dan Tidhar |
EMNLP | 2 |
| 2006 | Parallel Syntactic Annotation of Multiple Languages
Owen Rambow, Bonnie J. Dorr, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Flo Reeder, Advaith Siddharthan |
LREC | 12 |
| 2004 | Generating Referring Expressions in Open DomainsabstractWe present an algorithm for generating referring expressions in open domains. Existing algorithms work at the semantic level and assume the availability of a classification for attributes, which is only feasible for restricted domains. Our alternative works at the realisation level, relies on Word-Net synonym and antonym sets, and gives equivalent results on the examples cited in the literature and improved results for examples that prior approaches cannot handle. We believe that ours is also the first algorithm that allows for the incremental incorporation of relations. We present a novel corpus-evaluation using referring expressions from the Penn Wall Street Journal Treebank. Advaith Siddharthan, Ann A. Copestake |
ACL | 1 |
| 2004 | Syntactic Simplification for Improving Content Selection in Multi-Document Summarization
Advaith Siddharthan, Ani Nenkova, Kathy McKeown |
COLING | 1 |
| 2002 | Christopher D. Manning and Hinrich Schutze. Foundations of Statistical Natural Language Processing. MIT Press, 2000. ISBN 0-262-13360-1, 620 pp. $64.95/£44.95 (cloth)
Advaith Siddharthan |
Nat. Lang. Eng. | 1 |
| 2001 | Ehud Reiter and Robert Dale. Building Natural Language Generation Systems. Cambridge University Press, 2000. $64.95/£37.50 (Hardback), 234 pages
Advaith Siddharthan |
Nat. Lang. Eng. | 1 |
| 2001 | Inderjeet Mani and Mark T. Maybury (eds). Advances in Automatic Text Summarization. MIT Press, 1999. ISBN 0-262-13359-8, 442 pp. $47.95/£32.95 (paperback)
Advaith Siddharthan |
Nat. Lang. Eng. | 1 |