EDBT 2026 Demo / reviewers in the wild / expert
Amila Silva
dblp:220/0876
· DBLP profile ↗
11ranked-venue papers
10as first author
6since 2021 · last 2024
0000-0003-2042-9050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LipAT: Beyond Style Transfer for Controllable Neural Simulation of Lipstick using Cosmetic AttributesabstractLipstick virtual try-on (VTO) experiences have become widespread across the e-commerce sector and assist users in eliminating the guesswork of shopping online. However, such experiences still lack in both realism and accuracy. In this work, we propose LipAT, a neural framework that blends the strengths of Physics-Based Rendering (PBR) and Neural Style Transfer (NST) approaches to directly apply lipstick onto face images given lipstick attributes (e.g., colour, finish type). LipAT consists of a physics aware neural lipstick application module (LAM) to apply lipstick on face images given its attributes and Lipstick Refiner Module (LRM) to improve the realism by refining the imperfections. Unlike the NST approaches, LipAT allows precise and controllable lipstick attribute preservation, without requiring crude approximations and inference of various intertwined environment factors (e.g., scene lighting, face structure etc) involved in image generation that is required for accurate PBR. We propose an experimental framework with quantitative metrics to evaluate different desirable aspects of the lipstick attribute driven try-on alongside user studies to further validate our findings. Our results show that LipAT considerably outperforms fully-automated PBR approaches in preserving realism and the NST approaches in preserving various lipstick attributes such as finish types. Amila Silva, Olga Moskvyak, Alexander Long, Ravi Garg, Stephen Gould, Gil Avraham, Anton van den Hengel |
WACV | 1 |
| 2024 | Unsupervised Domain-Agnostic Fake News Detection Using Multi-Modal Weak SignalsabstractThe emergence of social media as one of the main platforms for people to access news has enabled the wide dissemination of fake news, having serious impacts on society. Thus, it is really important to identify fake news with high confidence in a timely manner, which is not feasible using manual analysis. This has motivated numerous studies on automating fake news detection. Most of these approaches are supervised, which requires extensive time and labour to build a labelled dataset. Although there have been limited attempts at unsupervised fake news detection, their performance suffers due to not exploiting the knowledge from various modalities related to news records and due to the presence of various latent biases in the existing news datasets (e.g., unrealistic real and fake news distributions). To address these limitations, this work proposes an effective framework for unsupervised fake news detection, which first embeds the knowledge available in four modalities (i.e., source credibility, textual content, propagation speed, and user credibility) in news records and then proposes$(UMD)^{2}$, a novel noise-robust self-supervised learning technique, to identify the veracity of news records from the multi-modal embeddings. Also, we propose a novel technique to construct news datasets minimizing the latent biases in existing news datasets. Following the proposed approach for dataset construction, we produce a Large-scale Unlabelled News Dataset consisting 419,351 news articles related to COVID-19, acronymed asLUND-COVID. We trained the proposed unsupervised framework usingLUND-COVIDto exploit the potential of large datasets, and evaluate it using a set of existing labelled datasets. Our results show that the proposed unsupervised framework largely outperforms existing unsupervised baselines for different tasks such as multi-modal fake news detection, fake news early detection and few-shot fake news detection, while yielding notable improvements for unseen domains during training. Amila Silva, Ling Luo 0002, Shanika Karunasekera, Christopher Leckie |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Noise-Robust Learning from Multiple Unsupervised Sources of Inferred LabelsabstractDeep Neural Networks (DNNs) generally require large-scale datasets for training. Since manually obtaining clean labels for large datasets is extremely expensive, unsupervised models based on domain-specific heuristics can be used to efficiently infer the labels for such datasets. However, the labels from such inferred sources are typically noisy, which could easily mislead and lessen the generalizability of DNNs. Most approaches proposed in the literature to address this problem assume the label noise depends only on the true class of an instance (i.e., class-conditional noise). However, this assumption is not realistic for the inferred labels as they are typically inferred based on the features of the instances. The few recent attempts to model such instance-dependent (i.e., feature-dependent) noise require auxiliary information about the label noise (e.g., noise rates or clean samples). This work proposes a theoretically motivated framework to correct label noise in the presence of multiple labels inferred from unsupervised models. The framework consists of two modules: (1) MULTI-IDNC, a novel approach to correct label noise that is instance-dependent yet not class-conditional; (2) MULTI-CCNC, which extends an existing class-conditional noise-robust approach to yield improved class-conditional noise correction using multiple noisy label sources. We conduct experiments using nine real-world datasets for three different classification tasks (images, text and graph nodes). Our results show that our approach achieves notable improvements (e.g., 6.4% in accuracy) against state-of-the-art baselines while dealing with both instance-dependent and class-conditional noise in inferred label sources. Amila Silva, Ling Luo 0002, Shanika Karunasekera, Christopher Leckie |
AAAI | 1 |
| 2021 | Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataabstractWith the rapid evolution of social media, fake news has become a significant social problem, which cannot be addressed in a timely manner using manual investigation. This has motivated numerous studies on automating fake news detection. Most studies explore supervised training models with different modalities (e.g., text, images, and propagation networks) of news records to identify fake news. However, the performance of such techniques generally drops if news records are coming from different domains (e.g., politics, entertainment), especially for domains that are unseen or rarely-seen during training. As motivation, we empirically show that news records from different domains have significantly different word usage and propagation patterns. Furthermore, due to the sheer volume of unlabelled news records, it is challenging to select news records for manual labelling so that the domain-coverage of the labelled dataset is maximised. Hence, this work: (1) proposes a novel framework that jointly preserves domain-specific and cross-domain knowledge in news records to detect fake news from different domains; and (2) introduces an unsupervised technique to select a set of unlabelled informative news records for manual labelling, which can be ultimately used to train a fake news detection model that performs well for many domains while minimizing the labelling cost. Our experiments show that the integration of the proposed fake news model and the selective annotation approach achieves state-of-the-art performance for cross-domain news datasets, while yielding notable improvements for rarely-appearing domains in news datasets. Amila Silva, Ling Luo 0002, Shanika Karunasekera, Christopher Leckie |
AAAI | 1 |
| 2021 | On Predicting Personal Values of Social Media Users using Community-Specific Language Features and Personal Value Correlation
Amila Silva, Pei-Chi Lo, Ee-Peng Lim |
ICWSM | 1 |
| 2021 | Propagation2Vec: Embedding partial propagation networks for explainable fake news early detection
Amila Silva, Yi Han 0003, Ling Luo 0002, Shanika Karunasekera, Christopher Leckie |
Inf. Process. Manag. | 1 |
| 2020 | METEOR: Learning Memory and Time Efficient Representations from Multi-modal Data StreamsabstractMany learning tasks involve multi-modal data streams, where continuous data from different modes convey a comprehensive description about objects. A major challenge in this context is how to efficiently interpret multi-modal information in complex environments. This has motivated numerous studies on learning unsupervised representations from multi-modal data streams. These studies aim to understand higher-level contextual information (e.g., a Twitter message) by jointly learning embeddings for the lower-level semantic units in different modalities (e.g., text, user, and location of a Twitter message). However, these methods directly associate each low-level semantic unit with a continuous embedding vector, which results in high memory requirements. Hence, deploying and continuously learning such models in low-memory devices (e.g., mobile devices) becomes a problem. To address this problem, we present METEOR, a novel MEmory and Time Efficient Online Representation learning technique, which: (1) learns compact representations for multi-modal data by sharing parameters within semantically meaningful groups and preserves the domain-agnostic semantics; (2) can be accelerated using parallel processes to accommodate different stream rates while capturing the temporal changes of the units; and (3) can be easily extended to capture implicit/explicit external knowledge related to multi-modal data streams. We evaluate METEOR using two types of multi-modal data streams (i.e., social media streams and shopping transaction streams) to demonstrate its ability to adapt to different domains. Our results show that METEOR preserves the quality of the representations while reducing memory usage by around 80% compared to the conventional memory-intensive embeddings. Amila Silva, Shanika Karunasekera, Christopher Leckie, Ling Luo 0002 |
CIKM | 1 |
| 2020 | JPLink: On Linking Jobs to Vocational Interest Types
Amila Silva, Pei-Chi Lo, Ee-Peng Lim |
PAKDD (2) | 1 |
| 2020 | OMBA: User-Guided Product Representations for Online Market Basket Analysis
Amila Silva, Ling Luo 0002, Shanika Karunasekera, Christopher Leckie |
ECML/PKDD (1) | 1 |
| 2019 | Understanding Multilingual Communities through Analysis of Code-switching Behaviors in Social Media DiscussionsabstractCurrently, the enormous span of social media usage - while providing valuable resources for linguistic behavior analysis - makes tracking and understanding these multilingual discussions a challenging task. We have undertaken a multidisciplinary comprehensive study of multilingual discussions via the development of specialized data collection techniques that discover and track multilingual users of social media, and their associated discussions, within a defined geographical region. To facilitate automatic discussion analysis of large numbers of discussions we generated a machine learning model based on ground truth data obtained from Amazon Turk. Our approach goes beyond analyzing social media posts in isolation, by analyzing them in the context of the discussion in which they appear. We show a selection of example discussions found using our approach which reveals a number of interesting socio-linguistic interactions in the communities that we sampled, in support of approach as a general methodology for multilingual community analysis. Aaron Harwood, Shanika Karunasekera, Michelle Vanni, Lucia Falzon, Prarthana Padia, Amila Silva |
IEEE BigData | 6 |
| 2019 | USTAR: Online Multimodal Embedding for Modeling User-Guided Spatiotemporal ActivityabstractBuilding spatiotemporal activity models for people's activities in urban spaces is important for understanding the ever-increasing complexity of urban dynamics. With the emergence of Geo-Tagged Social Media (GTSM) records, previous studies demonstrate the potential of GTSM records for spatiotemporal activity modeling. State-of-the-art methods for this task embed different modalities (location, time, and text) of GTSM records into a single embedding space. However, they ignore Non-GeoTagged Social Media (NGTSM) records, which generally account for the majority of posts (e.g., more than 95% in Twitter), and could represent a great source of information to alleviate the sparsity of GTSM records. Furthermore, in the current spatiotemporal embedding techniques, less focus has been given to the users, who exhibit spatially motivated behaviors. To bridge this research gap, this work proposes USTAR, a novel online learning method for User-guided SpatioTemporal Activity Representation, which (1) embeds locations, time, and text along with users into the same embedding space to capture their correlations; (2) uses a novel collaborative filtering approach to incorporate both NGTSM and GTSM records in learning; and (3) introduces a novel sampling technique to learn spatiotemporal representations in an online fashion to accommodate recent information into the embedding space, while avoiding overfitting to recent records and frequently appearing units in social media streams. Our results show that USTAR substantially improves the state-of-the-art for region retrieval and keyword retrieval and its potential to be applied to other downstream applications such as local event detection. Amila Silva, Shanika Karunasekera, Christopher Leckie, Ling Luo 0002 |
IEEE BigData | 1 |