Junhua Liu 0002

dblp:30/4261-2 · DBLP profile ↗
← Back
9ranked-venue papers in the field
5as first author
6since 2021 · last 2025
0000-0003-4477-7439ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4 (3 first)Big Data, Cloud & Distributed Data Systems · 3 (1 first)Information Retrieval & Web Search · 2 (1 first)
YearPublicationVenuePosition
2025 BGM-HAN: A Hierarchical Attention Network for Accurate and Fair Decision Assessment on Semi-structured Profiles
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001
ASONAM (2)1
2025 Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001
ASONAM (2)1
2025 From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
Junhua Liu 0002, Tan Yong Keat, Kwan Hui Lim 0001
CIKM1
2023 Photozilla: An Image Dataset of Photography Styles and its Application to Visual Embedding and Style Detection
abstract
The widespread sharing of digital photography and images have led to the rapid development of various vision-related applications, such as photography style detection. Towards this effort, we introduce a photography style dataset termed Photozilla, which comprises over 990k images belonging to 10 different photographic styles. We used Photozilla to train 3 classification models for categorizing images into the relevant style and achieve an accuracy of ~96%. To better detect new photography styles that are constantly emerging, we also present a Siamese-based network that uses the trained classification models as the base architecture to adapt and classify unseen styles with only 25 training samples. Experiment results show an accuracy of over 68% in terms of identifying 10 additional distinct categories of photography styles. This dataset can be found at https://trisha025.github.io/Photozilla/.
Trisha Singhal, Junhua Liu 0002, Wenchuan Mu, Luciënne T. M. Blessing, Kwan Hui Lim 0001
ASONAM2
2023 A Transformer-Based Framework for POI-Level Social Post Geolocation
Kwan Hui Lim 0001, Teng Guo 0002, Junhua Liu 0002
ECIR (1)4
2021 Analyzing Scientific Publications using Domain-Specific Word Embedding and Topic Modelling
abstract
The scientific world is changing a tarapid pace, with new technology being developed and new trends being set at an increasing frequency. This paper presents a framework for conducting scientific analyses of academic publications, which is crucial to monitor research trends and identify potential innovations. This framework adopts and combines various techniques of Natural Language Processing, such as word embedding and topic modelling. Word embedding is used to capture semantic meanings of domain-specific words. We propose two novel scientific publication embedding, i.e., P UB-G and P UB-W, which are capable of learning semantic meanings of general as well as domain-specific words in various research fields. Thereafter, topic modelling is used to identify clusters of research topics within these larger research fields. We curated apublication dataset consisting of two conferences and two journals from 1995 to 2020 from two research domains. Experimental results show that our PUB-G and PUB-W embeddings are superior in comparison to other baseline embeddings by a margin of ~0.18-1.03 based on topic coherence.
Trisha Singhal, Junhua Liu 0002, Luciënne T. M. Blessing, Kwan Hui Lim 0001
IEEE BigData2
2020 Urban Crowdsensing using Social Media: An Empirical Study on Transformer and Recurrent Neural Networks
abstract
An important aspect of urban planning is understanding crowd levels at various locations, which typically require the use of physical sensors. Such sensors are potentially costly and time consuming to implement on a large scale. To address this issue, we utilize publicly available social media datasets and use them as the basis for two urban sensing problems, namely event detection and crowd level prediction. One main contribution of this work is our collected dataset from Twitter and Flickr, alongside ground truth events. We demonstrate the usefulness of this dataset with two preliminary supervised learning approaches: firstly, a series of neural network models to determine if a social media post is related to an event and secondly a regression model using social media post counts to predict actual crowd levels. We discuss preliminary results from these tasks and highlight some challenges.
Jerome Heng, Junhua Liu 0002, Kwan Hui Lim 0001
IEEE BigData2
2020 EPIC30M: An Epidemics Corpus of Over 30 Million Relevant Tweets
abstract
Since the start of COVID-19, there has been several relevant corpora from various sources that were released to support research in this area. While these corpora are valuable in supporting analysis for this specific pandemic, researchers will benefit from additional benchmark corpora that contain other epidemics for better generalizability and to facilitate cross-epidemic pattern recognition and trend analysis tasks. During our research, we discover little disease related corpora in the literature that are sizable and rich enough to support such cross-epidemic analysis tasks. To address this issue, we present EPIC30M, a large-scale epidemic corpus that contains more than 30 million micro-blog posts, i.e., tweets crawled from Twitter, from year 2006 to 2020. EPIC30M contains a subset of 26.2 million tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 million tweets of six global epidemic outbreaks, including the 2009 H1N1 Swine Flu, 2010 Haiti Cholera, 2012 Middle-East Respiratory Syndrome (MERS), 2013 West African Ebola, 2016 Yemen Cholera and 2018 Kivu Ebola. Furthermore, we explore and discuss the properties of this corpus with statistics of key terms and hashtags and trends analysis for each subset. Finally, we discuss the potential value and impact that EPIC30M could generate through a discussion of multiple use cases of cross-epidemic research topics that attract growing interest in recent years. These use cases span multiple research areas, such as epidemiological modeling, pattern recognition, natural language understanding and economical modeling. The corpus is publicly available at https://www.github.com/junhua/epic.
Junhua Liu 0002, Trisha Singhal, Luciënne T. M. Blessing, Kristin L. Wood, Kwan Hui Lim 0001
IEEE BigData1
2020 Strategic and Crowd-Aware Itinerary Recommendation
Junhua Liu 0002, Kristin L. Wood, Kwan Hui Lim 0001
ECML/PKDD (4)1