VLDB 2026 Research / reviewers in the wild / expert
Junhua Liu 0002
dblp:30/4261-2
· DBLP profile ↗
12ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-4477-7439ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing Accuracy and Efficiency in Multi-Turn Intent Classification for LLM-Powered Dialog Systems in ProductionabstractAccurate multi-turn intent classification is critical for advancing conversational AI systems but remains challenging due to limited datasets and complex contextual dependencies across dialogue turns. This paper presents two novel approaches leveraging Large Language Models (LLMs) to enhance scalability and reduce latency in production dialogue systems. First, we introduce Symbol Tuning, which simplifies intent labels to reduce task complexity and improve performance in multi-turn dialogues. Second, we propose Consistency-aware, Linguistics Adaptive Retrieval Augmentation (CLARA), a framework that employs LLMs for data augmentation and pseudo-labeling to generate synthetic multi-turn dialogues. These enriched datasets are used to fine-tune a small, efficient model suitable for deployment. Experiments on multilingual dialogue datasets show that our methods result in notable gains in both accuracy and resource efficiency, with improvements of 5.09% in classification accuracy, a 40% reduction in annotation costs, and effective deployment in low-resource multilingual industrial settings. Junhua Liu 0002, Tan Yong Keat, Kwan Hui Lim 0001 |
AAAI | 1 |
| 2026 | Physics-Informed Autonomous LLM Agents for Explainable Power Electronics Modulation DesignabstractLLM-based autonomous agents have recently shown strong capabilities in solving complex industrial design tasks. However, in domains aiming for carbon neutrality and high-performance renewable energy systems, current AI-assisted design automation methods face critical challenges in explainability, scalability, and practical usability. To address these limitations, we introduce PHIA (Physics-Informed Autonomous Agent), an LLM-driven system that automates modulation design for power converters in Power Electronics Systems with minimal human intervention. In contrast to traditional pipeline-based methods, PHIA incorporates an LLM-based planning module that interactively acquires and verifies design requirements via a user-friendly chat interface. This planner collaborates with physics-informed simulation and optimization components to autonomously generate and iteratively refine modulation designs. The interactive interface also supports interpretability by providing textual explanations and visual outputs throughout the design process. Experimental results show that PHIA reduces standard mean absolute error by 63.2% compared to the second-best benchmark and accelerates the overall design process by over 33 times. A user study involving 20 domain experts further confirms PHIA’s superior design efficiency and usability, highlighting its potential to transform industrial design workflows in power electronics. Junhua Liu 0002, Fanfan Lin, Shuai Zhao 0003, Kwan Hui Lim 0001 |
AAAI | 1 |
| 2025 | BGM-HAN: A Hierarchical Attention Network for Accurate and Fair Decision Assessment on Semi-structured Profiles
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001 |
ASONAM (2) | 1 |
| 2025 | Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001 |
ASONAM (2) | 1 |
| 2025 | From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
Junhua Liu 0002, Tan Yong Keat, Kwan Hui Lim 0001 |
CIKM | 1 |
| 2024 | Spatial-Temporal Graph Representation Learning for Tactical Networks Future State PredictionabstractResource allocation in tactical ad-hoc networks presents unique challenges due to their dynamic and multi-hop nature. Accurate prediction of future network connectivity is essential for effective resource allocation in such environments. In this paper, we introduce the Spatial-Temporal Graph Encoder-Decoder (STGED) framework for Tactical Communication Networks that leverages both spatial and temporal features of network states to learn latent tactical behaviors effectively. STGED hierarchically utilizes graph-based attention mechanism to spatially encode a series of communication network states, leverages a recurrent neural network to temporally encode the evolution of states, and a fully-connected feed-forward network to decode the connectivity in the future state. Through extensive experiments, we demonstrate that STGED consistently outperforms baseline models by large margins across different time-steps input, achieving an accuracy of up to 99.2% for the future state prediction task of tactical communication networks. Junhua Liu 0002, Justin Albrethsen, Lincoln Goh, David K. Y. Yau, Kwan Hui Lim 0001 |
IJCNN | 1 |
| 2023 | Photozilla: An Image Dataset of Photography Styles and its Application to Visual Embedding and Style DetectionabstractThe widespread sharing of digital photography and images have led to the rapid development of various vision-related applications, such as photography style detection. Towards this effort, we introduce a photography style dataset termed Photozilla, which comprises over 990k images belonging to 10 different photographic styles. We used Photozilla to train 3 classification models for categorizing images into the relevant style and achieve an accuracy of ~96%. To better detect new photography styles that are constantly emerging, we also present a Siamese-based network that uses the trained classification models as the base architecture to adapt and classify unseen styles with only 25 training samples. Experiment results show an accuracy of over 68% in terms of identifying 10 additional distinct categories of photography styles. This dataset can be found at https://trisha025.github.io/Photozilla/. Trisha Singhal, Junhua Liu 0002, Wenchuan Mu, Luciënne T. M. Blessing, Kwan Hui Lim 0001 |
ASONAM | 2 |
| 2023 | A Transformer-Based Framework for POI-Level Social Post Geolocation
Kwan Hui Lim 0001, Teng Guo 0002, Junhua Liu 0002 |
ECIR (1) | 4 |
| 2021 | Analyzing Scientific Publications using Domain-Specific Word Embedding and Topic ModellingabstractThe scientific world is changing a tarapid pace, with new technology being developed and new trends being set at an increasing frequency. This paper presents a framework for conducting scientific analyses of academic publications, which is crucial to monitor research trends and identify potential innovations. This framework adopts and combines various techniques of Natural Language Processing, such as word embedding and topic modelling. Word embedding is used to capture semantic meanings of domain-specific words. We propose two novel scientific publication embedding, i.e., P UB-G and P UB-W, which are capable of learning semantic meanings of general as well as domain-specific words in various research fields. Thereafter, topic modelling is used to identify clusters of research topics within these larger research fields. We curated apublication dataset consisting of two conferences and two journals from 1995 to 2020 from two research domains. Experimental results show that our PUB-G and PUB-W embeddings are superior in comparison to other baseline embeddings by a margin of ~0.18-1.03 based on topic coherence. Trisha Singhal, Junhua Liu 0002, Luciënne T. M. Blessing, Kwan Hui Lim 0001 |
IEEE BigData | 2 |
| 2020 | Urban Crowdsensing using Social Media: An Empirical Study on Transformer and Recurrent Neural NetworksabstractAn important aspect of urban planning is understanding crowd levels at various locations, which typically require the use of physical sensors. Such sensors are potentially costly and time consuming to implement on a large scale. To address this issue, we utilize publicly available social media datasets and use them as the basis for two urban sensing problems, namely event detection and crowd level prediction. One main contribution of this work is our collected dataset from Twitter and Flickr, alongside ground truth events. We demonstrate the usefulness of this dataset with two preliminary supervised learning approaches: firstly, a series of neural network models to determine if a social media post is related to an event and secondly a regression model using social media post counts to predict actual crowd levels. We discuss preliminary results from these tasks and highlight some challenges. Jerome Heng, Junhua Liu 0002, Kwan Hui Lim 0001 |
IEEE BigData | 2 |
| 2020 | EPIC30M: An Epidemics Corpus of Over 30 Million Relevant TweetsabstractSince the start of COVID-19, there has been several relevant corpora from various sources that were released to support research in this area. While these corpora are valuable in supporting analysis for this specific pandemic, researchers will benefit from additional benchmark corpora that contain other epidemics for better generalizability and to facilitate cross-epidemic pattern recognition and trend analysis tasks. During our research, we discover little disease related corpora in the literature that are sizable and rich enough to support such cross-epidemic analysis tasks. To address this issue, we present EPIC30M, a large-scale epidemic corpus that contains more than 30 million micro-blog posts, i.e., tweets crawled from Twitter, from year 2006 to 2020. EPIC30M contains a subset of 26.2 million tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 million tweets of six global epidemic outbreaks, including the 2009 H1N1 Swine Flu, 2010 Haiti Cholera, 2012 Middle-East Respiratory Syndrome (MERS), 2013 West African Ebola, 2016 Yemen Cholera and 2018 Kivu Ebola. Furthermore, we explore and discuss the properties of this corpus with statistics of key terms and hashtags and trends analysis for each subset. Finally, we discuss the potential value and impact that EPIC30M could generate through a discussion of multiple use cases of cross-epidemic research topics that attract growing interest in recent years. These use cases span multiple research areas, such as epidemiological modeling, pattern recognition, natural language understanding and economical modeling. The corpus is publicly available at https://www.github.com/junhua/epic. Junhua Liu 0002, Trisha Singhal, Luciënne T. M. Blessing, Kristin L. Wood, Kwan Hui Lim 0001 |
IEEE BigData | 1 |
| 2020 | Strategic and Crowd-Aware Itinerary Recommendation
Junhua Liu 0002, Kristin L. Wood, Kwan Hui Lim 0001 |
ECML/PKDD (4) | 1 |