Wei Liu 0006

dblp:49/3283-6 · also Wei Vivian Liu · DBLP profile ↗
← Back
51ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0002-7409-0948ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 19 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2026 LMS-Retrieval: Layout-Aware, Modality-Aware, Structure-Aware Document Retrieval
Man Qin, Tim French 0002, Wei Liu 0006
ICDAR (3)3
2025 Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
abstract
Text-driven Human-Object Interaction (Text-to-HOI) generation is an emerging field with applications in animation, video games, virtual reality, and robotics. A key challenge in HOI generation is maintaining interaction consistency in long sequences. Existing Text-to-Motion-based approaches, such as discrete motion tokenization, cannot be directly applied to HOI generation due to limited data in this domain and the complexity of the modality. To address the problem of interaction consistency in long sequences, we propose an autoregressive diffusion model (ARDHOI) that predicts the next continuous token. Specifically, we introduce a Contrastive Variational Autoencoder (cVAE) to learn a physically plausible space of continuous HOI tokens, thereby ensuring that generated human-object motions are realistic and natural. For generating sequences autoregressively, we develop a Mamba-based context encoder to capture and maintain consistent sequential actions. Additionally, we implement an MLP-based denoiser to generate the subsequent token conditioned on the encoded context. Our model has been evaluated on the OMOMO and BEHAVE datasets, where it outperforms existing state-of-the-art methods in terms of both performance and inference speed. This makes ARDHOI a robust and efficient solution for text-driven HOI tasks.
Zichen Geng, Zeeshan Hayder, Wei Liu 0006, Ajmal Mian
AAAI3
2025 Differential Privacy on Large Language Models for Privacy Preserving Clinical Coding
abstract
Recent advancements in Large Language Models (LLMs) have significantly enhanced performance across various Natural Language Processing (NLP) tasks. In certain fields, particularly healthcare, the risk of data leakage in research data management is a critical concern when employing LLMs. To ensure data privacy, recent studies have adopted approaches, such as de-identification by masking out personal identifiable information. However, these anonymisation techniques remain vulnerable to various attacks, including linkage attacks, attribute inference attacks, and membership inference attacks. Differential privacy is a robust anonymisation technique that constrains the influence of individual data samples during model training to address data leakage. Nonetheless, the trade-off between utility and privacy protection remains challenging. Moreover, while differential privacy has been extensively studied in the context of tabular and image data, its application in NLP, especially with clinical data, is limited. In this paper, we explore the integration of differential privacy into the fine-tuning process of LLMs for clinical data, covering a range of model sizes and privacy standards within a healthcare context. We utilise these LLMs to generate synthetic medical notes and assess the privacy and utility of our differential privacy training approach by deploying these synthetic notes in a downstream clinical coding task. Our findings demonstrate that synthetic data from differential privacy-based LLMs achieve comparable or superior classification accuracy to non-differential privacy-based LLMs.
Ben Marshall, Shiv Akarsh Meka, Wei Liu 0006
IJCNN4
2025 DAG-Think-Twice: Causal Structure Guided Elicitation of Causal Reasoning in LLMs
Zheyuan Deng, Qiang Sun 0006, Jichunyang Li, Wei Liu 0006
PAKDD (7)4
2025 Graph-Based Multimodal Contrastive Learning for Chart Question Answering
abstract
Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode.This work introduces a novel joint multimodal scene graph framework that explicitly models the relationships among chart components and their underlying structures.The framework integrates both visual and textual graphs to capture structural and semantic characteristics, while a graph contrastive learning strategy aligns node representations across modalities-enabling their seamless incorporation into a transformer decoder as soft prompts.Moreover, a set of tailored Chain-of-Thought (CoT) prompts is proposed to enhance multimodal large language models (MLLMs) in zero-shot scenarios by mitigating hallucinations.Extensive evaluations on benchmarks including ChartQA, OpenCQA, and ChartX demonstrate significant performance improvements and validate the efficacy of the proposed approach.
Yue Dai 0006, Soyeon Caren Han, Wei Liu 0006
SIGIR3
2025 Spherical Embeddings for Atomic Relation Projection Reaching Complex Logical Query Answering
abstract
Projecting knowledge graph queries into an embedding space using geometric models (points, boxes and spheres) can help to answer queries for large incomplete knowledge graphs. In this work, we propose a symbolic learning-free approach using fuzzy logic to address the shape-closure problem that restricted geometric-based embedding models to only a few shapes (e.g. ConE) for answering complex logical queries. The use of symbolic approach facilitates non-closure geometric models (e.g. point, box) to handle logical operators (including negation). This enabled our newly proposed spherical embeddings (SpherE) in this work to use a polar coordinate system to effectively represent hierarchical relation. Results show that the SpherE model can answer existential positive first-order logic and negation queries. We show that SpherE significantly outperforms the point and box embeddings approaches while generating semantically meaningful hierarchy-aware embeddings.
Chau D. M. Nguyen, Tim French 0002, Michael Stewart 0006, Melinda R. Hodkiewicz, Wei Liu 0006
WWW5
2025 TriagedMSA: Triaging Sentimental Disagreement in Multimodal Sentiment Analysis
abstract
Existing multimodal sentiment analysis models are effective at capturing sentiment commonalities across different modalities and discerning emotions. However, these models still face significant challenges when analyzing samples with sentiment polarity differences across modalities. Neural networks struggle to process such divergent sentiment samples, particularly when they are scarce within datasets. While larger datasets could help address this limitation, collecting and annotating them is resource-intensive. To overcome this challenge, we proposeTriagedMSA, a multimodal sentiment analysis model with triage capability. Our model introduces theSentiment Disagreement Triage Network, which identifies sentiment disagreement between modalities within a sample. This triage mechanism reduces mutual influence by learning to distinguish between samples of sentiment agreement and disagreement. To process these two sample types, we develop theSentiment Selection Attention Networkand theSentiment Commonality Attention Network, both of which enhance modality interaction learning. Furthermore, we propose theAdaptive Polarity Detection (APD) algorithm, which ensures the generalizability of our model across different datasets, regardless of whether unimodal labels are available. The APD algorithm adaptively determines sentiment polarity disagreement or agreement between modalities. We conduct experiments on three multimodal sentiment analysis datasets:CMU-MOSI,CMU-MOSEIandCH-SIMS.v2. The results demonstrate that our proposed methodology outperforms existing state-of-the-art approaches.
Yuanyi Luo, Wei Liu 0006, Qiang Sun 0006, Jichunyang Li, Rui Wu 0002, Xianglong Tang
IEEE Trans. Affect. Comput.2
2024 MSG-Chart: Multimodal Scene Graph for ChartQA
abstract
Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in charts. To address this challenge, we design a joint multimodal scene graph for charts to explicitly represent the relationships between chart elements and their patterns. Our proposed multimodal scene graph includes a visual graph and a textual graph to jointly capture the structural and semantical knowledge from the chart. This graph module can be easily integrated with different vision transformers as inductive bias. Our experiments demonstrate that incorporating the proposed graph module enhances the understanding of charts' elements' structure and semantics, thereby improving performance on publicly available benchmarks, ChartQA and OpenCQA.
Yue Dai 0006, Soyeon Caren Han, Wei Liu 0006
CIKM3
2024 MaintIE: A Fine-Grained Annotation Schema and Benchmark for Information Extraction from Maintenance Short Texts
abstract
Maintenance short texts (MST), derived from maintenance work order records, encapsulate crucial information in a concise yet information-rich format. These user-generated technical texts provide critical insights into the state and maintenance activities of machines, infrastructure, and other engineered assets–pillars of the modern economy. Despite their importance for asset management decision-making, extracting and leveraging this information at scale remains a significant challenge. This paper presents MaintIE, a multi-level fine-grained annotation scheme for entity recognition and relation extraction, consisting of 5 top-level classes: PhysicalObject, State, Process, Activity and Property and 224 leaf entities, along with 6 relations tailored to MSTs. Using MaintIE, we have curated a multi-annotator, high-quality, fine-grained corpus of 1,076 annotated texts. Additionally, we present a coarse-grained corpus of 7,000 texts and consider its performance for bootstrapping and enhancing fine-grained information extraction. Using these corpora, we provide model performance measures for benchmarking automated entity recognition and relation extraction. The MaintIE scheme, corpus, and model are publicly available at https://github.com/nlp-tlp/maintie under the MIT license, encouraging further community exploration and innovation in extracting valuable insights from MSTs.
Tyler Bikaun, Tim French 0002, Michael Stewart 0006, Wei Liu 0006, Melinda R. Hodkiewicz
LREC/COLING4
2024 Are Graph Embeddings the Panacea? - An Empirical Survey from the Data Fitness Perspective
Qiang Sun 0006, Du Q. Huynh, Mark Reynolds 0001, Wei Liu 0006
PAKDD (2)4
2024 Spatio-Temporal Graph Representation Learning for Fraudster Group Detection
abstract
Motivated by potential financial gain, companies may hire fraudster groups to write fake reviews to either demote competitors or promote their own businesses. Such groups are considerably more successful in misleading customers, as people are more likely to be influenced by the opinion of a large group. To detect such groups, a common model is to represent fraudster groups' static networks, consequently overlooking the longitudinal behavior of a reviewer, thus, the dynamics of coreview relations among reviewers in a group. Hence, these approaches are incapable of excluding outlier reviewers, which are fraudsters intentionally camouflaging themselves in a group and genuine reviewers happen to coreview in fraudster groups. To address this issue, we propose "FGDT," a framework for "fraudster group detection through temporal relations." FGDT first capitalizes on the effectiveness of the HIN-recurrent neural network (RNN) in both reviewers' representation learning while capturing the collaboration between reviewers. The HIN-RNN models the coreview relations of reviewers in a group in a fixed time window of 28 days. We refer to this as spatial relation learning representation to signify the generalizability of this work to other networked scenarios. Then, we use an RNN on the spatial relations to predict the spatio-temporal relations of reviewers in the group. In the third step, a graph convolution network (GCN) refines the reviewers' vector representations using these predicted relations. These refined representations are then used to remove outlier reviewers. The average of the remaining reviewers' representation is then fed to a simple fully connected layer to predict if the group is a fraudster group or not. Exhaustive experiments of FGDT showed a 5% (4%), 12% (5%), and 12% (5%) improvement over three of the most recent approaches on precision, recall, and F1-value over the Yelp (Amazon) dataset, respectively.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Neural Networks Learn. Syst.3
2023 CylE: Cylinder Embeddings for Multi-hop Reasoning over Knowledge Graphs
abstract
Recent geometric-based approaches have been shown to efficiently model complex logical queries (including the intersection operation) over Knowledge Graphs based on the natural representation of Venn diagram.Existing geometric-based models (using points, boxes embeddings), however, cannot handle the logical negation operation.Further, those using cones embeddings are limited to representing queries by two-dimensional shapes, which reduced their effectiveness in capturing entities query relations for correct answers.To overcome this challenge, we propose unbounded cylinder embeddings (namely CylE), which is a novel geometric-based model based on threedimensional shapes.Our approach can handle a complete set of basic first-order logic operations (conjunctions, disjunctions and negations).CylE considers queries as Cartesian products of unbounded sector-cylinders and consider a set of nearest boxes corresponds to the set of answer entities.Precisely, the conjunctions can be represented via the intersections of unbounded sector-cylinders.Transforming queries to Disjunctive Normal Form can handle queries with disjunctions.The negations can be represented by considering the closure of complement for an arbitrary unbounded sector-cylinder.Empirical results show that the performance of multihop reasoning task using CylE significantly increases over state-of-the-art geometric-based query embedding models for queries without negation.For queries with negation operations, though the performance is on a par with the best performing geometric-based model, CylE significantly outperforms a recent distributionbased model.
Chau D. M. Nguyen, Tim French 0002, Wei Liu 0006, Michael Stewart 0006
EACL3
2023 Language Model Agnostic Gray-Box Adversarial Attack on Image Captioning
abstract
Adversarial susceptibility of neural image captioning is still under-explored due to the complex multi-model nature of the task. We introduce a GAN-based adversarial attack to effectively fool encoder-decoder based image captioning frameworks. Unique to our attack is the systematic disruption of the internal representation of an image at the encoder stage which allows control over the captions generated at the decoder stage. We cause the desired disruption with an input perturbation that promotes similarity between the features of the input image with a target image of our choice. The target image provides a convenient handle to control the incorrect captions in our method. We do not assume any knowledge of the decoder module, which makes our attack ‘gray-box’. Moreover, our attack remains agnostic to the type of decoder module, thereby proving effective for RNNs as well as Transformers as the language models. This makes our attack highly pragmatic.
Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Mubarak Shah, Ajmal Mian
IEEE Trans. Inf. Forensics Secur.3
2023 Dense Video Captioning With Early Linguistic Information Fusion
abstract
Dense captioning methods generally detect events in videos first and then generate captions for the individual events. Events are localized solely based on the visual cues while ignoring the associated linguistic information and context. Whereas end-to-end learning may implicitly take guidance from language, these methods still fall short of the power of explicit modeling. In this paper, we propose aVisual-Semantic Embedding (ViSE) Frameworkthat models the word(s)-context distributional properties over the entire semantic space and computes weights for all then-gramssuch that higher weights are assigned to the more informativen-grams. The weights are accounted for in learning distributed representations of all the captions to construct a semantic space. To perform the contextualization of visual information and the constructed semantic space in a supervised manner, we designVisual-Semantic Joint Modeling Network (VSJM-Net). The learnedViSEembeddings are then temporally encoded with aHierarchical Descriptor Transformer (HDT). For caption generation, we exploit a transformer architecture to decode the input embeddings into natural language descriptions. Experiments on the large-scale ActivityNet Captions dataset and YouCook-II dataset demonstrate the efficacy of our method.
Nayyer Aafaq, Ajmal Mian, Naveed Akhtar, Wei Liu 0006, Mubarak Shah
IEEE Trans. Multim.4
2023 HIN-RNN: A Graph Representation Learning Neural Network for Fraudster Group Detection With No Handcrafted Features
abstract
Social reviews are indispensable resources for modern consumers' decision making. For financial gain, companies pay fraudsters preferably in groups to demote or promote products and services since consumers are more likely to be misled by a large number of similar reviews from groups. Recent approaches on fraudster group detection employed handcrafted features of group behaviors without considering the semantic relation between reviews from the reviewers in a group. In this paper, we propose the first neural approach, HIN-RNN, a Heterogeneous Information Network (HIN) Compatible RNN for fraudster group detection that requires no handcrafted features. HIN-RNN provides a unifying architecture for representation learning of each reviewer, with the initial vector as the sum of word embeddings of all review text written by the same reviewer, concatenated by the ratio of negative reviews. Given a co-review network representing reviewers who have reviewed the same items with the same ratings and the reviewers' vector representation, a collaboration matrix is acquired through HIN-RNN training. The proposed approach is confirmed to be effective with marked improvement over state-of-the-art approaches on both the Yelp (22% and 12% in terms of recall and F1-value, respectively) and Amazon (4% and 2% in terms of recall and F1-value, respectively) datasets.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Neural Networks Learn. Syst.3
2022 E2EET: from pipeline to end-to-end entity typing via transformer-based embeddings
Michael Stewart 0006, Wei Liu 0006
Knowl. Inf. Syst.2
2022 ScoreGAN: A Fraud Review Detector Based on Regulated GAN With Data Augmentation
abstract
The promising performance of Deep Neural Networks (DNNs) in text classification has attracted researchers to use them for fraud review detection. However, the lack of trusted labeled data has limited the performance of the current solutions in detecting fraud reviews. The Generative Adversarial Network (GAN) as a semi-supervised method has been demonstrated to be effective for data augmentation purposes. The state-of-the-art solutions utilize GANs to overcome the data scarcity problem. However, they fail to incorporate the behavioral clues in fraud generation. Additionally, state-of-the-art approaches overlook the possible bot-generated reviews in the dataset. Finally, they also suffer from a common limitation in the generalization and stability of the GAN, slowing down the training procedure. In this work, we propose ScoreGAN for fraud review detection that makes use of both review text and review rating scores in the generation and detection process. Scores are incorporated through Information Gain Maximization (IGM) into the loss function for three reasons. One is to generate score-correlated reviews based on the scores given to the generator. Second, the generated reviews are employed to train the discriminator, allowing the discriminator to correctly label the possible bot-generated reviews through joint representations learned from the concatenation of GLobal Vector for Word representation (GLoVe) extracted from the text and the score. Finally, it can be used to improve the stability and generalization of the GAN. Results show that the proposed framework outperformed the existing state-of-the-art FakeGAN framework, in terms of AP by 7%, and 5% on the Yelp and TripAdvisor datasets, respectively.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Inf. Forensics Secur.3
2021 Adversarial Attacks and Defense on Deep Learning Classification Models using YCbCr Color Images
abstract
Deep neural network models are vulnerable to adversarial perturbations that are subtle but change the model predictions. Adversarial perturbations are generally computed for RGB images and are, hence, equally distributed among the RGB channels. We show, for the first time, that adversarial perturbations prevail in the Y-channel of the$\mathbf{YC}_{b}\mathbf{C}_{r}$> color space and exploit this finding to propose a defense mechanism. Our defense ResUpNet, which is end-to-end trainable, removes perturbations only from the Y-channel by exploiting ResNet features in a bottleneck free up-sampling framework. The refined Y-channel is combined with the untouched$\mathbf{C}_{b}\mathbf{C}_{r}$-channels to restore the clean image. We compare ResUpNet to existing defenses in the input transformation category and show that it achieves the best balance between maintaining the original accuracies on clean images and defense against adversarial attacks. Finally, we show that for the same attack and fixed perturbation magnitude, learning perturbations only in the Y-channel results in higher fooling rates. For example, with a very small perturbation magnitude$\epsilon=0.002$) the fooling rates of FGSM and PGD attacks on the ResNet50 model increase by 11.1% and 15.6% respectively, when the perturbations are learned only for the Y-channel.
Camilo Pestana, Naveed Akhtar, Wei Liu 0006, David G. Glance, Ajmal Mian
IJCNN3
2021 Defense-friendly Images in Adversarial Attacks: Dataset and Metrics for Perturbation Difficulty
abstract
Dataset bias is a problem in adversarial machine learning, especially in the evaluation of defenses. An adversarial attack or defense algorithm may show better results on the reported dataset than can be replicated on other datasets. Even when two algorithms are compared, their relative performance can vary depending on the dataset. Deep learning offers state-of-the-art solutions for image recognition, but deep models are vulnerable even to small perturbations. Research in this area focuses primarily on adversarial attacks and defense algorithms. In this paper, we report for the first time, a class of robust images that are both resilient to attacks and that recover better than random images under adversarial attacks using simple defense techniques. Thus, a test dataset with a high proportion of robust images gives a misleading impression about the performance of an adversarial attack or defense. We propose three metrics to determine the proportion of robust images in a dataset and provide scoring to determine the dataset bias. We also provide an ImageNet-R dataset of 15000+ robust images to facilitate further research on this intriguing phenomenon of image strength under attack. Our dataset, combined with the proposed metrics, is valuable for unbiased benchmarking of adversarial attack and defense algorithms.
Camilo Pestana, Wei Liu 0006, David G. Glance, Ajmal Mian
WACV2
2021 SubICap: Towards Subword-informed Image Captioning
Naeha Sharif, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah
WACV3
2021 Auto-labelling entities in low-resource text: a geological case study
Majigsuren Enkhsaikhan, Wei Liu 0006, Eun-Jung Holden, Paul Duuring
Knowl. Inf. Syst.2
2021 DFraud³: Multi-Component Fraud Detection Free of Cold-Start
abstract
Fraud review detection is a hot research topic in recent years. The Cold-start is a particularly new but significant problem referring to the failure of a detection system to recognize the authenticity of a new user. State-of-the-art solutions employ a translational knowledge graph embedding approach (TransE) to model the interaction of the components of a review system. However, these approaches suffer from the limitation of TransE in handling N-1 relations and the narrow scope of a single classification task, i.e., detecting fraudsters only. In this paper, we model a review system as a Heterogeneous Information Network (HIN) which enables a unique representation to every component and performs graph inductive learning on the review data through aggregating features of nearby nodes. HIN with graph induction helps to address the camouflage issue (fraudsters with genuine reviews) which has shown to be more severe when it is coupled with cold-start, i.e., new fraudsters with genuine first reviews. In this research, instead of focusing only on one component, detecting either fraud reviews or fraud users (fraudsters), vector representations are learned for each component, enabling multi-component classification. In other words, we can detect fraud reviews, fraudsters, and fraud-targeted items, thus the name of our approach DFraud3. DFraud3demonstrates a significant accuracy increase of 13% over the state of the art on Yelp.
Saeedreza Shehnepoor, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
IEEE Trans. Inf. Forensics Secur.3
2021 Deep fusion of multimodal features for social media retweet time prediction
Shuiqiao Yang, Wei Liu 0006, Jianxin Li 0001
World Wide Web4
2020 Seq2KG: An End-to-End Neural Model for Domain Agnostic Knowledge Graph (not Text Graph) Construction from Text
abstract
Knowledge Graph Construction (KGC) from text unlocks information held within unstructured text and is critical to a wide range of downstream applications. General approaches to KGC from text are heavily reliant on the existence of knowledge bases, yet most domains do not even have an external knowledge base readily available. In many situations this results in information loss as a wealth of key information is held within "non-entities". Domain-specific approaches to KGC typically adopt unsupervised pipelines, using carefully crafted linguistic and statistical patterns to extract co-occurred noun phrases as triples, essentially constructing text graphs rather than true knowledge graphs. In this research, for the first time, in the same flavour as Collobert et al.'s seminal work of "Natural language processing (almost) from scratch" in 2011, we propose a Seq2KG model attempting to achieve "Knowledge graph construction (almost) from scratch". An end-to-end Sequence to Knowledge Graph (Seq2KG) neural model jointly learns to generate triples and resolves entity types as a multi-label classification task through deep learning neural networks. In addition, a novel evaluation metric that takes both semantic and structural closeness into account is developed for measuring the performance of triple extraction. We show that our end-to-end Seq2KG model performs on par with a state of the art rule-based system which outperformed other neural models and won the first prize of the first Knowledge Graph Contest in 2019. A new annotation scheme and three high-quality manually annotated datasets are available to help promote this direction of research.
Michael Stewart 0006, Wei Liu 0006
KR2
2019 Enhanced Random Forest Algorithms for Partially Monotone Ordinal Classification
abstract
One of the factors hindering the use of classification models in decision making is that their predictions may contradict expectations. In domains such as finance and medicine, the ability to include knowledge of monotone (nondecreasing) relationships is sought after to increase accuracy and user satisfaction. As one of the most successful classifiers, attempts have been made to do so for Random Forest. Ideally a solution would (a) maximise accuracy; (b) have low complexity and scale well; (c) guarantee global monotonicity; and (d) cater for multi-class. This paper first reviews the state-of-theart from both the literature and statistical libraries, and identifies opportunities for improvement. A new rule-based method is then proposed, with a maximal accuracy variant and a faster approximate variant. Simulated and real datasets are then used to perform the most comprehensive ordinal classification benchmarking in the monotone forest literature. The proposed approaches are shown to reduce the bias induced by monotonisation and thereby improve accuracy.
Christopher Bartley, Wei Liu 0006, Mark Reynolds 0001
AAAI2
2019 Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning
abstract
Automatic generation of video captions is a fundamental challenge in computer vision. Recent techniques typically employ a combination of Convolutional Neural Networks (CNNs) and Recursive Neural Networks (RNNs) for video captioning. These methods mainly focus on tailoring sequence learning through RNNs for better caption generation, whereas off-the-shelf visual features are borrowed from CNNs. We argue that careful designing of visual features for this task is equally important, and present a visual feature encoding technique to generate semantically rich captions using Gated Recurrent Units (GRUs). Our method embeds rich temporal dynamics in visual features by hierarchically applying Short Fourier Transform to CNN features of the whole video. It additionally derives high level semantics from an object detector to enrich the representation with spatial dynamics of the detected objects. The final representation is projected to a compact space and fed to a language model. By learning a relatively simple language model comprising two GRU layers, we establish new state-of-the-art on MSVD and MSR-VTT datasets for METEOR and ROUGELmetrics.
Nayyer Aafaq, Naveed Akhtar, Wei Liu 0006, Syed Zulqarnain Gilani, Ajmal Mian
CVPR3
2019 ICDM 2019 Knowledge Graph Contest: Team UWA
abstract
We present an overview of our triple extraction system for the ICDM 2019 Knowledge Graph Contest. Our system uses a pipeline-based approach to extract a set of triples from a given document. It offers a simple and effective solution to the challenge of knowledge graph construction from domain-specific text. It also provides the facility to visualise useful information about each triple such as the degree, betweenness, structured relation type(s), and named entity types.
Michael Stewart 0006, Majigsuren Enkhsaikhan, Wei Liu 0006
ICDM3
2019 LCEval: Learned Composite Metric for Caption Evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu 0006, Syed Afaq Ali Shah
Int. J. Comput. Vis.4
2018 Towards Geological Knowledge Discovery Using Vector-Based Semantic Similarity
Majigsuren Enkhsaikhan, Wei Liu 0006, Eun-Jung Holden, Paul Duuring
ADMA2
2018 Finding Word Sense Embeddings of Known Meaning
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
CICLing (2)3
2018 A Novel Framework for Constructing Partially Monotone Rule Ensembles
abstract
In many machine learning applications there exists prior knowledge that the response variable should be non-decreasing in one or more of the features. For example, the chance of a tumour being malignant should not decrease with increasing diameter (all else being equal). While a number of classification algorithms make use of monotone knowledge, many are limited to full monotonicity (in all features). Taking inspiration from instance based classifiers, we present a framework for monotone additive rule ensembles that is the first to cater for partial monotonicity (in some features). We demonstrate it by developing a partially monotone instance based classifier based on L1 cones. Experiments show that the algorithm produces reasonable results on real data sets while ensuring perfect partial monotonicity.
Christopher Bartley, Wei Liu 0006, Mark Reynolds 0001
ICDE2
2018 Towards a multilayered permission-based access control for extending Android security
abstract
Summary This paper discusses security issues on the user equipment, which is the “last mile” of social networks. One of the main Achilles' heel of social networks is not the organization of networks themselves, but the user devices, typically Android ones. The existing system of privileges makes it easy to infiltrate the network via applications installed on users' devices. Conventional signature‐based and static analysis methods are vulnerable. Access to privacy‐ and security‐relevant parts of the application programming interface is controlled by the corresponding permission in a manifest file. While requesting access to permissions, it may offer opportunities to malicious codes, which will cause security issues. Few works among permission analysis, however, pay attention to the prevention of permission leakage on both hardware and software frameworks. In this paper we tackle the challenge of providing our multilayered permission‐based security extension scheme on Android platforms. We propose a usage and access control model and an effective method of preventing permission leakage based on ARM TrustZone security extension mechanism. In contrast to previous work, the proposed security architecture provides a permission‐based mandatory access control on Android middleware, Linux kernel, and hardware layers. The evaluation results demonstrate the effectiveness of the proposed scheme in mitigating permission leakage vulnerabilities.
Liehui Jiang, Wenzhi Chen, Hongqi He, Shuiqiao Yang, Wei Liu 0006
Concurr. Comput. Pract. Exp.7
2017 An Interactive Web-Based Toolset for Knowledge Discovery from Short Text Log Data
Michael Stewart 0006, Wei Liu 0006, Rachel Cardell-Oliver, Mark Griffin
ADMA2
2017 A Matrix-Vector Recurrent Unit Model for Capturing Compositional Semantics in Phrase Embeddings
abstract
The meaning of a multi-word phrase not only depends on the meaning of its constituent words, but also the rules of composing them to give the so-called compositional semantic. However, many deep learning models for learning compositional semantics target specific NLP tasks such as sentiment classification. Consequently, the word embeddings encode the lexical semantics, the weights of the networks are optimised for the classification task. Such models have no mechanisms to explicitly encode the compositional rules, and hence they are insufficient in capturing the semantics of phrases. We present a novel recurrent computational mechanism that specifically learns the compositionality by encoding the compositional rule of each word into a matrix. The network uses a recurrent architecture to capture the order of words for phrases with various lengths without requiring extra preprocessing such as part-of-speech tagging. The model is thoroughly evaluated on both supervised and unsupervised NLP tasks including phrase similarity, noun-modifier questions, sentiment distribution prediction, and domain specific term identification tasks. We demonstrate that our model consistently outperforms the LSTM and CNN deep learning models, simple algebraic compositions, and other popular baselines on different datasets.
Rui Wang 0116, Wei Liu 0006, Chris McDonald
CIKM2
2016 Temporal Interaction Biased Community Detection in Social Networks
Noha Alduaiji, Jianxin Li 0001, Amitava Datta, Xiaolu Lu 0002, Wei Liu 0006
ADMA5
2016 Effective Monotone Knowledge Integration in Kernel Support Vector Machines
Christopher Bartley, Wei Liu 0006, Mark Reynolds 0001
ADMA2
2016 Generating Bags of Words from the Sums of Their Word Embeddings
Lyndon White, Roberto Togneri, Wei Liu 0006, Mohammed Bennamoun
CICLing (1)3
2016 An incremental algorithm for discovering routine behaviours from smart meter data
Jin Wang 0002, Rachel Cardell-Oliver, Wei Liu 0006
Knowl. Based Syst.3
2015 An Investigation of Neural Embeddings for Coreference Resolution
Varun Godbole, Wei Liu 0006, Roberto Togneri
CICLing (1)2
2015 Efficient Discovery of Recurrent Routine Behaviours in Smart Meter Time Series by Growing Subsequences
Jin Wang 0002, Rachel Cardell-Oliver, Wei Liu 0006
PAKDD (2)3
2014 How Preprocessing Affects Unsupervised Keyphrase Extraction
Rui Wang 0116, Wei Liu 0006, Chris McDonald
CICLing (1)2
2014 Constructing Consumer-Oriented Medical Terminology from the Web A Supervised Classifier Ensemble Approach
Wei Liu 0006, Harrison J. Sweeney, Bo Chung, David G. Glance
PRICAI1
2011 An Investigation of Recursive Auto-associative Memory in Sentiment Detection
Saeed Danesh, Wei Liu 0006, Tim French 0002, Mark Reynolds 0001
ADMA (1)2
2010 An HMM-SVM-Based Automatic Image Annotation Approach
Yinjie Lei, Wilson Wong, Wei Liu 0006, Mohammed Bennamoun
ACCV (4)3
2009 Acquiring Semantic Relations Using the Web for Constructing Lightweight Ontologies
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun
PAKDD2
2009 Connecting the Real World with the Virtual World - Controlling AIBO through Second Life
Evan Wong, Wei Liu 0006
RoboCup2
2009 Introduction: Practical Cognitive Agents and Robots
Wei Liu 0006, Mary-Anne Williams
Auton. Agents Multi Agent Syst.2
2009 A probabilistic framework for automatic term recognition
abstract
Term recognition identifies domain-relevant terms which are essential for discovering domain concepts and for the construction of terminologies required by a wide range of natural language applications. Many techniques have been developed in an attem
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun
Intell. Data Anal.2
2008 Determining the Unithood of Word Sequences Using a Probabilistic Approach
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun
IJCNLP2
2007 Tree-Traversing Ant Algorithm for term clustering based on featureless similarities
Wilson Wong, Wei Liu 0006, Mohammed Bennamoun
Data Min. Knowl. Discov.2
2007 Internet collaboration and service composition as a loose form of teamwork
Lin Padgham, Wei Liu 0006
J. Netw. Comput. Appl.2