EDBT 2026 Demo / reviewers in the wild / expert
Hui Wang 0001
dblp:39/721-1
· DBLP profile ↗
37ranked-venue papers in the field
5as first author
7since 2021 · last 2026
0000-0003-2633-6015ORCID · conflict
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 12 (1 first)Information Retrieval & Web Search · 10 (1 first)Database Systems & Data Management · 8 (1 first)Other / Interdisciplinary · 4Data Mining & Knowledge Discovery · 2 (2 first)Business Process & Enterprise Data · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting market manipulation with dual-branch self-supervised learning: A unified framework integrating frequency-informed anomaly synthesis and domain-specific featuresabstractEffective detection of financial market manipulation is critically impeded by three fundamental challenges: signal concealment, data sparsity, the boundary vagueness. This paper introduces SD-FMM , a S elf-supervised D etection framework tailored for F inancial M arket M anipulation that addresses these fundamental challenges through three innovative components. First, our Amplification Component extracts and fuses domain-specific features grounded in market microstructure theory, substantially amplifying subtle manipulation signals that would otherwise remain concealed. Second, our Synthesis Component generates realistic synthetic anomalies through few-shot learning and dynamic frequency analysis using Discrete Wavelet Transform, enabling self-supervised training without relying on scarce labeled data. Third, our Detection Component employs a novel Dual-branch Contrastive Detection Neural Network that enhances sensitivity to manipulation boundaries through local contrastive learning and holistic modeling of temporal dependency. We evaluate SD-FMM using a newly collected proprietary dataset of 25 Chinese stock market manipulation cases and a public benchmark of 338 cryptocurrency pump-and-dump schemes. Extensive experiments against 12 state-of-the-art baselines demonstrate the significant superiority of SD-FMM. On the stock dataset, our method outperforms the second-best baseline by 47.61% in average precision metrics and reduces the false alarm rate by 47.46%. Meanwhile, it shortens the mean detection delay by 25.05%, enabling swift regulatory intervention. On the cryptocurrency dataset, SD-FMM exhibits remarkable sensitivity, achieving a Hit Rate@3 of 83.13% and Hit Rate@20 of 97.93%. Overall, our framework offers a generalized solution that can not only accurately distinguish manipulations from normal trading but also deliver a faster and stronger response to manipulations across diverse financial markets. Yongsheng Dai, Barry Quinn 0003, Fearghal Kearney, Ivor T. A. Spence, Karen Rafferty, Hui Wang 0001 |
Inf. Process. Manag. | 7 |
| 2025 | OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringabstractReliable factuality evaluation is critical for the iterative development of open-domain question answering (ODQA) systems, especially given the rise of large language models (LLMs) and their propensity for hallucination. However, state-of-the-art (SOTA) automatic metrics, which are mostly supervised, remain notably less reliable than humans. In this paper, we find two key challenges behind this gap: (1) length distribution mismatch between lengthy LLM answers and shorter training answers used by current metrics; and (2) reference incompleteness, where current metrics often misjudge valid system answers absent from given references-a challenge worsened by the diversity of LLM outputs. To address these issues, we present a new ODQA factuality evaluation dataset called OBELLA (Open-Book Evaluation for Long-form LLM Answers). OBELLA narrows the length distribution mismatch by significantly increasing the candidate answer length to align with LLM outputs. Moreover, it introduces a neutral class for plausible yet under-supported candidate answers to differentiate reference incompleteness from outright incorrectness, thus enabling flexible reevaluation by consulting external knowledge for more references. Based on OBELLA, we propose a novel metric named OBELLAM (OBELLA Metric). OBELLAM integrates a cross-attention mechanism to enhance long-form candidate answer representations and employs a dynamic closed-open book evaluation strategy to tackle reference incompleteness. Our OBELLAM sets a new SOTA in aligning with human judgments across two ODQA evaluation benchmarks, marking a promising step toward more robust ODQA factuality evaluation. Zhaoyu Zhang 0001, Hui Wang 0001, Karen Rafferty |
SIGIR | 3 |
| 2025 | A Survey of Change Point Detection in Dynamic GraphsabstractChange point detection is crucial for identifying state transitions and anomalies in dynamic systems, with applications in network security, health care, and social network analysis. Dynamic systems are represented by dynamic graphs with spatial and temporal dimensions. As objects and their relations in a dynamic graph change over time, detecting these changes is essential. Numerous methods for change point detection in dynamic graphs have been developed, but no systematic review exists. This paper addresses this gap by introducing change point detection tasks in dynamic graphs, discussing two tasks based on input data types: detection in graph snapshot series (focusing on graph topology changes) and time series on graphs (focusing on changes in graph entities with temporal dynamics). We then present related challenges and applications, provide a comprehensive taxonomy of surveyed methods, including datasets and evaluation metrics, and discuss promising research directions. Shang Gao 0005, Dandan Guo, Xiaohui Wei 0002, Jon G. Rokne, Hui Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | MVRMLM 2024: Multimodal Video Retrieval and Multimodal Language ModellingabstractAs the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example-based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi-modal content is essential. In designing our novel video retrieval framework named Multi-modal Video Search by Examples (MVSE)1, we focused on accuracy (precision and recall), efficiency (retrieval time in seconds), interactivity, and extensibility, with key components including advanced data processing and a user-friendly interface aimed at enhancing search effectiveness and user experience. With the advent of Large Language Models (LLMs), the interaction between multimodal data, including image and audio has been transformed with a significant leap forward towards a bigger goal of artificial general intelligence. This workshop aims to bring together experts from diverse domains to explore the possibilities of developing novel ways of multimodal data search, understanding and interaction. Hui Wang 0001, Josef Kittler, Mark J. F. Gales, Rob Cooper, Maurice D. Mulvenna, Wing W. Y. Ng, Yang Hua 0001, Richard Gault, Abbas Haider, Guanfeng Wu |
ICMR | 1 |
| 2024 | Optimal granularity selection based on algorithm stability with application to attribute reduction in rough set theory
Yue Gao 0016, Degang Chen 0002, Hui Wang 0001 |
Inf. Sci. | 3 |
| 2023 | Cluster-based data relabelling for classification
Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001, Xin Wei 0002 |
Inf. Sci. | 2 |
| 2022 | A dynamic rule-based classification model via granular computing
Jiaojiao Niu, Degang Chen 0002, Jinhai Li 0001, Hui Wang 0001 |
Inf. Sci. | 4 |
| 2020 | Ontology-based enriched concept graphs for medical document classification
Niloofer Shanavas, Hui Wang 0001, Zhiwei Lin 0002, Glenn I. Hawe |
Inf. Sci. | 2 |
| 2020 | Quantifying consensus of rankings based on q-support patterns
Zhengui Xue, Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
Inf. Sci. | 3 |
| 2019 | Structure-Based Supervised Term Weighting and Regularization for Text Classification
Niloofer Shanavas, Hui Wang 0001, Zhiwei Lin 0002, Glenn I. Hawe |
NLDB | 2 |
| 2019 | Multi-Level Matching Networks for Text MatchingabstractText matching aims to establish the matching relationship between two texts. It is an important operation in some information retrieval related tasks such as question duplicate detection, question answering, and dialog systems. Bidirectional long short term memory (BiLSTM) coupled with attention mechanism has achieved state-of-the-art performance in text matching. A major limitation of existing works is that only high level contextualized word representations are utilized to obtain word level matching results without considering other levels of word representations, thus resulting in incorrect matching decisions for cases where two words with different meanings are very close in high level contextualized word representation space. Therefore, instead of making decisions utilizing single level word representations, a multi-level matching network (MMN) is proposed in this paper for text matching, which utilizes multiple levels of word representations to obtain multiple word level matching results for final text level matching decision. Experimental results on two widely used benchmarks, SNLI and Scaitail, show that the proposed MMN achieves the state-of-the-art performance. Chunlin Xu, Zhiwei Lin 0002, Shengli Wu 0001, Hui Wang 0001 |
SIGIR | 4 |
| 2018 | Case-Based Decision Support System with Contextual Bandits Learning for Similarity Retrieval Model Selection
Booma Devi Sekar, Hui Wang 0001 |
KSEM (1) | 2 |
| 2015 | A New Dynamic Rule Activation Method for Extended Belief Rule-Based SystemsabstractData incompleteness and inconsistency are common issues in data-driven decision models. To some extend, they can be considered as two opposite circumstances, since the former occurs due to lack of information and the latter can be regarded as an excess of heterogeneous information. Although these issues often contribute to a decrease in the accuracy of the model, most modeling approaches lack of mechanisms to address them. This research focuses on an advanced belief rule-based decision model and proposes a dynamic rule activation (DRA) method to address both issues simultaneously. DRA is based on “smart” rule activation, where the actived rules are selected in a dynamic way to search for a balance between the incompleteness and inconsistency in the rule-base generated from sample data to achive a better performance. A series of case studies demonstrate how the use of DRA improves the accuracy of this advanced rule-based decision model, without compromising its efficiency, especially when dealing with multi-class classification datasets. DRA has been proved to be beneficial to select the most suitable rules or data instances instead of aggregating an entire rule-base. Beside the work performed in rule-based systems, DRA alone can be regarded as a generic dynamic similarity measurement that can be applied in different domains. Alberto Calzada, Jun Liu 0001, Hui Wang 0001, Anil Kashyap |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | An investigation into the application of ensemble learning for entailment classification
Niall Rooney, Hui Wang 0001, Philip S. Taylor |
Inf. Process. Manag. | 2 |
| 2014 | A linguistic multi-criteria decision making approach based on logical reasoning
Shuwei Chen 0001, Jun Liu 0001, Hui Wang 0001, Yang Xu 0001, Juan Carlos Augusto |
Inf. Sci. | 3 |
| 2013 | Interactive surveillance event detection at TRECVid2012abstractThis demonstration shows the integration of video analysis and search tools to facilitate the interactive retrieval of video segments depicting specific activities from surveillance footage. The implementation was developed by members of the SAVASA project for participation in the interactive surveillance event detection (SED) task of TRECVid 2012. This year, for the first time, the purpose of the interactive SED task was to evaluate systems' ability to support users in identifying video segments that depict a specific activity (event) in a large collection of surveillance video footage. Project partners worked together to analyse video and provide a query interface enabling users to search and identify matching video segments. The collaborative integration of components from multiple partners and the participation of end user partners in evaluating the system are the novel aspects of this work. Suzanne Little, Iveel Jargalsaikhan, Kathy M. Clawson, Marcos Nieto Doncel, Cem Direkoglu, Noel E. O'Connor, Alan F. Smeaton, Jun Liu 0001, Bryan W. Scotney, Hui Wang 0001, Seán Gaines, Aitor Rodriguez, Pedro J. Sánchez, Ana Martínez Llorens, Karina Villarroel Paniza, Roberto Gimenez, Raúl Santos de la Cámara, Anna Mereu, Celso Prados, Emmanouil Kafetzakis |
ICMR | 11 |
| 2013 | An information retrieval approach to identifying infrequent events in surveillance videoabstractThis paper presents work on integrating multiple computer vision-based approaches to surveillance video analysis to support user retrieval of video segments showing human activities. Applied computer vision using real-world surveillance video data is an extremely challenging research problem, independently of any information retrieval (IR) issues. Here we describe the issues faced in developing both generic and specific analysis tools and how they were integrated for use in the new TRECVid interactive surveillance event detection task. We present an interaction paradigm and discuss the outcomes from face-to-face end user trials and the resulting feedback on the system from both professionals, who manage surveillance video, and computer vision or machine learning experts. We propose an information retrieval approach to finding events in surveillance video rather than solely relying on traditional annotation using specifically trained classifiers. Suzanne Little, Iveel Jargalsaikhan, Kathy M. Clawson, Marcos Nieto Doncel, Cem Direkoglu, Noel E. O'Connor, Alan F. Smeaton, Bryan W. Scotney, Hui Wang 0001, Jun Liu 0001 |
ICMR | 10 |
| 2012 | Application of Evidence Theory and Discounting Techniques to Aerospace Design
Fiona Browne, David A. Bell, Weiru Liu, Yan Jin 0009, Colm Higgins, Niall Rooney, Hui Wang 0001, Jann Müller |
IPMU (3) | 7 |
| 2012 | Extended Twofold-LDA Model for Two Aspects in One Sentence
Nicola Burns, Yaxin Bi, Hui Wang 0001, Terry J. Anderson |
IPMU (2) | 3 |
| 2012 | A Knowledge-Driven Approach to Activity Recognition in Smart HomesabstractThis paper introduces a knowledge-driven approach to real-time, continuous activity recognition based on multisensor data streams in smart homes. The approach goes beyond the traditional data-centric methods for activity recognition in three ways. First, it makes extensive use of domain knowledge in the life cycle of activity recognition. Second, it uses ontologies for explicit context and activity modeling and representation. Third and finally, it exploits semantic reasoning and classification for activity inferencing, thus enabling both coarse-grained and fine-grained activity recognition. In this paper, we analyze the characteristics of smart homes and Activities of Daily Living (ADL) upon which we built both context and ADL ontologies. We present a generic system architecture for the proposed knowledge-driven approach and describe the underlying ontology-based recognition process. Special emphasis is placed on semantic subsumption reasoning algorithms for activity recognition. The proposed approach has been implemented in a function-rich software system, which was deployed in a smart home research laboratory. We evaluated the proposed approach and the developed system through extensive experiments involving a number of various ADL use scenarios. An average activity recognition rate of 94.44 percent was achieved and the average recognition runtime per recognition operation was measured as 2.5 seconds. Liming Chen 0001, Chris D. Nugent, Hui Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | A Multidimensional Sequence Approach to Measuring Tree SimilarityabstractTree is one of the most common and well-studied data structures in computer science. Measuring the similarity of such structures is key to analyzing this type of data. However, measuring tree similarity is not trivial due to the inherent complexity of trees and the ensuing large search space. Tree kernel, a state of the art similarity measurement of trees, represents trees as vectors in a feature space and measures similarity in this space. When different features are used, different algorithms are required. Tree edit distance is another widely used similarity measurement of trees. It measures similarity through edit operations needed to transform one tree to another. Without any restrictions on edit operations, the computation cost is too high to be applicable to large volume of data. To improve efficiency of tree edit distance, some approximations were introduced into tree edit distance. However, their effectiveness can be compromised. In this paper, a novel approach to measuring tree similarity is presented. Trees are represented as multidimensional sequences and their similarity is measured on the basis of their sequence representations. Multidimensional sequences have their sequential dimensions and spatial dimensions. We measure the sequential similarity by the all common subsequences sequence similarity measurement or the longest common subsequence measurement, and measure the spatial similarity by dynamic time warping. Then we combine them to give a measure of tree similarity. A brute force algorithm to calculate the similarity will have high computational cost. In the spirit of dynamic programming two efficient algorithms are designed for calculating the similarity, which have quadratic time complexity. The new measurements are evaluated in terms of classification accuracy in two popular classifiers (k-nearest neighbor and support vector machine) and in terms of search effectiveness and efficiency in k-nearest neighbor similarity search, using three different data sets from natural language processing and information retrieval. Experimental results show that the new measurements outperform the benchmark measures consistently and significantly. Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | A Twofold-LDA Model for Customer Review AnalysisabstractThe Latent Dirichlet Allocation model is an unsupervised generative model that is widely used for topic modelling in text. We propose to add supervision to the model in the form of domain knowledge to direct the focus of topics to more relevant aspects than the topics produced by standard LDA. Experimental results demonstrate the effectiveness of our method. We also propose a novel Twofold-LDA model to improve the current output of LDA in order to visualize results in graphical form, which can ultimately be used by potential customers. Experiments show the benefit of this new output, with the ability to produce topics focused on our desired aspects in a user friendly chart. Nicola Burns, Yaxin Bi, Hui Wang 0001, Terry J. Anderson |
Web Intelligence | 3 |
| 2011 | Concordance and consensus
Cees H. Elzinga, Hui Wang 0001, Zhiwei Lin 0002 |
Inf. Sci. | 2 |
| 2010 | Weighting common syntactic structures for natural language based information retrievalabstractNatural Language Processing (NLP) techniques are believed to hold the potential to assist "bag-of-words" Information Retrieval (IR) in terms of retrieval accuracy. In this paper, we report a natural language based IR approach where the common syntactic structures between documents and the query is regarded to as a query-dependent feature for documents. Specifically, a "structural weight" is proposed for query terms, which can be seen as a weight to model the degree of term's involvement in the common syntactic structures. This structural weight is used together with the TF-IDF weighting scheme, which results in a new ranking function. The accumulation of this structural weight of all the query terms in the new ranking function will be seen as a measure of how much a document and a query share the common syntactic structures. The experimental results show that by using this ranking function, significant improvements in the retrieval performance are achieved. Hui Wang 0001, Sally I. McClean, Epaminondas Kapetanios, Denis Carroll |
CIKM | 2 |
| 2010 | Facilitating Experience Reuse: Towards a Task-Based Approach
Liming Chen 0001, David Patterson 0002, Hui Wang 0001 |
KSEM | 5 |
| 2010 | Behavioural Rule Discovery from Swarm Systems
David Stoops, Hui Wang 0001, George Moore, Yaxin Bi |
KSEM | 2 |
| 2010 | Measuring Tree Similarity for Natural Language Processing Based Information Retrieval
Zhiwei Lin 0002, Hui Wang 0001, Sally I. McClean |
NLDB | 2 |
| 2010 | Mass function derivation and combination in multivariate data spaces
Hui Wang 0001, Jun Liu 0001, Juan Carlos Augusto |
Inf. Sci. | 1 |
| 2009 | Data Driven Rank Ordering and Its Application to Financial Portfolio Construction
Maria Dobrska, Hui Wang 0001, William Blackburn |
KSEM | 2 |
| 2008 | A Study of the Neighborhood Counting SimilarityabstractThe neighborhood counting measure is the number of all common neighborhoods between a pair of data points. It can be used as a similarity measure for different types of data through the notion of neighborhood: multivariate, sequence, and tree data. It has been shown that this measure is closely related to a secondary probability G, which is defined in terms of a primary probability P of interest to a problem. It has also been shown that the G probability can be estimated by aggregating neighborhood counts. The following questions can be asked: What is the relationship between this similarity measure and the primary probability P, especially for the task of classification? How does this similarity measure compare with the euclidean distance, since they are directly comparable? How does the G probability estimation compare with the popular kernel density estimation for the task of classification? These questions are answered in this paper, some theoretically and some experimentally. It is shown that G is a linear function of P and, therefore, a G-based Bayes classifier is equivalent to a P-based Bayes classifier. It is also shown that a weighted k-nearest neighbor classifier equipped with the neighborhood counting measure is, in fact, an approximation of the G-based Bayes classifier. It is further shown that the G probability leads to a probability estimator similar in spirit to the kernel density estimator. New experimental results are presented in this paper, which show that this measure compares favorably with the euclidean distance not only on multivariate data but also on time-series data. New experimental results are also presented regarding probability/density estimation. It was found that the G probability estimation can outperform the kernel density estimation in classification tasks. Hui Wang 0001, Fionn Murtagh |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Influences of Functional Dependencies on Bucket-Based Rewriting Algorithms
Qingyuan Bai, Jun Hong 0001, Hui Wang 0001, Michael F. McTear |
WAIM | 3 |
| 2004 | Classification Decision Combination for Text Categorization: An Experimental Study
Yaxin Bi, David A. Bell, Hui Wang 0001, Gongde Guo, Werner Dubitzky |
DEXA | 3 |
| 2004 | Contextual Probability-Based Classification
Gongde Guo, Hui Wang 0001, David A. Bell, Zhining Liao |
ER | 2 |
| 2004 | Bucket-Based Query Rewriting with Disjunctive Data SourceabstractMany algorithms for query rewriting using views have been proposed. Most of them are used for rewriting conjunctive queries with conjunctive views only. Only one inverse rule-based rewriting algorithm has been presented to deal with query rewriting with disjunctive views, in which a set of disjunctive inverse rules are created and used to generate the query rewritings. However, there has no been bucket-based algorithm for this issue. In this paper we apply the ideas of the previous bucket-based algorithms to deal with query rewriting using views in the presence of disjunctions in view definitions. We create a set of buckets over either a conjunctive view or a disjunctive view. Specially, when a bucket is over a disjunctive view, we present a method to remove the tuples that are in the view but not required by a given query. The rewritings obtained by our algorithm are contained in the original query. Qingyuan Bai, Jun Hong 0001, Michael F. McTear, Hui Wang 0001 |
Web Intelligence | 4 |
| 2002 | Data Reduction and Noise Filtering for Predicting Times Series
Gongde Guo, Hui Wang 0001, David A. Bell |
WAIM | 2 |
| 2001 | Classification through Maximizing DensityabstractThis paper presents a novel method for classification, which makes use of models built by the lattice machine (LM). The LM approximates data resulting in, as a model of data, a set of hyper tuples that are equilabelled, supported and maximal. The method presented uses the LM model of data to classify new data with a view to maximising the density of the model. Experiments show that this method, when used with the LM, outperforms the C2 algorithm and is comparable to the C5.0 classification algorithm. Hui Wang 0001, Ivo Düntsch, David A. Bell, Dayou Liu |
ICDM | 1 |
| 1998 | Data Reduction Based on Hyper Relations
Hui Wang 0001, Ivo Düntsch, David A. Bell |
KDD | 1 |