EDBT 2026 Demo / reviewers in the wild / expert
Tiejun Ma
dblp:61/3846
· DBLP profile ↗
22ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 5 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Security and privacy · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?abstractLarge Language Models (LLMs) have recently been leveraged for asset pricing tasks and stock trading applications, enabling AI agents to generate investment decisions from unstructured financial data. However, most evaluations of LLM timing-based investing strategies are conducted on narrow timeframes and limited stock universes, overstating effectiveness due to survivorship and data-snooping biases. We critically assess their generalizability and robustness by proposing FINSABER, a backtesting framework evaluating timing-based strategies across longer periods and a larger universe of symbols. Systematic backtests over two decades and 100+ symbols reveal that previously reported LLM advantages deteriorate significantly under broader cross-section and over a longer-term evaluation. Our market regime analysis further demonstrates that LLM strategies are overly conservative in bull markets, underperforming passive benchmarks, and overly aggressive in bear markets, incurring heavy losses. These findings highlight the need to develop LLM strategies that are able to prioritise trend detection and regime-aware risk controls over mere scaling of framework complexity. Weixian Waylon Li, Hyeonjun Kim, Mihai Cucuringu, Tiejun Ma |
KDD (1) | 4 |
| 2026 | Learn to Rank Risky Investors: A Case Study of Predicting Retail Traders' Behaviour and ProfitabilityabstractIdentifying risky traders with high profits in financial markets is crucial for market makers, such as trading exchanges, to ensure effective risk management through real-time decisions on regulation compliance and hedging. However, capturing the complex and dynamic behaviours of individual traders poses significant challenges. Traditional classification and anomaly detection methods often establish a fixed risk boundary, failing to account for this complexity and dynamism. To tackle this issue, we propose a profit-aware risk ranker (PA-RiskRanker) that reframes the problem of identifying risky traders as a ranking task using LEarning-TO-Rank (LETOR) algorithms. Our approach features a Profit-Aware binary cross-entropy (PA-BCE) loss function and a transformer-based ranker enhanced with a Self-Cross-Trader Attention pipeline. These components effectively integrate profit and loss (P&L) considerations into the training process while capturing intra- and inter-trader relationships. Our research critically examines the limitations of existing deep learning-based LETOR algorithms in trading risk management, which often overlook the importance of P&L in financial scenarios. By prioritising P&L, our method improves risky trader identification, achieving an 8.4% increase in F1 score compared to state-of-the-art (SOTA) ranking models like Rankformer. Additionally, it demonstrates a 10%–17% increase in average profit compared to all benchmark models. Weixian Waylon Li, Tiejun Ma |
ACM Trans. Inf. Syst. | 2 |
| 2025 | TSPRank: Bridging Pairwise and Listwise Methods with a Bilinear Travelling Salesman ModelabstractTraditional Learning-To-Rank (LETOR) approaches, including pairwise methods like RankNet and LambdaMART, often fall short by solely focusing on pairwise comparisons, leading to sub-optimal global rankings. Conversely, deep learning based listwise methods, while aiming to optimise entire lists, require complex tuning and yield only marginal improvements over robust pairwise models. To overcome these limitations, we introduce Travelling Salesman Problem Rank (TSPRank), a hybrid pairwise-listwise ranking method. TSPRank reframes the ranking problem as a Travelling Salesman Problem (TSP), a well-known combinatorial optimisation challenge that has been extensively studied for its numerous solution algorithms and applications. This approach enables the modelling of pairwise relationships and leverages combinatorial optimisation to determine the listwise ranking. TSPRank can be directly integrated as an additional component into embeddings generated by existing backbone models to enhance ranking performance. Our extensive experiments across three backbone models on diverse tasks, including stock ranking, information retrieval, and historical events ordering, demonstrate that TSPRank significantly outperforms both pure pairwise and listwise methods. Our qualitative analysis reveals that TSPRank's main advantage over existing methods is its ability to harness global information better while ranking. TSPRank's robustness and superior performance across different domains highlight its potential as a versatile and effective LETOR solution. Weixian Waylon Li, Yftah Ziser, Yifei Xie 0002, Shay B. Cohen, Tiejun Ma |
KDD (1) | 5 |
| 2025 | Pre-training Time Series Models with Stock Data CustomizationabstractStock selection, which aims to predict stock prices and identify the most profitable ones, is a crucial task in finance. While existing methods primarily focus on developing model structures and building graphs for improved selection, pre-training strategies remain underexplored in this domain. Current stock series pre-training follows methods from other areas without adapting to the unique characteristics of financial data, particularly overlooking stock-specific contextual information and the non-stationary nature of stock prices. Consequently, the latent statistical features inherent in stock data are underutilized. In this paper, we propose three novel pre-training tasks tailored to stock data characteristics: stock code classification, stock sector classification, and moving average prediction. We develop the Stock Specialized Pre-trained Transformer (SSPT) based on a two-layer transformer architecture. Extensive experimental results validate the effectiveness of our pre-training methods and provide detailed guidance on their application. Evaluations on five stock datasets, including four markets and two time periods, demonstrate that SSPT consistently outperforms the market and existing methods in terms of both cumulative investment return ratio and Sharpe ratio. Additionally, our experiments on simulated data investigate the underlying mechanisms of our methods, providing insights into understanding price series. Our code is publicly available at: https://github.com/astudentuser/Pre-training-Time-Series-Models-with-Stock-Data-Customization. Tiejun Ma, Shay B. Cohen |
KDD (2) | 2 |
| 2024 | MANA-Net: Mitigating Aggregated Sentiment Homogenization with News Weighting for Enhanced Market PredictionabstractIt is widely acknowledged that extracting market sentiments from news data benefits market predictions. However, existing methods of using financial sentiments remain simplistic, relying on equal-weight and static aggregation to manage sentiments from multiple news items. This leads to a critical issue termed "Aggregated Sentiment Homogenization'', which has been explored through our analysis of a large financial news dataset from industry practice. This phenomenon occurs when aggregating numerous sentiments, causing representations to converge towards the mean values of sentiment distributions and thereby smoothing out unique and important information. Consequently, the aggregated sentiment representations lose much predictive value of news data. To address this problem, we introduce the Market Attention-weighted News Aggregation Network (MANA-Net), a novel method that leverages a dynamic market-news attention mechanism to aggregate news sentiments for market prediction. MANA-Net learns the relevance of news sentiments to price changes and assigns varying weights to individual news items. By integrating the news aggregation step into the networks for market prediction, MANA-Net allows for trainable sentiment representations that are optimized directly for prediction. We evaluate MANA-Net using the S&P 500 and NASDAQ 100 indices, along with financial news spanning from 2003 to 2018. Experimental results demonstrate that MANA-Net outperforms various recent market prediction methods, enhancing Profit & Loss by 1.1% and the daily Sharpe ratio by 0.252. Tiejun Ma |
CIKM | 2 |
| 2024 | Parallel Byzantine Consensus Based on Hierarchical Architecture and Trusted HardwareabstractByzantine fault-tolerant (BFT) state machine replication (SMR) is adopted to support blockchain consensus by tolerating arbitrarily faulty behaviours. However, the inherent complexity of BFT protocols makes existing BFT protocols hard to adapt to large-scale applications that require high scalability and performance. In this paper, we propose a BFT parallelism protocol designed to enhance its scalability by using a hierarchical multi-committee architecture. It also encompasses a cross-layer consensus operation flow to improve safety and support trusted execution environments (TEEs). Our proposed approach allows the lower bound on the number of peers to be reduced to$2f+1$. We show the value of our proposed protocol in comparison to other state-of-the-art BFT protocols through experiments and performance evaluations on a testbed built on a cloud platform. The proposed protocol demonstrates a remarkable level of scalability, capable of accommodating a growing number of peers. Additionally, it exhibits improved performance when contrasted with HotStuff and FastBFT, with approximately 100% and 200% enhancements, respectively. Xiao Chen 0003, Tiejun Ma, Btissam Er-Rahmadi, Jane Hillston, Guanxu Yuan |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | ParBFT: An Optimized Byzantine Consensus Parallelism SchemeabstractByzantine fault-tolerance (BFT) consensus is a fundamental building block of distributed systems such as blockchains. However, implementations based on classic PBFT and most linear PBFT-variants still suffer from message communication complexity, restricting the scalability and performance of BFT algorithms when serving large-scale systems with growing numbers of peers. To tackle the scalability and performance challenges, we proposeParBFT, a new Byzantine consensus parallelism scheme combining classic BFT protocols and a novel Bilevel Mixed-Integer Linear Programming (BL-MILP)-based optimisation model. The core aim of ParBFT is to improve scalability via parallel consensus while providing enhanced safety (i.e. ensuring consistent total order across all correct replicas). Another core novelty is the integration of the BL-MILP model into ParBFT. The BL-MILP allows us to compute optimal numerical decisions for parallel committees (i.e. the optimal number of committees and peer allocation for each committee) and improve consensus performance while ensuring security. Finally, we test the performance of the proposed ParBFT on Microsoft Azure Cloud systems with 20 to 300 peers and find that ParBFT can achieve significant improvement compared to the state-of-the-art protocols. Xiao Chen 0003, Btissam Er-Rahmadi, Tiejun Ma, Jane Hillston |
IEEE Trans. Computers | 3 |
| 2023 | Data Augmentation to Address Various Rotation Errors of Wearable Sensors for Robust Pre-impact Fall DetectionabstractOBJECTIVE: This paper proposes a novel application of data augmentation to address various rotation errors of wearable sensors for robust pre-impact fall detection. In such systems, sensor rotation errors are inevitable because of loose attachment and body movement during long deployment. METHODS: Two augmented models with uniform and normal strategies were compared with a non-augmented model on the original dataset (no rotation error) and a validation dataset (with rotation error). The validation dataset was constructed with three types of rotation errors, namely, pitch, roll, and compound roll and pitch (CRP) at three levels of range (low: 15°, medium: 30°, and high: 45°). RESULTS: Five-fold cross validation showed the two augmented models maintained accuracy (>98.5%) as high as the non-augmented model on the original dataset but showed considerable improvements of 6.11% and 6.50% on the validation dataset, respectively. CRP error negatively affected the model accuracy the most, followed by pitch and then roll errors. In addition, the normal model had advantages over the uniform model in the low-to-medium range of error, which is expected to be the typical error range in practical applications. As for lead time, similarly, the augmented models achieved performance similar to the non-augmented model on the original dataset but showed significant improvements on the validation dataset. CONCLUSION AND SIGNIFICANCE: Data augmentation had notable capacities to address sensor rotation errors for practical applications and augmented models especially the normal model showed good potential to be embedded in a wearable system for robust pre-impact fall detection and injury prevention. Xiaoqun Yu, Tiejun Ma, Jaehyuk Jang, Shuping Xiong |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Experiments and Analyses of Anonymization Mechanisms for Trajectory Data Publishing
She Sun, Shuai Ma 0001, Jinghe Song, Wen-Hai Yue, Xuelian Lin, Tiejun Ma |
J. Comput. Sci. Technol. | 6 |
| 2020 | Comparing the effectiveness of deep feedforward neural networks and shallow architectures for predicting stock price indices
Larry Olanrewaju Orimoloye, Ming-Chien Sung, Tiejun Ma, Johnnie E. V. Johnson |
Expert Syst. Appl. | 3 |
| 2017 | An Ensemble Approach to Link PredictionabstractA network with n nodes contains O(n2) possible links. Even for networks of modest size, it is often difficult to evaluate all pairwise possibilities for links in a meaningful way. Further, even though link prediction is closely related to missing value estimation problems, it is often difficult to use sophisticated models such as latent factor methods because of their computational complexity on large networks. Hence, most known link prediction methods are designed for evaluating the link propensity on a specified subset of links, rather than on the entire networks. In practice, however, it is essential to perform an exhaustive search over the entire networks. In this article, we propose an ensemble enabled approach to scaling up link prediction, by decomposing traditional link prediction problems into subproblems of smaller size. These subproblems are each solved with latent factor models, which can be effectively implemented on networks of modest size. By incorporating with the characteristics of link prediction, the ensemble approach further reduces the sizes of subproblems without sacrificing its prediction accuracy. The ensemble enabled approach has several advantages in terms of performance, and our experimental results demonstrate the effectiveness and scalability of our approach. Liang Duan, Shuai Ma 0001, Charu C. Aggarwal, Tiejun Ma, Jinpeng Huai |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | Single-view determinacy and rewriting completeness for a fragment of XPath queries
Lixiao Zheng, Shuai Ma 0001, Tiejun Ma |
Sci. China Inf. Sci. | 4 |
| 2016 | Bridging the divide in financial market forecasting: machine learners vs. financial economists
Ming-Wei Hsu, Stefan Lessmann, Ming-Chien Sung, Tiejun Ma, Johnnie E. V. Johnson |
Expert Syst. Appl. | 4 |
| 2013 | Combining POS Tagging, Lucene Search and Similarity Metrics for Entity Linking
Shujuan Zhao, Chune Li, Shuai Ma 0001, Tiejun Ma, Dianfu Ma |
WISE (1) | 4 |
| 2011 | Rule-Based Verification of Network Protocol Implementations Using Symbolic ExecutionabstractThe secure and correct implementation of network protocols for resource discovery, device configuration and network management is complex and error-prone. Protocol specifications contain ambiguities, leading to implementation flaws and security vulnerabilities in network daemons. Such problems are hard to detect because they are often triggered by complex sequences of packets that occur only after prolonged operation. The goal of this work is to find semantic bugs in network daemons. Our approach is to replay a set of input packets that result in high source code coverage of the daemon and observe potential violations of rules derived from the protocol specification. We describe SYMNV, a practical verification tool that first symbolically executes a network daemon to generate high coverage input packets and then checks a set of rules constraining permitted input and output packets. We have applied SYMNV to three different implementations of the Zeroconf protocol and show that it is able to discover non-trivial bugs. Jaeseung Song, Tiejun Ma, Cristian Cadar, Peter R. Pietzuch |
ICCCN | 2 |
| 2011 | Femtocell Coverage Optimisation Using Statistical Verification
Tiejun Ma, Peter R. Pietzuch |
Networking (1) | 1 |
| 2010 | Characterising a grid site's trafficabstractGrid computing has been widely adopted for intensive high performance computing. Since grid resources are distributed over complex large-scale infrastructures, understanding grid site data traffic behaviour is important for efficient resource utilisation, performance optimisation, and the design of future grid sites as well as traffic-aware grid applications. In this paper, we study and analyse the traffic generated at a grid site in the Large Hadron Collider (LHC) Computing Grid (LCG). We find that most of the generated traffic is TCP-based and that a small set of grid applications generate significant amounts of the data. Upon analysing the different traffic metrics, we also find that the traffic exhibits long-range dependence and self-similarity. We also investigate packet-level metrics such as throughput, packet rate, round trip time (RTT) and packet loss. Our study establishes that these metrics can be well represented by Gaussian mixture models. The findings we present in this paper will enable accurate grid site traffic monitoring and potentially on-the-fly traffic modelling and prediction. It will also lead to a better understanding of grid site's traffic behaviour and contribute to more efficient grid site planning, traffic management, data transmission protocol optimisation, and data-aware grid application design. Tiejun Ma, Yehia El-khatib, Michael Mackay 0001, Christopher Edwards |
HPDC | 1 |
| 2010 | On the Quality of Service of Crash-Recovery Failure DetectorsabstractWe model the probabilistic behavior of a system comprising a failure detector and a monitored crash-recovery target. We extend failure detectors to take account of failure recovery in the target system. This involves extending QoS measures to include the recovery detection speed and proportion of failures detected. We also extend estimating the parameters of the failure detector to achieve a required QoS to configuring the crash-recovery failure detector. We investigate the impact of the dependability of the monitored process on the QoS of our failure detector. Our analysis indicates that variation in the MTTF and MTTR of the monitored process can have a significant impact on the QoS of our failure detector. Our analysis is supported by simulations that validate our theoretical results. Tiejun Ma, Jane Hillston, Stuart Anderson 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2009 | Shibboleth Access for Resources on the National Grid Service (SARoNGS)abstractThe national grid service (NGS) provides access to compute and data resources for UK academics. Currently users are required to have an X.509 certificate from the UK e-science certification authority (CA) or one of its international peers to access the NGS. The CA must satisfy the requirements for internationally agreed assurance levels and some users find the processes of obtaining and managing certificates difficult. Shibboleth, an implementation of federation identity based authentication, has been widely deployed in academic environments in the UK. The SARoNGS project, was proposed to integrate the Shibboleth and X.509 based infrastructures, to deliver a production level service for accessing the NGS in a user friendly way. This paper describes an architecture by which users are authenticated by the UK access management federation to acquire low assurance credentials to access Grid resources on the NGS. Users can login to NGS resources via NGS portal, using their local institution's authentication system. Mike Jones 0002, Jens Jensen, Andrew Richards, David Wallom, Tiejun Ma, Robert Frank 0003, David Spence, Claire Devereux, Neil Geddes |
IAS | 6 |
| 2009 | An Image Processing Portal and Web-Service for the Study of Ancient DocumentsabstractLinking up two projects that are dedicated to facilitate the work of documentary scholars, this paper presents image processing algorithms tailored to the study of ancient documents and how they have been made available to the users through a portal that calls upon a web-service exploiting grid computational power. To that end, image processing algorithms were wrapped to fit into the National Grid Service (NGS) Uniform Execution Environment; the data model of an existing Virtual Research Environment (VRE-SDM) was extended; JSR-168 compliant portlets were developed to facilitate secure and seamless distributed image analysis; and a GridSAM interface between the portal and the NGS-installed algorithms was developed. The outcomes of the project include: a web-based application, a proof of concept for the usability of the VRE-SDM platform, an opportunity for wider dissemination for the image processing algorithms, and a proof of feasibility for the use of the NGS for Humanities applications. Ségolène M. Tarte, David Wallom, Pin Hu, Kang Tang, Tiejun Ma |
eScience | 5 |
| 2008 | Adding Standards Based Job Submission to a Commodity Grid BrokerabstractThe Condor matchmaker provides a powerful mechanism for matching together both users job requirements and resource providers requirements in such a manner that not only is a paring selected which satisfies both requirements but is optimal for both. This has made the Condor system a good choice for use as a meta-scheduler within the Grid. Integrating Condor within a wider Grid context is possible but only through modification to the Condor source code. Here we describe how the standards for job submission and resource descriptions can be integrated into the Condor system to allow arbitrary Grid resources which support these standards to be brokered through Condor. Dave Colling, A. Stephen McGough, Jazz Mack Smith, Vesso Novov, Tiejun Ma, David Wallom |
eScience | 5 |
| 2007 | On the Quality of Service of Crash-Recovery Failure DetectorsabstractIn this paper, we study and model a crash-recovery target and its failure detector's probabilistic behavior. We extend quality of service (QoS) metrics to measure the recovery detection speed and the proportion of the detected failures of a crash-recovery failure detector. Then the impact of the dependability of the crash-recovery target on the QoS bounds for such a crash-recovery failure detector is analysed by adopting general dependability metrics such as MTTF and MTTR. In addition, we analyse how to estimate the failure detector's parameters to achieve the QoS from a requirement based on Chen's NFD-S algorithm. We also demonstrate how to execute the configuration procedure of this crash-recovery failure detector. The simulations are based on the revised NFD-S algorithm with various MTTF and MTTR. The simulation results show that the dependability of a recoverable monitored target could have significant impact on the QoS of such a failure detector and match our analysis results. Tiejun Ma, Jane Hillston, Stuart Anderson 0001 |
DSN | 1 |