Mark Junjie Li

dblp:36/996 · also Mark Jun Jie Li · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-7252-5346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dual-Space Contrastive Learning with Abnormal Edge Suppression for Graph Anomaly Detection
abstract
Graph Anomaly Detection identifies nodes in a graph that deviate from normal behavior and finds wide applications in finance, social networks, and cybersecurity. Recent studies focus on capturing rich contrastive information between positive and negative samples by constructing multi-view contrast patterns through data augmentation, achieving notable performance gains. Nevertheless, existing methods often suffer from abnormal information diffusion, where anomalies propagate along abnormal edges and contaminate neighboring nodes, ultimately compromising the semantic consistency between the target node and its positive subgraph. Furthermore, most existing approaches learn node representations solely in Euclidean space, limiting their ability to capture the hierarchical structure prevalent in real-world graphs. To address these challenges, we propose a novel Dual-space Contrastive Learning Framework with Abnormal Edge Suppression, named DC-AES. By incorporating hyperbolic space, our framework preserves the hierarchical structure of the graph, while the abnormal edge suppression module mitigates anomaly diffusion by filtering out anomalous edges. Extensive experiments on six real datasets demonstrate the effectiveness of our approach compared to existing SOTA methods, with a maximum improvement of 6.63% in AUC.
Mark Junjie Li, Shiyang He, Jinren Li, Sunjie Huang
ECAI1
2025 LaPNER: Label-Aware Prompt Learning with Non-Entity Clustering Regularization for Few-Shot NER
abstract
Despite the recent success of prompt-based methods in few-shot named entity recognition (FSNER), most approaches rely on manually created prompts (e.g., templates or label words) that fail to capture sufficient semantic information. This limits generalization, particularly in cross-domain settings. Additionally, conventional FSNER methods typically assign tokens that are not meant to be recognized to a single non-entity category, even though this category encompasses both irrelevant tokens and other entities that are not needed to be recognized. Treating such a semantically diverse group as a single category introduces noise and confusion. In this paper, we propose an approach called Label-Aware Prompt learning with Non-Entity clustering Regularization for Few-Shot NER (LaPNER). We introduce a learnable prompt pool to enrich the semantic representations of label words. Additionally, we employ embedding clustering regularization to more effectively distinguish the heterogeneous tokens within the non-entity category. Comprehensive experiments on multiple benchmarks demonstrate that LaPNER consistently outperforms prior methods in various settings, highlighting its effectiveness in improving generalization across tasks.
Mark Junjie Li, Sunjie Huang, Yigang Lin, Qilong Gong
ECAI1
2025 Anomaly detection in attributed networks via local multi-order contrastive learning and global topology awareness
Mark Junjie Li, Sunjie Huang, Qin Zhang 0011, Meiting Li
Neurocomputing1
2024 A Payment Transaction Pre-training Model for Fraud Transaction Detection
abstract
The surge in merchant fraud poses a significant threat to market order and consumer security. Effective security monitoring for merchants is crucial in safeguarding the digital life ecosystem and users' financial well-being. Detecting daily fraudulent payment transactions, a challenging task for current methods, requires efficient transformation of transactions into embeddings, especially in representing merchants based on their behavioral transactions. To address this, we propose the Grouping Sampling-based Sequence Generation (GSSG) method to generate meaningful sequences, enabling interactions among correlated transactions. We introduce Hierarchical Embedding Learning (HEL) and Hierarchical Masking pre-training (HMP) for the effective representation of hierarchical structures within flat transaction sequences. Pretrained on WeChat Pay data, our model, PTP, demonstrates superior performance in downstream fraud transaction detection, especially in few-shot learning scenarios, showcasing great potential in payment transaction scenarios.
Wenxi Huang, Zhangyi Zhao, Xiaojun Chen 0006, Qin Zhang 0011, Mark Junjie Li, Hanjing Su, Qingyao Wu
CIKM5
2024 Boosting Attributed Graph Anomaly Detection via Negative Sample Awareness
Meiting Li, Mark Junjie Li, Lingxuan Zhu
ICANN (5)3
2024 Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-Labeling
abstract
Learning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure.
Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang
IEEE Trans. Multim.5
2023 RSP-gcForest: A Distributed Deep Forest via Random Sample Partition
abstract
Deep Forest, a powerful alternative to deep neural networks, has gained much attention due to its advantages, such as low complexity, minimal hyperparameter requirements, and strong application performance. In the current big data environment, where data volumes and model complexities are growing rapidly, distributed computing is needed to increase computational efficiency. Recently, a distributed deep forest approach, called BLB-gcForest (Bag of Little Bootstraps-gcForest), has been successful in reducing training instances within cascade forests, combining BLB and granularity segmentation, thus improving the computational efficiency and scalability of distributed deep forests. However, it still transmits the entire data set between layers and requires double sampling with BLB, limiting the amount of data and the scalability of resource utilization. This paper introduces a novel algorithm, RSP-gcForest, based on Random Sample Partition (RSP) to improve distributed deep forests computational efficiency and scalability. RSP-gcForest uses block-level samples that replace the full dataset, significantly reducing interlayer instance transmission and prediction within cascade forests. Additionally, RSP blocks are integrated with the segmentation granularity of cascade forests for ensemble learning, effectively addressing computational efficiency and resource constraints. We conducted experiments on four extensive datasets using Spark and evaluated performance across five key metrics. The results clearly show that RSP-gcForest, while maintaining high classification quality, surpasses state-of-the-art methods in terms of computational efficiency and resource utilization. Furthermore, it achieves superior load balancing, demonstrating its potential as a powerful tool in big data and distributed computing.
Mark Junjie Li, Wenzhu Cai, Yigang Lin, Sunjie Huang, Joshua Zhexue Huang, Patrick Xiaogang Peng
IEEE Big Data1
2023 Anomaly Detection in Directed Dynamic Graphs via RDGCN and LSTAN
Mark Junjie Li, Zukang Gao, Xianyu Bao, Meiting Li
ICANN (3)1
2023 Topic Modeling for Short Texts via Adaptive P$\acute{o}$lya Urn Dirichlet Multinomial Mixture
Mark Junjie Li, Rui Wang 0136, Xianyu Bao, Jueying He, Lijuan He
ICONIP (14)1
2023 Encouraging Sparsity in Neural Topic Modeling with Non-Mean-Field Inference
Rui Wang 0136, Jueying He, Mark Junjie Li
ECML/PKDD (4)4
2022 Semi-automatic Data Annotation System for Multi-Target Multi-Camera Vehicle Tracking
abstract
Multi-target multi-camera tracking (MTMCT) plays an important role in intelligent video analysis, surveillance video retrieval, and other application scenarios. Nowadays, the deep-learning-based MTMCT has been the mainstream and has achieved fascinating improvements regarding tracking accuracy and efficiency. However, according to our investigation, the lacking of datasets focusing on real-world application scenarios limits the further improvements for current learning-based MTMCT models. Specifically, the learning-based MTMCT models training by common datasets usually cannot achieve satisfactory results in real-world application scenarios. Motivated by this, this paper presents a semi-automatic data annotation system to facilitate the real-world MTMCT dataset establishment. The proposed system first employs a deep-learning-based single-camera trajectory generation method to automatically extract trajectories from surveillance videos. Subsequently, the system provides a recommendation list in the following manual cross-camera trajectory matching process. The recommendation list is generated based on side information, including camera location, timestamp relation, and background scene. In the experimental stage, extensive results further demonstrate the efficiency of the proposed system.
Haohong Liao, Silin Zheng, Xuelin Shen, Mark Junjie Li, Xu Wang 0006
DSAA4
2022 DuSAG: An Anomaly Detection Method in Dynamic Graph Based on Dual Self-attention
Weiqin Lin, Xianyu Bao, Mark Junjie Li, Zukang Gao
ICANN (1)3
2022 Multi-knowledge Embeddings Enhanced Topic Modeling for Short Texts
Jueying He, Mark Junjie Li
ICONIP (3)3
2022 HSGAN: Reducing mode collapse in GANs by the latent code distance of homogeneous samples
Simin Yu, Kuntian Zhang, Chuan Xiao 0001, Joshua Zhexue Huang, Mark Junjie Li, Makoto Onizuka
Comput. Vis. Image Underst.5
2021 CmaGraph: A TriBlocks Anomaly Detection Method in Dynamic Graph Using Evolutionary Community Representation Learning
Weiqin Lin, Xianyu Bao, Mark Junjie Li
ICANN (1)3
2021 BTGAN: Training GAN with Balanced Triplet Loss and Two-Branch Architecture
abstract
TripletGAN is a variant of Generative Adversarial Network (GAN) by replacing the classification loss of discriminator with a triplet loss. Although TripletGAN delivers better mode coverage than vanilla GAN thanks to the characteristics of adversarial triplet loss that maximizes the embedding distance between generated samples, its adversarial training method suffers from the drawback that some generated images tend to deviate from the real sample distribution and noisy images are produced as we increase the number of iterations of training. In this paper, we propose an adversarially balanced triplet loss with four dynamic coefficients to achieve a trade-off between the quality and the diversity of generated samples. We also design a novel network architecture to provide GANs with an auto-encoding ability. Extensive experiments demonstrate the effectiveness of our proposed methods in terms of alleviating the problem in TripletGAN and the superiority in terms of reconstruction over some methods that directly train generator and encoder such as O-GAN.
Simin Yu, Kuntian Zhang, Chuan Xiao 0001, Xianyu Bao, Joshua Zhexue Huang, Mark Junjie Li
IJCNN6
2019 ASP-based Discovery of Semi-Markovian Causal Models under Weaker Assumptions
abstract
In recent years the possibility of relaxing the so-called Faithfulness assumption in automated causal discovery has been investigated. The investigation showed (1) that the Faithfulness assumption can be weakened in various ways that in an important sense preserve its power, and (2) that weakening of Faithfulness may help to speed up methods based on Answer Set Programming. However, this line of work has so far only considered the discovery of causal models without latent variables. In this paper, we study weakenings of Faithfulness for constraint-based discovery of semi-Markovian causal models, which accommodate the possibility of latent variables, and show that both (1) and (2) remain the case in this more realistic setting.
Zhalama, Jiji Zhang, Frederick Eberhardt, Wolfgang Mayer, Mark Junjie Li
IJCAI5
2015 A New Feature Sampling Method in Random Forests for Predicting High-Dimensional Data
Thanh-Tung Nguyen, He Zhao 0009, Joshua Zhexue Huang, Thi Thuy Nguyen, Mark Junjie Li
PAKDD (2)5
2014 Extensions to Quantile Regression Forests for Very High-Dimensional Data
Nguyen Thanh Tung, Joshua Zhexue Huang, Imran Khan 0009, Mark Junjie Li, Graham J. Williams
PAKDD (2)4
2012 Scalable Random Forests for Massive Data
Bingguo Li, Xiaojun Chen 0006, Mark Junjie Li, Joshua Zhexue Huang, Shengzhong Feng
PAKDD (1)3
2012 Hybrid Random Forests: Advantages of Mixed Trees in Classifying Text Data
Baoxun Xu, Joshua Zhexue Huang, Graham J. Williams, Mark Junjie Li, Yunming Ye
PAKDD (1)4
2010 CPLDP: An Efficient Large Dataset Processing System Built on Cloud Platform
Zhiyong Zhong, Mark Junjie Li, Jin Chang, Joshua Zhexue Huang, Shengzhong Feng
ADMA (2)2
2010 SKM-SNP: SNP markers detection method
Yang Liu 0100, Mark Junjie Li, Yiu-Ming Cheung, Pak Chung Sham, Michael Kwok-Po Ng
J. Biomed. Informatics2
2010 On cluster tree for nested and multi-density data clustering
Xutao Li 0003, Yunming Ye, Mark Junjie Li, Michael Kwok-Po Ng
Pattern Recognit.3
2008 Agglomerative Fuzzy K-Means Clustering Algorithm with Selection of Number of Clusters
abstract
In this paper, we present an agglomerative fuzzy $k$-means clustering algorithm for numerical data, an extension to the standard fuzzy $k$-means algorithm by introducing a penalty term to the objective function to make the clustering process not sensitive to the initial cluster centers. The new algorithm can produce more consistent clustering results from different sets of initial clusters centers. Combined with cluster validation techniques, the new algorithm can determine the number of clusters in a data set, which is a well known problem in $k$-means clustering. Experimental results on synthetic data sets (2 to 5 dimensions, 500 to 5000 objects and 3 to 7 clusters), the BIRCH two-dimensional data set of 20000 objects and 100 clusters, and the WINE data set of 178 objects, 17 dimensions and 3 clusters from UCI, have demonstrated the effectiveness of the new algorithm in producing consistent clustering results and determining the correct number of clusters in different data sets, some with overlapping inherent clusters.
Mark Junjie Li, Michael Kwok-Po Ng, Yiu-Ming Cheung, Joshua Zhexue Huang
IEEE Trans. Knowl. Data Eng.1
2007 On the Impact of Dissimilarity Measure in k-Modes Clustering Algorithm
abstract
This correspondence describes extensions to the k-modes algorithm for clustering categorical data. By modifying a simple matching dissimilarity measure for categorical objects, a heuristic approach was developed in [4], [12] which allows the use of the k-modes paradigm to obtain a cluster with strong intrasimilarity and to efficiently cluster large categorical data sets. The main aim of this paper is to rigorously derive the updating formula of the k-modes clustering algorithm with the new dissimilarity measure and the convergence of the algorithm under the optimization framework.
Michael Kwok-Po Ng, Mark Junjie Li, Joshua Zhexue Huang, Zengyou He
IEEE Trans. Pattern Anal. Mach. Intell.2