Yijun Bei

dblp:81/411 · DBLP profile ↗
← Back
34ranked-venue papers
15as first author
25since 2021 · last 2026
0000-0001-5720-6374ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 AWDTR: An industrial time series anomaly detection model based on adaptive wavelet multiscale decomposition and time-frequency domain reconstruction
Yijun Bei, Guoning Xu, Kewei Gao, Nianshu Wang, Qinqin Chen
Neurocomputing1
2026 Mhdaunet: enhancing semantic consistency in diffusion models via dual-path noise alignment
Yijun Bei, Sicheng Zuo, Nianshu Wang, Qinqin Chen
Mach. Vis. Appl.1
2026 Efficient discovery of co-movement patterns from video data
Yijun Bei, Dongxiang Zhang, Junnan Hu, Kian-Lee Tan, Gang Chen 0001
VLDB J.1
2025 CAMERA: Cross-Feature Aligned Model for Recognizing Top Antimicrobial Peptides
abstract
Antimicrobial peptides (AMPs) represent a mechanistically distinct therapeutic strategy to combat antimicrobial resistance, yet their discovery is hindered by sequence diversity, experimental noise in assays, and strain-specific activity. We present CAMERA, a cross-feature aligned ranking framework that fuses protein-language embeddings, predicted secondary structure and physicochemical descriptors via an information-maximising alignment loss. CAMERA employs a Mixture-of-Experts (MoE) router to dynamically activate taskadaptive subnetworks, enabling efficient multi-species generalization. To benchmark fine-grained antibacterial efficacy, we curate Bac11, an 11-species dataset with horizontally comparable minimum-inhibitory-concentration labels. On Bac11 and two public single-species datasets, CAMERA improves NDCG@5 by 36.66 percentage points and AUC by 3.57 percentage points over state-of-the-art baselines. Ablation studies demonstrate that cross-modal alignment and an auxiliary attribute regressor contribute substantial, complementary improvements. Our results indicate that multimodal learning-to-rank with conditional computation offers a scalable approach for strain-aware AMP discovery and can be extended to other peptide-driven therapeutic tasks. All code and public datasets used in this paper are available at https://github.com/hydrogen599/CAMERA.
Zilun Yan, Yanchao Xu, Kewei Gao, Yijun Bei
BIBM4
2025 Spatial-Temporal Forgery Trace Based Forgery Image Identification
Zunlei Feng, Jiachi Wang, Hengrui Lou, Binjia Zhou, Jie Lei 0002, Mingli Song, Yijun Bei
ICCV8
2025 STD-FD: Spatio-Temporal Distribution Fitting Deviation for AIGC Forgery Identification
abstract
With the rise of AIGC technologies, particularly diffusion models, highly realistic fake images that can deceive human visual perception has become feasible. Consequently, various forgery detection methods have emerged. However, existing methods treat the generation process of fake images as either a black-box or an auxiliary tool, offering limited insights into its underlying mechanisms. In this paper, we propose Spatio-Temporal Distribution Fitting Deviation (STD-FD) for AIGC forgery detection, which explores the generative process in detail. By decomposing and reconstructing data within generative diffusion models, initial experiments reveal temporal distribution fitting deviations during the image reconstruction process. These deviations are captured through reconstruction noise maps for each spatial semantic unit, derived via a super-resolution algorithm. Critical discriminative patterns, termed DFactors, are identified through statistical modeling of these deviations. Extensive experiments show that STD-FD effectively captures distribution patterns in AIGC-generated data, demonstrating strong robustness and generalizability while outperforming state-of-the-art (SOTA) methods on major datasets. The source code is available at [this link](https://github.com/HengruiLou/STDFD).
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Jie Lei 0002, Lechao Cheng, Jie Song 0011, Mingli Song, Yijun Bei
ICML9
2025 CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
abstract
With the swift progression of image generation technology, the widespread emergence of facial deepfakes poses significant challenges to the field of security, thus amplifying the urgent need for effective deepfake detection. Existing techniques for face forgery detection can broadly be categorized into two primary groups: visual-based methods and multimodal approaches. The former often lacks clear explanations for forgery details, while the latter, which merges visual and linguistic modalities, is more prone to the issue of hallucinations.To address these shortcomings, we introduce a visual detail enhanced self-correction framework, designated CorrDetail, for interpretable face forgery detection. CorrDetail is meticulously designed to rectify authentic forgery details when provided with error-guided questioning, with the aim of fostering the ability to uncover forgery details rather than yielding hallucinated responses. Additionally, to bolster the reliability of its findings, a visual fine-grained detail enhancement module is incorporated, supplying CorrDetail with more precise visual forgery details. Ultimately, a fusion decision strategy is devised to further augment the model's discriminative capacity in handling extreme samples, through the integration of visual information compensation and model bias reduction. Experimental results demonstrate that CorrDetail not only achieves state-of-the-art performance compared to the latest methodologies but also excels in accurately identifying forged details, all while exhibiting robust generalization capabilities.
Binjia Zhou, Hengrui Lou, Lizhe Chen, Dawei Luo, Jie Lei 0002, Zunlei Feng, Yijun Bei
IJCAI9
2025 A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Lechao Cheng, Jie Lei 0002, Jie Song 0011, Mingli Song, Yijun Bei
ACM Multimedia9
2025 STDC: Sparse Transformer Deep Collaboration Prompt Tuning for Industrial Multimodal Large Models
Yijun Bei
PRCV (5)1
2025 One-Step Multi-Frame Inpainting Framework for Real-Time Lip-Sync Digital Human Generation
abstract
In recent times, audio-driven lip-synching generation for digital humans has attracted considerable attention. However, the prevailing methodologies frequently encounter challenges pertaining to elevated computational complexity and deficient real-time performance. Although the MuseTalk framework has achieved notable progress in inference efficiency through its end-to-end, latent-space-based single-step generation algorithm, it still suffers from noticeable lip jitter and insufficient synchronization between audio and lip movements. To address these limitations, we propose an enhanced multi-frame inpainting framework that integrates Variational Autoencoders (VAE) and a multi-scale U-Net architecture. Specifically, our approach directly synthesizes the occluded lip region by leveraging multi-frame visual references combined with corresponding audio embeddings, thereby effectively improving lip synchronization and maintaining identity consistency. Furthermore, we introduce a landmark-guided multi-frame sampling strategy designed to enhance model attention towards lip dynamics. To facilitate deeper feature extraction and fusion, we propose a hierarchical latent-space feature fusion network (FusionNet), incorporating global and local residual connections and an enhanced Convolutional Block Attention Module. Additionally, a frame interpolation technique is employed during inference to further smooth lip movements and significantly mitigate lip jitter. The model has been trained on a large-scale Chinese dataset and comprehensively evaluated using both Chinese and English datasets. The experimental results demonstrate that the proposed framework achieves high visual accuracy, consistent lip synchronization, and efficient real-time inference, highlighting its strong cross-lingual generalization capability.
Yijun Bei, Yunze Qi, Hengrui Lou, Erteng Liu, Hongchang Zhang
Int. J. Pattern Recognit. Artif. Intell.1
2025 Behavior capture guided engagement recognition
Yijun Bei, Songyuan Guo, Kewei Gao, Zunlei Feng
Pattern Recognit.1
2025 Target-Directed Progressive Gradient Adjusting for transfer learning
Yijun Bei, Kewei Gao, Zhuoyang Zhao, Erteng Liu, Zunlei Feng
Pattern Recognit.1
2025 Mining Platoon Patterns from Traffic Videos
abstract
Discovering co-movement patterns from urban-scale video data sources has emerged as an attractive topic. This task aims to identify groups of objects that travel together along a common route, which offers effective support for government agencies in enhancing smart city management. However, the previous work has made a strong assumption on the accuracy of recovered trajectories from videos and their co-movement pattern definition requires the group of objects to appear across consecutive cameras along the common route. In practice, this often leads to missing patterns if a vehicle is not correctly identified from a certain camera due to object occlusion or vehicle mis-matching. To address this challenge, we propose a relaxed definition of co-movement patterns from video data, which removes the consecutiveness requirement in the common route and accommodates a certain number of missing captured cameras for objects within the group. Moreover, a novel enumeration framework called Max-Growth is developed to efficiently retrieve the relaxed patterns. Unlike previous filter-and-refine frameworks comprising both candidate enumeration and subsequent candidate verification procedures, MaxGrowth incurs no verification cost for the candidate patterns. It treats the co-movement pattern as an equivalent sequence of clusters, enumerating candidates with increasing sequence length while avoiding the generation of any false positives. Additionally, we also propose two effective pruning rules to efficiently filter the non-maximal patterns. Extensive experiments are conducted to validate the efficiency of MaxGrowth and the quality of its generated co-movement patterns. Our MaxGrowth runs up to two orders of magnitude faster than the baseline algorithm. It also demonstrates high accuracy in real video dataset when the trajectory recovery algorithm is not perfect.
Yijun Bei, Dongxiang Zhang, Sai Wu, Kian-Lee Tan, Gang Chen 0001
Proc. VLDB Endow.1
2024 ViT-Calibrator: Decision Stream Calibration for Vision Transformer
abstract
A surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing internal model architecture designs that often entail significant trial and error with high burdens. In this work, we propose a new paradigm dubbed Decision Stream Calibration that boosts the performance of general Vision Transformers. To achieve this, we shed light on the information propagation mechanism in the learning procedure by exploring the correlation between different tokens and the relevance coefficient of multiple dimensions. Upon further analysis, it was discovered that 1) the final decision is associated with tokens of foreground targets, while token features of foreground target will be transmitted into the next layer as much as possible, and the useless token features of background area will be eliminated gradually in the forward propagation. 2) Each category is solely associated with specific sparse dimensions in the tokens. Based on the discoveries mentioned above, we designed a two-stage calibration scheme, namely ViT-Calibrator, including token propagation calibration stage and dimension propagation calibration stage. Extensive experiments on commonly used datasets show that the proposed approach can achieve promising results.
Zhijie Jia, Lechao Cheng, Yang Gao 0001, Jie Lei 0002, Yijun Bei, Zunlei Feng
AAAI6
2024 Critical Feature Sifting and Dynamic Aggregation for Anomalous Audio Sequence Detection
Erteng Liu, Kewei Gao, Jianhai Chen, Yijun Bei, Zunlei Feng
ICONIP (1)6
2024 Discriminative Feature Decoupling Enhancement for Speech Forgery Detection
Yijun Bei, Erteng Liu, Yang Gao 0001, Kewei Gao, Zunlei Feng
IJCAI1
2024 Hundredfold Accelerating for Pathological Images Diagnosis and Prognosis through Self-reform Critical Region Focusing
Xiaotian Yu, Haoming Luo, Jiacong Hu, Xiuming Zhang, Yijun Bei, Mingli Song, Zunlei Feng
IJCAI7
2024 Manufacturing Domain QA with Integrated Term Enhanced RAG
abstract
Large Language Models (LLMs) have demonstrated powerful capabilities, yet LLMs face issues like hallucination in certain domain-specific areas. Consequently, an increasing number of domain-specific models are emerging. The current paradigm for domain-specific models involves training with domain data, followed by the employment of Retrieval-Augmented Generation (RAG) to mitigate hallucination issues. However, in precision-critical domains such as manufacturing, if the knowledge documents are of low quality or contain noise, the context retrieved through simple semantic matching by RAG may not necessarily benefit model output. Additionally, there can be issues like getting "lost in the middle" due to irrelevant or excessive context. To overcome this, we introduce the Integrated Term Enhancement Methodology (ITEM). Inspired by Chinese educational methods focused on key term elucidation, ITEM extracts and explains critical terms precisely from knowledge documents to form a comprehensive Term Dictionary for retrieving terms and explanations to enhance query capabilities. This methodology refines query responses by providing more accurate and contextually relevant information. To assess ITEM's effectiveness, we utilize the Chinese Mould Manufacturing Dataset (CMMD) and Contextualized Adaptive Response Assessment (CARA) metric method. Our experiment demonstrates that ITEM significantly outperforms existing retrieval enhancement Dense Retrievers by over 17.0% in accuracy while requiring only 80% of their token length. Moreover, the accuracy of our method exceeded that of GPT-4 by 5.0%. This advancement represents a significant leap in context-specific retrieval in LLMs, especially beneficial for specialized domains. The results underscore ITEM's potential as a transformative method in the field, offering new perspectives on integrating domain-specific knowledge into LLMs.
Yijun Bei, Zhibin Fang, Shenyu Mao, Shuyi Yu, Yining Tong, Weimin Cai
IJCNN1
2024 Fire and Smoke Detection with Burning Intensity Representation
Xiaoyi Han, Yanfei Wu, Nan Pu, Zunlei Feng, Qifei Zhang 0001, Yijun Bei, Lechao Cheng
MMAsia6
2024 Behavior Capture Based Explainable Engagement Recognition
Yijun Bei, Songyuan Guo, Kewei Gao, Zunlei Feng, Yining Tong, Weimin Cai, Lechao Cheng
PRCV (10)1
2024 Benchmarking Multi-Scene Fire and Smoke Detection
Xiaoyi Han, Nan Pu, Zunlei Feng, Yijun Bei, Qifei Zhang 0001, Lechao Cheng
PRCV (11)4
2024 Multi-layer Tuning CLIP for Few-Shot Image Classification
Jinsong Geng, Cenyu Liu, Zunlei Feng, Yijun Bei
PRCV (5)7
2024 EC-PEFT: An Expertise-Centric Parameter-Efficient Fine-Tuning Framework for Large Language Models
Xuelin Cheng, Yijun Bei
PRICAI (2)4
2023 CKR-Calibrator: Convolution Kernel Robustness Evaluation and Calibration
Yijun Bei, Jinsong Geng, Erteng Liu, Kewei Gao, Wenqi Huang 0002, Zunlei Feng
ICONIP (5)1
2023 Co-movement Pattern Mining from Videos
abstract
Co-movement pattern mining from GPS trajectories has been an intriguing subject in spatial-temporal data mining. In this paper, we extend this research line by migrating the data source from GPS sensors to surveillance cameras, and presenting the first investigation into co-movement pattern mining from videos. We formulate the new problem, re-define the spatial-temporal proximity constraints from cameras deployed in a road network, and theoretically prove its hardness. Due to the lack of readily applicable solutions, we adapt existing techniques and propose two competitive baselines using Apriori-based enumerator and CMC algorithm, respectively. As the principal technical contributions, we introduce a novel index called temporal-cluster suffix tree (TCS-tree), which performs two-level temporal clustering within each camera and constructs a suffix tree from the resulting clusters. Moreover, we present a sequence-ahead pruning framework based on TCS-tree, which enables the concurrent utilization of all pattern constraints to filter candidate paths. Finally, to reduce verification cost on the candidate paths, we propose a sliding-window based co-movement pattern enumeration strategy and a hashing-based dominance eliminator, both of which are effective in avoiding redundant operations. We conduct extensive experiments for scalability and effectiveness analysis. Our results validate the efficiency of the proposed index and mining algorithm, which runs remarkably faster than the two baseline methods. Additionally, we construct a video database with 1169 cameras and perform an end-to-end pipeline analysis to study the performance gap between GPS-driven and video-driven methods. Our results demonstrate that the derived patterns from the video-driven approach are similar to those derived from groundtruth trajectories, providing evidence of its effectiveness.
Dongxiang Zhang, Junnan Hu, Yijun Bei, Kian-Lee Tan, Gang Chen 0001
Proc. VLDB Endow.4
2014 Graph indexing for large networks: A neighborhood tree-based approach
Yijun Bei
Knowl. Based Syst.2
2013 HBase System-Based Distributed Framework for Searching Large Graph Databases
abstract
In recent years, graphs have become popular in a variety of domains, with the volume of graph databases increasing rapidly. Therefore, large amounts of graph data need to be indexed, queried, and stored. To handle problems on graph querying and storing for large graph databases, we present a distributed graph searching framework based on the well-known storage system HBase. We improve the querying and indexing methods in GiS [4] to process the querying problem in a distributed environment. To ensure that graph data are well distributed, we design a distributed graph indexing technique that uses a line graph signature to index graphs in HBase. Experiments in both real and synthetic databases demonstrate that the proposed framework is an efficient distributed solution for querying sub graphs in a large volume of graph data.
Yijun Bei
SNPD1
2012 Evaluating geo-social influence in location-based social networks
abstract
The emerging location-based social network (LBSN) services not only allow people to maintain cyber links with their friends, but also enable them to share the events happening on them at different locations. The geo-social correlations among event participants make it possible to quantify mutual user influence for various events. Such a quantification of influence could benefit a wide spectrum of real-life applications such as targeted advertising and viral marketing.
Chao Zhang 0014, Lidan Shou, Ke Chen 0005, Gang Chen 0001, Yijun Bei
CIKM5
2011 PISA: A framework for integrating uncooperative peers into P2P-based federated search
Gang Chen 0001, Zujie Ren, Lidan Shou, Ke Chen 0005, Yijun Bei
Comput. Commun.5
2010 HAPS: Supporting Effective and Efficient Full-Text P2P Search with Peer Dynamics
Zujie Ren, Ke Chen 0005, Lidan Shou, Gang Chen 0001, Yijun Bei
J. Comput. Sci. Technol.5
2009 PISA: Federated Search in P2P Networks with Uncooperative Peers
Zujie Ren, Lidan Shou, Gang Chen 0001, Chun Chen 0001, Yijun Bei
DEXA5
2009 Bottom-up discovery of frequent rooted unordered subtrees
Yijun Bei, Gang Chen 0001, Lidan Shou, Jinxiang Dong
Inf. Sci.1
2007 A Caching System for XML Queries Using Frequent Query Patterns
abstract
In this paper, we present an optimization framework for querying XML data by caching frequent query patterns. In this framework, frequent query patterns are mined online from user queries, and these query results are cached for future use. The mining process of frequent query patterns is launched automatically when user queries meet predefined requirements. To deal with queries that are similar to but not exactly same as the cached ones, a novel technique named query rewriting is adopted. This technique is able to handle four kinds of similar queries namely exact matching, exact containment, semantic matching and semantic containment. A cache replacement scheme that utilizes both the query pattern support and query pattern accessing time is employed to perform a fine-grained region purging. Experiments are carried out on the XMARK datasets. The results show that the proposed methods are both effective and efficient in improving the performance of XML queries.
Yijun Bei, Gang Chen 0001, Tianlei Hu, Jinxiang Dong
CSCWD1
2007 Accelerating Parent-Child Path Matching in XML
abstract
With the rapidly increasing popularity of XML as a data format, there is a large demand for efficient XML structural matching techniques. Normally, data in XML are stored in a tree-like structure where nodes (with data) are located using the path relations in the tree. This paper proposes a coding policy named path code to accelerate an important category of structure matching in XML, namely the Parent-Child paths. A Parent-Child path is the path expression that contains only parent-child relationships. The proposed path code employs a partial prefix Path of XML elements using a special compression technique. Based on the path code, we present a Boosting algorithm, which has low-linear time complexity and little I/O cost, to match Parent-Child paths. In addition, we propose two heuristic policies to optimize the Boosting algorithm further by reducing the memory consumption. Our experiments show that the Boosting algorithm considerably outperforms previous algorithms in run-time Parent-Child path matching.
Gang Chen 0001, Lihua Yu, Yijun Bei, Jinxiang Dong
CSCWD4