VLDB 2026 Research / reviewers in the wild / expert
Shijia Li
dblp:208/8194
· DBLP profile ↗
15ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | tsLLM-SG: Time Series Forecasting by Fine-Tuning LLMs with Dynamic Soft Prompts and Gated Frequency Transformation Adapters
Dongdong Mao, Shijia Li |
DASFAA (4) | 2 |
| 2025 | Galene: A Toolkit for Encapsulating Software Engineering Practices in Marine Science AI ApplicationsabstractMarine scientists increasingly build AI applications but often lack software engineering expertise, resulting in poor code quality and maintenance issues. Traditional platforms provide model repositories but do not address broader software engineering concerns. We present Galene, a Web-based toolkit that transparently integrates software engineering best practices into AI application development for marine scientists. Galene provides a reusable paradigm for encapsulating SE practices in domain-specific AI tools through integrated capabilities including seven key features. Scientists follow an intuitive 6step workflow to develop production-ready applications without explicit SE knowledge. Field deployment at a national marine science research institute suggests high ratings on perceived usefulness, ease of use, and future use intention. We demonstrate Galene through a sea surface temperature prediction scenario, showcasing the complete development process from data selection to deployment. The corresponding demo video can be viewed at https://www.youtube.com/watch?v=u-FVK-n5WvM. A fullyfunctional prototype is accessible at http://47.92.254.164:8419. Shijia Li |
APSEC | 1 |
| 2025 | Adversarially Robust Assembly Language Model for Packed Executables DetectionabstractDetecting packed executables is a critical component of large-scale malware analysis and antivirus engine workflows, as it identifies samples that warrant computationally intensive dynamic unpacking to reveal concealed malicious behavior. Traditionally, packer detection techniques have relied on empirical features, such as high entropy or specific binary patterns. However, these empirical, feature-based methods are increasingly vulnerable to evasion by adversarial samples or unknown packers (e.g., low-entropy packers). Furthermore, the dependence on expert-crafted features poses challenges in sustaining and evolving these methods over time. Shijia Li, Jiang Ming 0002, Lanqing Liu, Longwei Yang, Chunfu Jia |
CCS | 1 |
| 2025 | Aircraft EWIS safety risk level classification based on Multi-EDA and MHATT-BiLSTM
Yiqin Sang, Hongjuan Ge, Shijia Li |
Adv. Eng. Informatics | 4 |
| 2025 | An adaptive graph neural network-based intrusion detection system for airborne network
Shijia Li, Yiqin Sang, Hongjuan Ge |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A Transformer Model Incorporating Dynamic Chunking Strategy for Multivariate Time Series ClassificationabstractMultivariate time series classification tasks play a crucial role in the field of data mining and find wide applications in areas such as audio, healthcare, and transportation. The core challenge in multivariate time series classification tasks lies in capturing the relationships between variables and effectively learning the intricate details within the sequences. In recent years, multivariate time series classification methods based on deep learning have continued to emerge and have achieved certain results. However, existing methods commonly suffer from high computational complexity and an inability to capture the hidden relationships between variables fully. To address these issues, this paper proposes a method named DCSformer for multivariate time series classification, based on the Transformer framework. To reduce computational complexity, DCSformer utilizes a dynamic chunking strategy, processing sequences in patch form. It dynamically adjusts the chunk sizes and positions based on the internal feature values within the sequence. This strategy aims to not only decrease computational complexity but also better ensure the integrity of crucial sequence information during chunking. To comprehensively capture potential associations among variables, DCSformer utilizes a multivariate information fusion method, processing multivariate sequences dimension by dimension. It learns the positional relationships among various variables through the network and embeds these relationships into self-attention mechanisms to adjust weights, thereby fully considering and integrating interactions among variables. A series of experiments on electrocardiogram signals, electroencephalogram signals, human activity signals, and traffic signals in the UEA dataset were conducted in this study, comparing against baseline models. Experimental results demonstrate that DCSformer achieves superior classification performance across seven datasets in the evaluation. Shijia Li, Ziyuan Cheng, Zhongheng Sun, Cuiling Jiang, Yongjing Wan |
IJCNN | 2 |
| 2024 | Anomaly detection of aviation data bus based on SAE and IMD
Yiqin Sang, Hongjuan Ge, Shijia Li |
Comput. Secur. | 5 |
| 2023 | PackGenome: Automatically Generating Robust YARA Rules for Accurate Malware Packer DetectionabstractBinary packing, a widely-used program obfuscation style, compresses or encrypts the original program and then recovers it at runtime. Packed malware samples are pervasive---they conceal arresting code features as unintelligible data to evade detection. To rapidly respond to large-scale packed malware, security analysts search specific binary patterns to identify corresponding packers. The quality of such packer patterns or signatures is vital to malware dissection. However, existing packer signature rules severely rely on human analysts' experience. In addition to expensive manual efforts, these human-written rules (e.g., YARA) also suffer from high false positives: as they are designed to search the pattern of bytes rather than instructions, they are very likely to mismatch with unexpected instructions. Shijia Li, Jiang Ming 0002, Pengda Qiu, Qiyuan Chen 0006, Lanqing Liu, Huaifeng Bao, Qiang Wang 0059, Chunfu Jia |
CCS | 1 |
| 2022 | Chosen-Instruction Attack Against Commercial Code Virtualization Obfuscators
Shijia Li, Chunfu Jia, Pengda Qiu, Qiyuan Chen 0006, Jiang Ming 0002, Debin Gao |
NDSS | 1 |
| 2022 | Secure Repackage-Proofing Framework for Android Apps Using Collatz ConjectureabstractApp repackaging has been raising serious concerns about the health of the Android ecosystem, and repackage-proofing is an important mitigation against threat of such attacks. However, existing app repackage-proofing schemes were only evaluated against trivial adversaries simulated using analyzers for other purposes (e.g., disclosing privacy leakage vulnerabilities), hence were shown “effective” mainly because their key programming features were not even supported by those toolkits. Furthermore, existing works have also neglected dynamic adversaries capable of manipulating victim apps at runtime, making them vulnerable against such stronger opponents. In this article, we propose a novel repackage-proofing framework, which deploys distributed detection and response sites into the subject app's native partition to cross-verify all its code files. The detection sites transmit obtained integrity metrics to response sites via secure communication channels built on the subject app's own control flows using a specialized obfuscation technique based on Collatz conjecture, turning the repackage-proofing process into complicated implicit flows that are intrinsically difficult to be resolved due to the conjecture's nonlinear dynamical behaviors. We evaluated our framework using sophisticated Android data-flow analyzers. Results showed that our prototype effectively impeded analyses aiming to trace the information flows of its cross-verification. Shijia Li, Debin Gao, Chunfu Jia |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Active Warden Attack: On the (In)Effectiveness of Android App Repackage-ProofingabstractApp repackaging has raised serious concerns to the Android ecosystem with the repackage-proofing technology attracting attention in the Android research community. In this article, we first show that existing repackage-proofing schemes rely on a flawed security assumption, and then propose a new class ofactive warden attackthat intercepts and falsifies the metrics used by repackage-proofing for detecting the integrity violations during repackaging. We develop a proof-of-concept toolkit to demonstrate that all the existing repackage-proofing schemes can be bypassed by our attack toolkit. On the positive side, our analysis further identifies a new integrity metric in the Android ART runtime that can robustly and efficiently indicate bytecode tampering caused by either repackaging or active warden attacks. By associating this new metric with two supplemental verification mechanisms, we construct a multi-party verification framework that significantly raises the bar of repackage-proofing and identify conditions under which the proposed framework could detect app repackaging without getting compromised by active warden attacks. Shijia Li, Debin Gao, Daoyuan Wu, Qiaowen Jia, Chunfu Jia |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | Effective passenger flow forecasting using STL and ESN based on two improvement strategies
Lan Qin, Weide Li, Shijia Li |
Neurocomputing | 3 |
| 2019 | Xmark: Dynamic Software Watermarking Using Collatz ConjectureabstractDynamic software watermarking is one of the major countermeasures against software licensing violations. However, conventional dynamic watermarking approaches have exhibited a number of weaknesses including exploitable payload semantics, exploitable embedding/recognition procedures, and weak correlation between payload and subject software. This paper presents a novel dynamic watermarking method, Xmark, which leverages a well-known unsolved mathematical problem referred to as the Collatz conjecture. Our method works by transforming selected conditional constructs (which originally belonged to the software to be watermarked) with a control flow obfuscation technique based on Collatz conjecture. These obfuscation routines are built in a particular way such that they are able to express a watermark in the form of iteratively executed branching activities occurred during computing the aforementioned conjecture. Exploiting the one-to-one correspondence between natural numbers and their orbits computed by the conjecture (also known as the “Hailstone sequences”), Xmark's watermark-related activities are designed to be insignificant without the pre-defined secret input. Meanwhile, being integrated with obfuscation techniques, our method is able to resist attacks based on various reverse engineering techniques on both syntax and semantic levels. Analyses and simulations indicated that Xmark could evade detections via pattern matching and model checking, and meanwhile effectively prohibit dynamic symbolic execution. We have also shown that our method could remain robust even if a watermarked software is compromised via re-obfuscation using approaches like control flow flattening. Chunfu Jia, Shijia Li, Wantong Zheng, Dinghao Wu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | A Hierarchical Model with Pseudoinverse Learning Algorithm Optimazation for Pulsar Candidate SelectionabstractPulsars search has always been one of the most concerned problem in the field of astronomy. Nowadays, with the development of astronomical instruments and observation technology, the amount of data is getting bigger and bigger. Radio pulsar surveys have generated and will generate vast amounts of data. To handle big data, developing new technologies and frameworks to efficiently and accurately analyze these data become increasing urgent. The number of positive and negative samples in pulsar candidate data set is very unbalanced, if we only use these a few positive samples to train a deep neural network (DNN), the trained DNN is prone because of the problem of overfitting and will affect the generalization ability. Motivated by the mixtures of experts network architecture, we proposed a hierarchical model for pulsar candidate selection which assembles a set of trained base classifiers. Moreover, training a neural network always takes a lot of time because of using gradient descent (GD) based algorithm. In this work, we utilize the pseudoinverse learning algorithm instead of GD based algorithm to train proposed model. With the designed network architecture and adopted training algorithm, our model has the advantages not only with high steady-state precision but also good generalization performance. Shijia Li, Sibo Feng, Ping Guo 0002, Qian Yin 0001 |
CEC | 1 |
| 2017 | Image Recognition with Histogram of Oriented Gradient Feature and Pseudoinverse Learning AutoEncoders
Sibo Feng, Shijia Li, Ping Guo 0002, Qian Yin 0001 |
ICONIP (6) | 2 |