EDBT 2026 Demo / reviewers in the wild / expert
Liqun Yang
dblp:72/9582
· DBLP profile ↗
36ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Computer networks · 6 · 3 first-author · 6 since 2021Security and privacy · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task Offloading and Resource Scheduling in a Vehicle-RSU-Cloud Resource Environment
Liqun Yang |
J. Grid Comput. | 1 |
| 2026 | Dynamic Searchable Symmetric Encryption With Efficient and Complete Access Control for Multi-User Cloud ComputingabstractSearchable symmetric encryption (SSE) enables the storage and retrieval of encrypted data on untrusted cloud servers, while dynamic searchable symmetric encryption (DSSE) further supports updating encrypted data. To date, in multi-user environments, most DSSE schemes cannot achieve simultaneous access control for both keyword retrieval and data updates. To address this issue, we propose a new DSSE scheme with efficient and complete(keyword retrieval and update)access control for multi-user environments, named EFCAM. Our work has simultaneously achieved efficient, flexible, and fine-grained access control for keyword retrieval and updating, this is extremely rare in existing research. For update operations, we combine file index encoding and homomorphic encryption (HE) technology, so that EFCAM optimizes the calculation; to achieve flexible access control, we adopt an equality test scheme that can supports three types of update authorization. For retrieval operations, users do not need to share keys. By executing a single query, the users can effectively retrieve all the data that they have permission to access. To enhance system security and operational efficiency, we have extended EFCAM with a dynamic policy update mechanism for flexible and real-time adjustment of access control policies. We formally analyze the security of EFCAM to prove that our scheme has forward security (FS) and backward security (BS). Experimental results show that, EFCAM maintains outstanding efficiency in encrypted data retrieval and update operations within multi-user environments, while also exhibiting strong scalability. Liqun Yang, Yuze Yang, Dusit Niyato, Zhoujun Li 0001, Wanxu Xia, Liang Sun 0007 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction TuningabstractRecent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-ranging code-related tasks. However, most previous existing methods mainly view each programming language in isolation and ignore the knowledge transfer among different programming languages. To bridge the gap among different programming languages, we introduce a novel multi-agent collaboration framework to enhance multilingual instruction tuning for code LLMs, where multiple language-specific intelligent agent components with generation memory work together to transfer knowledge from one language to another efficiently and effectively. Specifically, we first generate the language-specific instruction data from the code snippets and then provide the generated data as the seed data for language-specific agents. Multiple language-specific agents discuss and collaborate to formulate a new instruction and its corresponding solution (A new programming language or existing programming language), To further encourage the cross-lingual transfer, each agent stores its generation history as memory and then summarizes its merits and faults. Finally, the high-quality multilingual instruction data is used to encourage knowledge transfer among different programming languages to train Qwen2.5-xCoder. Experimental results on multilingual programming benchmarks demonstrate the superior performance of Qwen2.5-xCoder in sharing common knowledge, highlighting its potential to reduce the cross-lingual gap. Jian Yang 0003, Wei Zhang 0021, Yibo Miao, Shanghaoran Quan, Zhenhe Wu, Qiyao Peng 0006, Liqun Yang, Tianyu Liu 0001, Zeyu Cui, Binyuan Hui, Junyang Lin |
ACL (1) | 7 |
| 2025 | Breaking Size Barrier: Enhancing Reasoning for Large-Size Table Question Answering
Xianjie Wu, Di Liang, Jian Yang 0037, Xianfu Cheng, Linzheng Chai, Tongliang Li, Liqun Yang, Zhoujun Li 0001 |
DASFAA (2) | 7 |
| 2025 | CodeArena: Evaluating and Aligning CodeLLMs on Human PreferenceabstractJian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin |
EMNLP | 7 |
| 2025 | McEval: Massively Multilingual Code EvaluationabstractCode large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval. Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001 |
ICLR | 16 |
| 2025 | A New Perspective on CNN-Based Encrypted Traffic Classification: Data Preprocessing and Generalization Performance AnalysisabstractGiven the increasing prevalence of HTTPS communication on the Internet, the identification and classification of encrypted network traffic has become critical challenges for information security. Although deep learning models have demonstrated substantial performance in this area, existing research often relies on public datasets and prioritizes model architecture optimization over robust data preprocessing, limiting generalization performance when evaluated on external independent test datasets. To address this issue, this paper proposes a novel session-based approach for traffic dataset partitioning and preprocessing, utilizing the initial packets of each session as an input sample to enhance the model generalization performance. Furthermore, a Multi-Scale Convolutional Residual Network (MSCRNet) is designed and rigorously evaluated on both ISCXVPN2016 and the newly constructed LENS dataset. The experimental results show that the proposed initial-packets based preprocessing method combined with MSCRNet achieves an optimal generalization F1-score of 87.66% on the independent test dataset, representing a significant average improvement of 41.86% in generalization F1-score compared to conventional random packet partitioning schemes. Finally, we present a detailed empirical analysis of the impact of key preprocessing parameters—initial packet count, truncation length, and session quantity—on model performance, offering valuable insights for improving the robustness and efficacy of encrypted traffic classification systems. Xianfeng Ye, Tongge Xu, Liqun Yang |
LCN | 3 |
| 2025 | Efficient and secure multi-party computation protocol supporting deep learningabstractAbstract Privacy-preserving deep learning based on secure multi-party computation (MPC) has emerged as a critical research focus in recent years. While existing approaches predominantly employ additive secret sharing with a fixed number of parties, they have yet to fully leverage the more efficient Shamir-based schemes. However, the adoption of Shamir secret sharing faces two key challenges: limitations of decimal computation and signed number representation. Furthermore, current solutions often lack optimization for specific computational modules and rely on conventional methods ill-suited for MPC environments. To address these issues, this paper proposes a fixed-point decimal-supported Shamir secret sharing scheme. A key innovation is our truncation algorithm, which effectively manages the expanded decimal digits resulting from multiplication operations, enabling comprehensive fixed-point arithmetic within the Shamir-based MPC framework. Extensive large-scale simulations validate the accuracy of our truncation method. Moreover, we introduce optimized protocols for two crucial deep learning operations: convolution and Softmax function computation. Our convolution protocol leverages the Winograd algorithm to significantly reduce multiplication gate count, yielding over 50% performance improvement. For Softmax computation, we extend existing two-party protocols to a multi-party Shamir setting, developing the nQSMax algorithm. This algorithm achieves exceptional accuracy exceeding 99% within seconds, requiring only a few iterations. Shancheng Zhang, Zongyang Zhang, Minzhe Huang, Haochun Jin, Liqun Yang |
Cybersecur. | 6 |
| 2025 | FuzzCoder: Code Large Language Model-Based Fuzz Testing for Industrial IoT ProgramsabstractFuzz testing is an dynamic program analysis technique designed for discovering vulnerabilities in IoT systems. The core goal is to deliberately feed maliciously crafted inputs into an IoT device or service, triggering vulnerabilities such as system crashes, buffer overflow exploits, and memory corruption, etc. Efficiently generating malicious inputs remains challenging, with leading methods often relying on randomly mutating existing valid inputs. In this work, we propose to adopt fine-tuned large language models (FuzzCoder) to learn patterns in the input files from successful attacks to guide future fuzzing explorations. Specifically, we develop a framework that leverages code LLMs to guide the mutation process to perform meaningful input mutations. We formulate the mutation process as the sequenceto-sequence modeling, where LLM receives a sequence of bytes and outputs the mutated byte sequence. FuzzCoder is fine-tuned on our created instruction dataset (FuzzInstruct), where the successful fuzzing history is collected from the heuristic fuzzing tool. FuzzCoder can predict mutation positions and strategies for input files to trigger abnormal behaviors of the program. Most importantly, the experiment reveals results that FuzzCoder achieves better fuzzing performance compared to traditional and other AFL-based fuzzers, such as AFL, AFL++, AFLSmart, etc. On average, FuzzCoder achieves an improvement in code coverage of more than 20%, along with a significant increase in the number of crashes. 1 Liqun Yang, Chaoren Wei, Jian Yang 0030, Wanxu Xia, Yuze Yang, Dusit Niyato, Liang Sun 0007, Zhiquan Liu 0001 |
IEEE Internet Things J. | 1 |
| 2025 | ANT-ET: An end-to-end multimodal framework for fine-grained encrypted traffic fingerprintingabstractThe widespread use of encryption protocols and increasing privacy demands have significantly increased encrypted traffic, creating new challenges for network monitoring and threat detection. Current methods struggle with diverse scenarios and distinguish between subtle traffic patterns within webpages of the same application. To address these challenges, we introduce ANT-ET, an end-to-end multimodal framework designed for fine-grained encrypted webpage traffic fingerprinting. ANT-ET leverages a transformer to model payload semantics and constructs a traffic interaction graph to capture both temporal and spatial characteristics of packet interactions. Additionally, ANT-ET incorporates a gradient reversal layer to improve generalization by facilitating domain-invariant feature learning across related webpages. Experimental results demonstrate ANT-ET’s superior performance compared to various baseline models, which were evaluated using a proprietary encrypted webpage traffic dataset and three public datasets. Ablation studies confirm the effectiveness of different framework components, while sensitivity and complexity analyses further validate ANT-ET’s robustness and flexibility. He Kong 0003, Liqun Yang, Jingguo Ge, Tong Li 0012, Hui Li 0098 |
J. Comput. Secur. | 2 |
| 2025 | DPRFuzz: Enhancing Vulnerability Mining With Two-Stage Reinforcement LearningabstractAmerican fuzzy lop (AFL), as a representative tool for fuzzing, is capable of uncovering security vulnerabilities in industrial systems. It suffers from consuming a large amount of computational resources during the mutation. To improve the performance of AFL, researchers adopt algorithms, such as particle swarm optimization and long short-term memory, to optimize mutation operator selection. However, challenges persist in these approaches integrated with AFL, including optimization model complexity, insufficient accuracy, and poor generalization scalability. To address these issues, the article proposes a new fuzzer calledDPRFuzzto optimize AFL’s mutation phases. First, in the deterministic mutation strategy mutation phase, deep Q network and trust region policy optimization are leveraged to precisely generate effective mutated samples through perceiving mutation process in a relatively short time. Then, to boost the efficiency of the Havoc random mutation phase, we improve the Thompson sampling algorithm based on a multiagent strategy to generate an overall optimal mutation strategy chain. Finally, the approach is tested on eight programs, such asreadelf,tcpdump,andnm, and the advantages ofDPRFuzzare analyzed. Most importantly, the experiment reveals results thatDPRFuzzachieves better fuzzing performance compared to the traditional and other AFL-based fuzzers, such as AFL, AFL++, AFLSmart, etc. On average,DPRFuzzachieves an improvement in code coverage of over 10%, along with a significant increase in the number of crashes. Liqun Yang, Ruihao Li 0010, Chaoren Wei, Jian Yang 0030, Yuze Yang, Liang Sun 0007, Dong Zhao 0004, Zhoujun Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Joint UL-DL Power Allocation for Massive MIMO URLLC IoT Networks: A Comparative Study of Different Pilot PatternsabstractIn this paper, we employ massive multiple-input and multiple-output (MIMO) technology to support multiple Internet-of-Things devices with ultra-reliability and low-latency communication (URLLC) industrial applications. Specifically, we first derive lower bounds (LBs) on the achievable uplink (UL) and downlink (DL) data rates under the finite blocklength (FBL) and pilot contamination, where each base station (BS) employs maximum-ratio transmission (MRT) in the DL and maximum-ratio combining (MRC) in the UL detection. In addition, the LB rates are derived for two types of pilot of the regular pilot (RP) and superimposed pilot (SP). We study joint UL-DL power allocation optimization where the objective is to maximize the UL-DL overall average weighted sum rate (WSR) for the systems individually with RP and SP schemes. We propose to employ successive convex approximation to transform the original problems into a series of geometric program problems. Then, an iterative algorithm is proposed to jointly optimize the UL and DL pilot and data payload power allocation. Simulation results are shown to compare the performances of the systems with RP and SP schemes for different settings. Simulation results also verify that the derived LB rates tightly match the corresponding ergodic rates and confirm the rapid convergence speed of the proposed iterative algorithms. Liang Sun 0007, Yuanwei Liu, Liqun Yang |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | UniCoder: Scaling Code Large Language Model via Universal CodeabstractTao Sun, Linzheng Chai, Jian Yang, Yuwei Yin, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, Zhoujun Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tao Sun 0016, Linzheng Chai, Jian Yang 0030, Yuwei Yin, Hongcheng Guo, Liqun Yang, Zhoujun Li 0001 |
ACL (1) | 8 |
| 2024 | m3P: Towards Multimodal Multilingual Translation with Multimodal PromptabstractMultilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario. Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001 |
LREC/COLING | 9 |
| 2024 | Incorporating Dynamic Temperature Estimation into Contrastive Learning on GraphsabstractContrastive learning, a powerful self-supervised learning paradigm, has shown its efficacy in learning embed dings from independent and identically distributed (IID) as well as non-IID data without relying on label information. Since high-quality discriminative embeddings form a rich embedding space, which benefits model performance on downstream tasks, it is necessary to study how to improve the quality of contrastive node embeddings in graph contrastive learning. However, there has been limited research on this area. In this paper, we investigate how to generate high-quality contrastive node embeddings based on an in-depth analysis of graph contrastive losses. Firstly, we propose a novel and effective method, GLATE, for estimating the temperatures in three mainstream graph contrastive losses during the training phase. Secondly, we conduct the derivation of GLATE, and the derivation results reveal the specific relationship between the quality of contrastive node embeddings and tem-peratures. Finally, the extensive experiments on 16 benchmark datasets demonstrate that GLATE consistently outperforms the state-of-the-art graph contrastive learning models in terms of both model performance and training efficiency. Ziyang Liu 0004, Chaokun Wang, Liqun Yang, Yunkai Lou, Hao Feng 0007, Cheng Wu 0004, Kai Zheng 0001, Yang Song 0008 |
ICDE | 3 |
| 2024 | OWL: A Large Language Model for IT OperationsabstractWith the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition, machine translation, and dialogue systems. Recently, Large Language Models (LLMs) have achieved significant improvements across various domain-specific areas. However, there is a noticeable gap in the development of specialized Large Language Models (LLMs) tailored for IT operations. In this paper, we introduce the OWL, a large language model trained on our constructed Owl-Instruct with a wide range of IT-related information. Specifically, limited by the maximum input length, we propose the \textbf{H}omogeneous \textbf{M}arkov \textbf{C}ontext \textbf{E}xtension method (HMCE). The mixture-of-adapter strategy is leveraged to improve the parameter-efficient tuning across different domains or tasks.
Further, we evaluate the performance of OWL on the Owl-Bench established by us and open IT-related benchmarks. OWL demonstrates superior performance results on IT tasks, which outperforms existing models by significant margins. Moreover, we hope that the findings of our work will provide more insights to revolutionize the techniques of IT operations with specialized LLMs. Hongcheng Guo, Jian Yang 0030, Liqun Yang, Linzheng Chai, Jiaqi Bai 0001, Junran Peng, Xiaorong Hu, Dongfeng Zhang, Xu Shi 0005, Tieqiao Zheng, Liangfan Zheng, Bo Zhang 0096, Ke Xu 0001, Zhoujun Li 0001 |
ICLR | 4 |
| 2024 | mt4CrossOIE: Multi-stage tuning for cross-lingual open information extraction
Tongliang Li, Linzheng Chai, Jian Yang 0030, Jiaqi Bai 0001, Yuwei Yin, Hongcheng Guo, Liqun Yang, Hebboul Zine El Abidine, Zhoujun Li 0001 |
Expert Syst. Appl. | 9 |
| 2023 | GanLM: Encoder-Decoder Pre-training with an Auxiliary DiscriminatorabstractJian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001 |
ACL (1) | 8 |
| 2023 | HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei |
DASFAA (3) | 4 |
| 2023 | QURG: Question Rewriting Guided Context-Dependent Text-to-SQL Semantic Parsing
Linzheng Chai, Dongling Xiao, Jian Yang 0030, Liqun Yang, Qian-Wen Zhang, Yunbo Cao, Zhoujun Li 0001 |
PRICAI (2) | 5 |
| 2023 | GTrans: Grouping and Fusing Transformer Layers for Neural Machine TranslationabstractTransformer structure, stacked by a sequence of encoder and decoder network layers, achieves significant development in neural machine translation. However, vanilla Transformer mainly exploits the top-layer representation, assuming the lower layers provide trivial or redundant information and thus ignoring the bottom-layer feature that is potentially valuable. In this work, we propose theGroup-Transformer model (GTrans) that flexibly divides multi-layer representations of both encoder and decoder into different groups and then fuses these group features to generate target words. To corroborate the effectiveness of the proposed method, extensive experiments and analytic experiments are conducted on three bilingual translation benchmarks and three multilingual translation tasks, including the IWLST-14, IWLST-17, LDC, WMT-14, WMT-21 and OPUS-100 benchmark. Experimental and analytical results demonstrate that our model outperforms its Transformer counterparts by a consistent gain. Furthermore, it can be successfully scaled up to 60 encoder layers and 36 decoder layers. Jian Yang 0030, Yuwei Yin, Liqun Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Furu Wei, Zhoujun Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Sessionvideo: A Novel Approach for Encrypted Traffic Classification via 3D-CNN ModelabstractToday encrypted traffic has been used widely on the internet, such as HTTPS and SSH. Network traffic classification plays an important role in network resource management and cyberspace security, and previous machine learning methods face the challenge of identifying and classifying the encrypted traffic. In this paper, we propose a deep learning method based on a 3D convolutional neural network (3D-CNN), called SessionVideo. The new approach integrates traffic payload and time feature extraction for application classification. Specifically, our proposed scheme converts the raw traffic into a simple grayscale video, which enables the model to capture the temporal characteristics of the traffic. To verify the effectiveness of this scheme, we build a traffic dataset of 20 applications. The experimental results demonstrate that our 3D-CNN model achieves an accuracy of 97.89% and a weighted average precision of 97.96%. Further, the evaluation metrics of the 3D-CNN model significantly outperform 1D-CNN and 2D-CNN models in the comparison experiments. Tongge Xu, Jian Yang 0030, Lijin Wu, Liqun Yang |
APNOMS | 5 |
| 2022 | Cdga: A GAN-based Controllable Domain Generation AlgorithmabstractRecently Command and Control (C&C) servers have attracted considerable attention in botnets and domain generation algorithms (DGAs) further enhance the stealth of C&C servers. However, Algorithmically Generated Domains (AGDs) generated by DGAs can be easily detected by previous DGA detection approaches. More specifically, the previous DGAs are hard to satisfy domain name rules, low repetition rate, and anti-detection in practical scenarios simultaneously. Designing an outstanding DGA has become a crucial issue from the botnet owner’s perspective. To mitigate these problems, we propose Cdga, a Controllable DGA via Generative Adversarial Networks (GAN), which is a popular backbone model for text generation in the natural language processing (NLP) community.Controllable text generation approaches are adopted by Cdga to ensure no repetition in the generated domain names and compliance with the domain rules. In addition to cheating DGA detectors, GANs are exploited to equip Cdga with a powerful anti-detection ability. Furthermore, our proposed method uses the technique of NLP to force the AGDs to meet language rules, where the generated domain names are difficult for recognition by human. By utilizing the time-dependent seed, Cdga can dynamically generate domain names, ensuring that the malware can connect to the C&C server conditioned on a specific time stamp. Experimental results demonstrate that the domain names generated by our method are realistic enough to be resistant to the state-of-the-art DGA detectors. You Zhai, Jian Yang 0030, Longtao He, Liqun Yang, Zhoujun Li 0001 |
TrustCom | 5 |
| 2022 | A new methodology for anomaly detection of attacks in IEC 61850-based substation system
Liqun Yang, You Zhai, Zhoujun Li 0001, Tongge Xu |
J. Inf. Secur. Appl. | 1 |
| 2021 | Slice-sampling based 3D Object ClassificationabstractMultiview-based 3D object detection achieved great success in the past years. However, for some complex models with complex inner structures, the performances of these methods are not satisfactory. This paper provides a method based on slide sampling for 3D object classification. First, we slice and sample the model from the different depths and directions to get the model’s features. Then, a deep neural network designed based on the attention mechanism is used to classify the input data. The experiments show that the performance of our method is competitive on ModelNet. Moreover, for some special models with simple surfaces and complex inner structures, the performance of our method is outstanding and stable. Xiangwen Zhao, Yi-Jun Yang, Wei Zeng 0019, Liqun Yang |
ACML | 4 |
| 2021 | An improved ELM-based and data preprocessing integrated approach for phishing detection considering comprehensive features
Liqun Yang, Jiawei Zhang 0001, Xiaozhe Wang, Zhi Li 0045, Zhoujun Li 0001, Yueying He |
Expert Syst. Appl. | 1 |
| 2021 | Deep learning for online AC False Data Injection Attack detection in smart grids: An approach using LSTM-Autoencoder
Liqun Yang, You Zhai, Zhoujun Li 0001 |
J. Netw. Comput. Appl. | 1 |
| 2020 | A11 Your PLCs Belong to Me: ICS Ransomware Is RealisticabstractRansomware is a new business model for cybercrime which mainly targets individual users and machines. Many events have shown how profitable the technique can be. Industrial control systems (ICS) are becoming the next domain. More and more researchers and attackers have become the focus on this field and presented some ICS ransomware. But existing ICS ransomware is theoretically feasible and has a limited effect on real ICS. In this work, we present ICS-BROCK, a full-fledged ICS ransomware that can compromise a real-world. To demonstrate the capability of ICS-BROCK, we use SIEMENS S7-300 PLC, one of the most widely used devices in ICSs, to build a real water treatment environment. The results empirically demonstrate the feasibility of launching ICS ransomware attacks in a practical setting. In the end, we give some suggestions on ICS ransomware to aid in future study and defenses. Liqun Yang, Zhoujun Li 0001, Qiang Zeng 0001, Yueying He, Xiaoming Zhang 0001 |
TrustCom | 3 |
| 2020 | Detecting bi-level false data injection attack based on time series analysis method in smart grid
Liqun Yang, Xiaoming Zhang 0001, Zhi Li 0045, Zhoujun Li 0001, Yueying He |
Comput. Secur. | 1 |
| 2020 | Mixture distribution modeling for scalable graph-based semi-supervised learning
Zhi Li 0045, Chaozhuo Li, Liqun Yang, Philip S. Yu, Zhoujun Li 0001 |
Knowl. Based Syst. | 3 |
| 2019 | Large-scale Detection of Privacy Leaks for BAT Browsers Extensions in ChinaabstractAlthough browser extensions bring users a better experience, it creates a hidden danger of privacy leakage. A common privacy leakage detection method is realized through detecting private data transmission. However, only the unintended transmission is considered to be a privacy leak. Therefore, the real challenge is to determine whether or not the transmission is user intended. In order to address this problem, we check the rationality of private data transmission by establishing a privacy model based on classification for extensions to confirm the scope of private data that can be uploaded and domains that can be sent to. Furthermore, we present BEDS (Browser Extension Detection System), a Chromium based extension dynamic detection system. BEDS first builds a privacy model for each extension and then records the extension's network logs and browser API logs when accessing specified pages. Finally, BEDS determines whether there exists a privacy leak according to the strict privacy leakage judgment rules. We test our implementation in large scale on extensions in browsers developed by China's three major Internet companies and complete 15 months of continuous tracking. After examining a total of 14,487 extensions, 1,897 privacy leaks are identified, all results have been inspected by manual and the accuracy of BEDS is over 97%. A number of domains that illegally collect private user data are discovered and tracked. Our results show that about 47,000 Chinese IPs upload private information to suspicious servers every day. Longtao He, Zhoujun Li 0001, Liqun Yang, Yu Wang 0206 |
TASE | 4 |
| 2015 | A technology mapper for depth-constrained FPGA logic cellsabstractIn the last decade, progress in logic synthesis has brought about new advantageous circuit representations. These representations, such as And-Inverter Graphs in the ubiquitous open-source synthesizer ABC, have inspired new designs of Field Programmable Gate Arrays (FPGAs), which, instead of using Look-Up Tables (LUTs), mimic the topology of the circuit representation in the basic logic cells. More recent examples are Majority-Inverter Graphs, another uniform representation which has triggered considerable interest in synthesis and which naturally suggests new logic cells. Yet, in this paper we observe how naïvely adapting technology mapping solutions for classic LUT-based FPGAs to these new architectures incurs severe shortcomings. The key issue is that LUTs are inherently input-constrained (the logic function they implement is irrelevant) and have generally a single output; on the other hand, logic cells made of uniform networks of some fundamental logic function (e.g., And-Invert) are constrained in terms of logic depth and multiple outputs are an integral feature. We introduce novel and effective solutions to address these differences; the result is a highly versatile mapper—thus enabling further research in these new architectures—with a significantly better performance than what is described in literature for one such architecture. Specifically, when we compare with the state of the art on one sample architecture, we obtain a significant decrease in area (on average 18% over several benchmarks) while also improving slightly the critical path (a reduction of 3%). Zhenghong Jiang, Grace Zgheib, Colin Yu Lin, David Novo, Liqun Yang, Haigang Yang, Paolo Ienne |
FPL | 6 |
| 2014 | Revisiting and-inverter conesabstractAnd-Invert Cones (AICs) have been suggested as an alternative to the ubiquitous Look-Up Tables (LUTs) used in commercial FPGAs. The original article suggesting the new architecture made some untested assumptions on the circuitry needed to implement AIC architectures and did not develop completely the toolset necessary to assess comprehensively the idea. In this paper, we pick up the architecture that some of us proposed in the original AIC paper and try to implement it as thoroughly as we can afford. We build all components for the logic cluster at transistor level in a 40~nm technology as well as a LUT-based architecture inspired by Altera's Stratix~IV. We first determine that the characteristics of our LUT-based architecture are reasonably similar to those of the commercial counterpart. Then, we compare the AIC architecture to the baseline on a number of benchmarks, and we find a few difficulties that had been overlooked before. We thus explore other design possibilities around the original design point and show their detailed impact. Finally, we discuss how the very structure of current logic clusters seems not perfectly appropriate for getting the best out of AICs and conclude that, even though they are not confirmed as an immediate blessing today, AICs still offer rich research opportunities. Grace Zgheib, Liqun Yang, David Novo, Hadi Parandeh-Afshar, Haigang Yang, Paolo Ienne |
FPGA | 2 |
| 2014 | Exploring architecture parameters for dual-output LUT based FPGAsabstractDual-output lookup tables (LUTs) are mainstream in the design of commercial FPGA products. A detailed exploration of architectural parameters of FPGAs based on dualoutput LUTs is presented. Different from traditional single-output LUT based architecture, “shared inputs” between the sub-LUTs is a new parameter specific to dual-output architecture. In this paper, we focus on the effect of ratio of shared inputs on the performance and area-efficiency. First, we study the required cluster inputs and derive a relationship between cluster inputs, LUT size and cluster size under different ratios of shared inputs. Secondly, our evaluation results show that a FPGA with 4-LUTs and a shared input ratio of two thirds is preferred for area-efficiency, while a large LUT size of 9 with no shared inputs achieves best performance. Finally, we determine that a LUT size of 4, a cluster size from 3 to 8, and a shared input ratio between 1/3 and 2/3, provide the best area-delay product for dual-output LUT based FPGAs. Zhenghong Jiang, Colin Yu Lin, Liqun Yang, Haigang Yang |
FPL | 3 |
| 2014 | A semi-supervised modeling approach for performance characterization of FPGA architecturesabstractAn approach to estimate the performance of FPGA architectures is proposed based on semi-supervised model tree algorithm. The proposed approach avoids synthesizing, mapping, packing, placing and routing, which are essential steps in a traditional flow to obtain the performance of FPGA. Thus it is time efficient while the performance predicted maintains quite close to the result obtained through the traditional method (a tool flow called VTR). This can be utilized effectively during the early FPGA design stage to choose an optimal architecture under a certain metric. Comparisons are made between the performance obtained by the proposed approach and by VTR on a commercial 40nm technology. Results show that the proposed approach has MRE below 7.62% compared to VTR, and improves the time cost by thousands of times when utilized in architecture design space exploration. Liqun Yang, Haigang Yang, Wei Li 0008, Colin Yu Lin |
FPL | 1 |
| 1993 | BSD/I18N - Internationalization of the 4.3BSD UNIX system
Jianqiang Zhou, Liqun Yang, Shilei Pan, Hong Tan |
J. Comput. Sci. Technol. | 2 |