EDBT 2026 Demo / reviewers in the wild / expert
Litao Li
dblp:233/6749
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Systems and software security · 70% Malware analysis · 30% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Systems and software security
binary analysis |
0.9 | 1 | 2025 | Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection · NeurIPS 2025 |
Systems and software security › binary analysis
binary code similarity detection |
0.9 | 1 | 2025 | Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection · NeurIPS 2025 |
Program analysis › code representation learning
code embedding |
0.9 | 1 | 2025 | Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection · NeurIPS 2025 |
Malware analysis › malware classification
malware family classification |
0.2 | 1 | 2022 | Adversarial Variational Modality Reconstruction and Regularization for Zero-Day Malware Variants Similarity Detection · ICDM 2022 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.7fine-tuning · 1.7data augmentation · 1.7contrastive learning · 1.7siamese network · 0.6multi-modality learning · 0.6generative adversarial network · 0.6conditional variational autoencoder · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GDFDNet: A Novel Graph-Based Dynamically Fused Dual-Stream Network for Accuracy Prohibited Items DetectionabstractIn various security inspection scenarios, prohibited items detection in X-ray images is of great significance for safeguarding public safety and effectively reducing the potential risks of crimes and terrorist activities. However, existing detection methods still face the challenge of overlapping image features when distinguishing prohibited items from complex backgrounds. To tackle this issue, this paper proposes a novel Graph-based Dynamically Fused Dual-Stream Network (GDFDNet). This network introduces highly decoupled HSV color space features, which are dynamically fused with RGB color space features to fully explore the edge and material characteristics of overlapping prohibited items. Specifically, we adopt a dynamically fused RGB-HSV dual-stream framework and design two key modules: a Graph-based Edge-Aware module (GEA) and a Graph-based Material-Aware Fusion module (GMAF). The former is utilized to enhance multi-scale edge features of prohibited items, while the latter is responsible for performing the dynamic fusion of features in dual color spaces guided by edge features, and then extracting valuable material information of prohibited items from the overlapping foreground and background features. Extensive experiments on the SIXray and PIDray datasets demonstrate that our proposed network significantly outperforms the existing state-of-the-art methods. Hongxia Gao, Yaobin Huang, Zhenming Guan, Litao Li |
ICASSP | 5 |
| 2025 | Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionabstractCybersecurity and software research have crossed paths with modern deep learning research for a few years. The power of large language models (LLMs) in particular has intrigued us to apply them to understanding binary code. In this paper, we investigate some of the many ways LLMs can be applied to binary code similarity detection, as it is a significantly more difficult task compared to source code similarity detection due to the sparsity of information and less meaningful syntax. It also has great practical implications, such as vulnerability and malware detection. We find that pretrained LLMs are mostly capable of detecting similar binary code, even with a zero-shot setting. Our main contributions and findings are to provide several supervised fine-tuning methods that, when combined, significantly surpass zero-shot LLMs and state-of-the-art binary code similarity detection methods. Specifically, we up-train the model through data augmentation, translation-style causal learning, LLM2Vec, and cumulative GTE loss. With a complete ablation study, we show that our training method can transform a generic language model into a powerful binary similarity expert, and is also robust and general enough for cross-optimization, cross-architecture, and cross-obfuscation detection. Litao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland |
NeurIPS | 1 |
| 2025 | Correction Method for Low-Frequency Time-Varying Relative Radiometric Errors in Optical Remote Sensing: A Case Study of JiTian-03abstractWhile systematic radiometric errors can be mitigated through on-orbit calibration, correcting low-frequency, time-varying errors often induced by complex imaging dynamics remains a greater challenge. This study introduces a correction method tailored for optical remote sensing platforms to overcome such errors in JiTian-03 imagery, particularly those arising from high-agility curve imaging. The process begins with relative radiometric calibration to correct high-frequency inconsistencies among sensor detectors, thereby suppressing striping and banding noise in the imagery. Next, the image is decomposed into low-frequency background and high-frequency detail components. A column-wise moment-matching algorithm is then applied to the low-frequency background, aligning each column’s first- and second-order central moments with global statistics to derive correction coefficients. This approach effectively compensates for low-frequency radiometric errors while preserving high-frequency texture details in the image. Experimental results confirm that, following correction of both high- and low-frequency radiometric errors, JT-03 imagery exhibits a substantial reduction in striping, banding, and brightness nonuniformity. The correction yields a uniform radiometric response across the entire field of view while preserving ground texture fidelity and image detail. As a result, the radiometric uniformity of JT-03 image products is significantly enhanced. Quantitatively, post high-frequency calibration, the striping index across all JT-03 bands improved to within 0.1% and the standard deviation of column means to 4.7%. Subsequent low-frequency correction reduced the standard deviation of column means by an average of 73.89%, achieving a final value of better than 1.3%. Litao Li, Haoqi Liang, Yonghua Jiang 0001, Xin Shen 0001, Jiayang Cao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Multistrip Stitching Imaging Mission Planning Method for SAR Satellite Regional Mapping Considering Onboard Energy ConsumptionabstractTo meet the demands of regional mapping with synthetic aperture radar (SAR) satellites, this study proposes a multistrip stitching imaging mission planning (MSIMP) method that considers energy consumption. Accounting for both regional coverage rate and satellite energy consumption, an MSIMP model was constructed. The model maximizes the region coverage benefit (RCB), minimizes the total imaging time (TIT) as the optimization objectives, and sets the indices of orbit selection, side-swing angle coefficients, and determination coefficients of the start time and end times of each orbit as the decision variables to realize the overall optimal scheduling of available orbital resources (ORs), side-swing angles, and imaging time of satellites. To address the problem of the large-scale and diverse types of decision variables in the MSIMP model, an improved particle swarm optimization (PSO) algorithm with a hybrid local search and differential evolution (LSDE) strategy (LSDE-PSO) that implements differential operations and local searches on two types of particles at various stages of the evolution process to enhance the model’s solution efficiency was proposed. The proposed method was validated using three simulation scenarios with different regional sizes. The experimental results showed that the proposed method demonstrated consistent adaptability to different-scale tasks, achieving the synchronized optimization of coverage revenue and on-orbit energy consumption. Compared with existing algorithms, the proposed LSDE-PSO algorithm can obtain the optimal imaging scheme with a higher RCB with lower TIT at almost the same computational cost, which provides significant technical support for mission planning for regional SAR satellite imaging in practical applications. Xin Shen 0001, Zezhong Lu, Litao Li, Yaxin Chen, Xufei Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | 114Xray: A Large-Scale X-Ray Security Detection Benchmark and Aware Enhance Network for Real-World Prohibited Item Inspection in Baggage
Hongxia Gao, Zhenming Guan, Yaobin Huang, Hongyu Liao, Hongzhen Zheng, Runze Lin, Litao Li, Haolin Tang, Guoyuan Lin, Zhanhong Chen |
PRCV (11) | 9 |
| 2023 | VulANalyzeR: Explainable Binary Vulnerability Detection with Multi-task Learning and Attentional Graph ConvolutionabstractSoftware vulnerabilities have been posing tremendous reliability threats to the general public as well as critical infrastructures, and there have been many studies aiming to detect and mitigate software defects at the binary level. Most of the standard practices leverage both static and dynamic analysis, which have several drawbacks like heavy manual workload and high complexity. Existing deep learning-based solutions not only suffer to capture the complex relationships among different variables from raw binary code but also lack the explainability required for humans to verify, evaluate, and patch the detected bugs. We propose VulANalyzeR, a deep learning-based model, for automated binary vulnerability detection, Common Weakness Enumeration-type classification, and root cause analysis to enhance safety and security. VulANalyzeR features sequential and topological learning through recurrent units and graph convolution to simulate how a program is executed. The attention mechanism is integrated throughout the model, which shows how different instructions and the corresponding states contribute to the final classification. It also classifies the specific vulnerability type through multi-task learning as this not only provides further explanation but also allows faster patching for zero-day vulnerabilities. We show that VulANalyzeR achieves better performance for vulnerability detection over the state-of-the-art baselines. Additionally, a Common Vulnerability Exposure dataset is used to evaluate real complex vulnerabilities. We conduct case studies to show that VulANalyzeR is able to accurately identify the instructions and basic blocks that cause the vulnerability even without given any prior knowledge related to the locations during the training phase. Litao Li, Steven H. H. Ding, Yuan Tian 0008, Benjamin C. M. Fung, Philippe Charland, Weihan Ou, Leo Song, Congwei Chen |
ACM Trans. Priv. Secur. | 1 |
| 2022 | Adversarial Variational Modality Reconstruction and Regularization for Zero-Day Malware Variants Similarity DetectionabstractMatching malware variants in the same malware family has always been a significant challenge for Cyber Threat Intelligence (CTI). For zero-day malware that does not belong to an existing family, a timely matching of its variants is essential for effective threat tracing and prompt response to the cyber incident. However, malware variants are of diverse forms that make them difficult to match. Additionally, the information extracted from a given malware sample is inaccurate, especially on zero-day malware. Existing malware solutions only focus on detecting known malware or find if two samples are similar without creating any reusable representation of the samples. In this paper, we propose the first practical and efficient solution for zero-day malware variant matching with reconstruction. By combining multi-modality learning and a Siamese-based structure, our model can navigate across different modalities and match zero-day variants. To address the missing or noisy modality issue, we propose a Conditional Variable Autoencoder with a Generative Adversarial Network for heightened resolution. We trained the model on 100,000 malware triplet pairs. Our experiments on real-world noisy samples show that the model out-performs the state-of-the-art and can accurately match not only zero-day malware, but also out-of-sample benign binaries of the same category. Christopher Molloy, Jeremy Banks, Steven H. H. Ding, Philippe Charland, Andrew Walenstein, Litao Li |
ICDM | 6 |
| 2022 | TASR: Adversarial learning of topic-agnostic stylometric representations for informed crisis response through social mediaabstractThe impact of crisis events can be devastating in a multitude of ways, many of which are unpredictable due to the suddenness in which they occur. The evolution of social media (for example Twitter) has given directly affected individuals or those with valuable information a platform to effectively share their stories to the masses. As a result, these platforms have become vast repositories of helpful information for emergency organizations. However, different crisis events often contain event-specific keywords, which results in the difficult extraction of useful information with a single model. In this paper, we put forward TASR, which stands for Topic-Agnostic Stylometric Representations, a novice deep learning architecture that uses stylometric and adversarial learning to remove topical bias to better manage the unknown surrounding unseen events. As an alternative to domain adaptive approaches requiring data from the unseen event, it reduces the work for those responding to the onset of a crisis. Overall, we conduct a comprehensive study of the situational properties of TASR, the benefits of its architecture including its topic-agnostic and explainable properties, and how it improves upon comparable models in past research. From two experiments, on average, TASR is able to outperform state-of-the-art methods such as transfer learning and domain adoption by 11% in AUC. The ablation study illustrates how different architecture choices of TASR impact the results and that TASR has been optimized for this task. Finally, we conduct a case study to show that explainable results from our model can be used to help guide human analysts through crisis information extraction. Litao Li, Rylen Sampson, Steven H. H. Ding, Leo Song |
Inf. Process. Manag. | 1 |
| 2022 | An Improved On-Orbit Relative Radiometric Calibration Method for Agile High-Resolution Optical Remote-Sensing Satellites With Sensor Geometric DistortionabstractNoticeable striping artifacts in collected satellite imagery, caused by the inconsistency of pushbroom sensor detector responses, degrade the image quality. Relative radiometric calibration aims to calibrate inconsistencies in terms of detector responses, thereby eliminating detector-level striping artifacts. Yaw calibration has become the preferred imagery-based calibration method for satellites without on-board calibration equipment, owing to its high accuracy and convenience. However, it often relies on a large uniform field on the Earth’s surface and is a linear calibration method. This method does not consider acquisition-related geometric errors associated with the sensor, resulting in reduced calibration accuracy. In this study, we propose an improved method for yaw calibration, which accounts for the geometric distortion of sensor imaging and does not require a uniform field. A fast geometric-positioning algorithm was used to remove geometric distortion, followed by high-precision extraction of the calibration reference for each sensor’s detector. The dynamic range of the sensor calibration was extended, followed by the calibration of the sensor nonlinear response model. Our results based on Yaogan-25 images suggest the following: 1) the improved method effectively eliminates the “sawtooth” caused by the yaw calibration method without considering the geometric distortion and 2) it outperforms other imagery-based calibration methods with respect to visual destriping effects such that the root-mean-square deviations of the corrected imagery experienced a decrease of 0.65 percentage point, compared with that of the statistical method, and 0.17 percentage point for the yaw calibration. Such improvements will promote high-quality applications of remote-sensing images. Litao Li, Guo Zhang 0001, Yonghua Jiang 0001, Xin Shen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Translution-SNet: A Semisupervised Hyperspectral Image Stripe Noise Removal Based on Transformer and CNNabstractHyperspectral remote sensing images (HSIs) have been applied in urban planning, environmental monitoring, and other fields. However, they are susceptible to noise interference, such as Gaussian noise, stripe, and mixed noises, from various factors in the imaging process, which greatly limits their applications. Although previous efforts to improve HSI quality have achieved remarkable results, there are still many challenges to be solved. To avoid the poor generalization ability and improve the stripe removal performance of the network in real scenarios. In this paper, we proposed a novel deep learning model (Translution-SNet) for HSI stripe noise removal based on a semi-supervised training strategy that applies a convolution and transformer for feature extraction. Moreover, we used an unbiased estimation method to calculate the loss function of the unsupervised part from noisy data without a clean image. The semi-supervised method improved the ability of Translution-SNet to deal with various complex stripe noises during stripe removal and strengthened its robustness and generalization ability. Our experimental results showed that Translution-SNet could robustly handle stripe noise of images with different loads and achieve satisfactory results, proving its feasibility and effectiveness. In addition, Translution-SNet showed good generalization ability. Miaozhong Xu, Yonghua Jiang 0001, Guo Zhang 0001, Hao Cui 0002, Litao Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Multiscale Intensity Propagation to Remove Multiplicative Stripe Noise From Remote Sensing ImagesabstractSensor instability, dark currents, and other factors often cause stripe noise corruption in hyperspectral remote sensing images and severely limit their application in practical purposes. Previous studies have proposed numerous destriping algorithms that have yielded impressive results. Although most destriping algorithms are based on the premise of additive noise, a few studies have focused directly on multiplicative stripe noise. This article fully analyzes the characteristics of the stripe noise of OHS-01 images and proposes a multiplicative stripe noise removal method. Specifically, stripe noise is tackled by performing radiometric normalization of different columns in the image. First, the relative gain coefficients of adjacent columns are separated based on prior knowledge. Second, the local relative intensity correspondence of the image columns are established by means of intensity propagation, intensity connection, and so on. Finally, the above-mentioned process is iterated in multiscale space, and the accumulated gain correction coefficient maps were used to correct the radiation of the original image. The results of extensive experiments on simulated and real remote sensing image data demonstrate that the proposed method can, in most cases, yield desirable results. In certain cases, the results are even better, visually, and quantitatively, than those obtained using classical algorithms. Moreover, the proposed method has high robustness and efficiency. Thus, it can conform to the requirements of engineering applications. Hao Cui 0002, Peng Jia 0006, Guo Zhang 0001, Yonghua Jiang 0001, Litao Li, Jingyin Wang, Xiaoyun Hao |
IEEE Trans. Geosci. Remote. Sens. | 5 |