Runhao Li

dblp:242/7076 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unigaussian: Driving Scene Reconstruction From Multiple Camera Models Via Unified Gaussian Representations
abstract
Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras. In fact, how to effectively simulate fisheye cameras in driving scene remains an unsolved problem. In this work, we propose UniGaussian, a novel approach that learns a unified 3D Gaussian representation from multiple camera models for urban scene reconstruction in autonomous driving. Our contributions are two-fold. First, we propose a new differentiable rendering method that distorts 3D Gaussians using a series of affine transformations tailored to fisheye camera models. This addresses the compatibility issue of 3D Gaussian splatting with fisheye cameras, which is hindered by light ray distortion caused by lenses or mirrors. Besides, our method maintains real-time rendering while ensuring differentiability. Second, built on the differentiable rendering method, we design a new framework that learns a unified Gaussian representation from multiple camera models. By applying affine transformations to adapt different camera models and regularizing the shared Gaussians with supervision from different modalities, our framework learns a unified 3D Gaussian representation with input data from multiple sources and achieves holistic driving scene understanding. As a result, our approach models multiple sensors (pinhole and fisheye cameras) and modalities (depth, semantic, normal and LiDAR point clouds). Our experiments show that our method achieves superior rendering quality and fast rendering speed for driving scene simulation.
Guile Wu, Runhao Li, Zheyuan Yang, Tongtong Cao, Xingxin Chen
3DV3
2025 Chat4seed: Semantic-Awareness Highly Structured Seed Generation for Fuzzing
abstract
Fuzzing, one of the most popular methods for enhancing software security and quality, relies heavily on the quality of its initial seed corpus to effectively uncover vulnerabilities. Traditional methods for generating initial seed corpora, whether crawl-based or generation-based, struggle with programs that handle highly structured formats with complex semantics, leading to low testing coverage and reduced fuzzer effectiveness. Although researchers have proposed leveraging Large Language Models (LLMs) to create high-quality seed corpora, current approaches are limited to text-based seed files, such as JavaScript code. To address the limitations, we propose Chat4Seed, a novel approach that extends the capabilities of LLMs to produce not only text but also highly-structured binary seed files. Chat4Seed leverages LLMs to extract and interpret the semantic constraints embedded within formatspecific branches of programs. While existing approaches focus solely on generating text-based seeds, Chat4Seed goes further by utilizing LLMs to generate functional library invocation code to produce binary seeds. It also employs binarylevel manipulation to handle unsupported corner cases, achieving robust seed generation for both text and binary formats. Our experiments and evaluations on 12 real-world programs demonstrate that after 48 hours of fuzz testing using AFL++, the seed corpus generated by Chat4Seed achieves an average increase in coverage of 28.78% compared to traditional crawl-based methods and 39.98% compared to generation-based methods. Additionally, it facilitates the discovery of 84.24% and 392.7% more crashes than seed corpus generated by traditional approaches, respectively, underscoring the potential of Chat4Seed to enhance the efficacy of fuzzing.
Jiarui Chen, Jiongyi Chen, Runhao Li, Chaojing Tang
QRS4
2025 Online weighted hashing for cross-modal retrieval
Zining Jiang, Zhenyu Weng, Runhao Li, Huiping Zhuang, Zhiping Lin 0001
Pattern Recognit.3
2025 Class-Specific Prompt Learning for Vision-Language Models
abstract
The use of learning prompts to adapt pretrained vision-language models (VLMs) for downstream tasks has gained significant attention due to its potential to reduce training costs compared to model fine-tuning through few-shot learning. Most existing methods rely on a universal prompt for all classes, as it generally delivers consistent performance across various datasets. However, a universal prompt cannot capture class-specific discriminative information. To overcome this limitation, we propose class-specific prompt learning (CPL). CPL represents the context of a prompt using two components: a base vector shared among all classes and a class-specific vector designed for individual classes. This method combines the generalization ability of the base context with the adaptability of the class-specific context. Furthermore, we introduce contrastive CPL, which enhances the ability of the prompt to capture discriminative features unique to each class. Also, we adopt the self-consistency loss to regularize the base context, enhancing its generalization ability. As a result, CPL effectively learns tailored prompts for each class. Extensive experiments demonstrate that CPL achieves superior performance over existing methods in both base-class classification and new class generalization.
Runhao Li, Yongming Chen, Zhenyu Weng, Zhiping Lin 0001, Yap-Peng Tan
IEEE Trans. Neural Networks Learn. Syst.1
2024 Joint-Neighborhood Product Quantization for Unsupervised Cross-Modal Retrieval
abstract
Product quantization (PQ) is a technique that transforms high-dimensional data into compact binary codes to reduce data storage and improve search efficiency. However, existing PQ methods separate the learning of modality-specific features from the learning of quantization codewords, resulting in suboptimal performance in cross-modal retrieval tasks. In this paper, we propose a joint-neighborhood product quantization (JNPQ) method to simultaneously learn modality-specific features and quantization codewords. To achieve this, we first introduce a cross-modal quantization contrastive learning module that preserves the inter-modal neighborhood of the original data and reduces the quantization error. Then, we design a self-neighbor contrastive learning module that enhances the intra-modal neighborhood within individual modalities. Extensive experiments demonstrate that JNPQ achieves state-of-the-art results in crossmodal retrieval when compared with other unsupervised crossmodal quantization methods.
Runhao Li, Zhenyu Weng, Yongming Chen, Huiping Zhuang, Yap-Peng Tan, Zhiping Lin 0001
VCIP1
2023 Neighborhood Learning from Noisy Labels for Cross-Modal Retrieval
abstract
Cross-modal retrieval methods are developed to retrieve relevant data across different modalities. Usually, super-vised cross-modal retrieval methods can achieve higher accuracy than unsupervised methods because they can utilize the semantic information provided by clean labels. However, training data with noisy labels will lead to the performance degradation of supervised cross-modal retrieval methods. In this work, we present a novel framework called Neighborhood Learning for Cross-Modal Retrieval (NLCMR) that is robust against noisy labels by exploiting the information contained in the neighbor-hood. Our NLCMR contains two main components: Clustering with Neighborhood Alignment and Neighborhood Contrastive Learning. The first component focuses on reducing the impact of noisy labels and improving clustering robustness, and the second component learns from noisy data by exploring pairwise and neighborhood information. Extensive experiments are conducted on three multi-modal datasets to demonstrate the effectiveness of NLCMR.
Runhao Li, Zhenyu Weng, Huiping Zhuang, Yongming Chen, Zhiping Lin 0001
ISCAS1
2023 Towards Automatic and Precise Heap Layout Manipulation for General-Purpose Programs
Runhao Li, Jiongyi Chen, Wenfeng Lin, Chao Feng 0002, Chaojing Tang
NDSS1
2023 Automated Exploitable Heap Layout Generation for Heap Overflows Through Manipulation Distance-Guided Fuzzing
Jiongyi Chen, Runhao Li, Chao Feng 0002, Ruilin Li 0002, Chaojing Tang
USENIX Security Symposium3
2022 Tri-AoA: Robust AoA Estimation of Mobile RFID Tags With COTS Devices
abstract
Radio frequency identification (RFID) is a form of wireless communication that has received much attention in recent years due to low costs of passive RFID tags and availability of commercial-off-the-shelf (COTS) RFID devices. Existing indoor localization and tracking methods based on RFID do not perform well in dynamic environments with severe multi-path interference. In this paper, we propose a robust Angle of Arrival (AoA) estimation method for mobile RFID tags in a rich multi-path environment with a large feasible area. The proposed method Tri-AoA consists of three essential modules, phase likelihood estimation, Received Signal Strength Indicator (RSSI) likelihood estimation and a deep learning algorithm. The phase likelihood estimation module exploits the concept of an antenna array to provide a basic estimation of an AoA, but with an ambiguity. The RSSI likelihood estimation module helps alleviate the ambiguity. To achieve a more robust estimation of AoA for mobile RFID tags, we construct a 2-dimensional feature image that contains AoA estimation from the phase and RSSI modules. We then develop a deep learning algorithm to analyze this image to improve the AoA tracking accuracy as well as the robustness by suppressing the multi-path interference. The experimental results show that our system outperforms existing approaches by achieving a median error of$2.36^{\mathrm{o}}$in a$3m\times 4m$area using four COTS RFID antennas. We also show that our system can realize real-time performance on a personal computer.
Runhao Li, Rongzihan Song, Benaya Christo, Lei Sun 0006, Zhiping Lin 0001
GLOBECOM2