Longfei Huang

dblp:216/8320 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diffusion for automatic matting
Longfei Huang, Hao Zhang 0063, Xinning Zhu, Lunde Chen
Neurocomputing1
2026 DyCoT-RE: Chain-of-Thought-enhanced LLM reward engineering with dual-dynamic optimization for reinforcement learning
Xinning Zhu, Jinxin Du, Longfei Huang, Lunde Chen
Neurocomputing3
2025 Multimodal Semantic Decoupled Prompt for Zero-Shot Referring Expression Comprehension
abstract
Large-scale Vision-Language Models (VLMs) have demonstrated impressive zero-shot performance in sample-level downstream tasks (e.g., image classification), driven by their powerful generalization ability. However, they still struggle in instance-level tasks, e.g., zero-shot Referring Expression Comprehension (REC), which requires precisely locating the target instance in an image based on a provided text caption. To address this issue, we propose Multimodal Semantic Decoupled Prompting (MSDP), a simple yet effective prompt engineering approach that contains both textual- and visual-focused instance-level understanding prompting. Specifically, we first propose a novel textual restructure strategy to eliminate the impact of task-irrelevant semantic information, steering the model’s attention at the textual understanding level. Meanwhile, we design a united visual prompt at the visual understanding level that maximally activates the instance-level understanding capabilities of VLMs. Experiments on several benchmarks reveal that the proposed approach outperforms state-of-the-art (SOTA) methods. The code is available at repository.
Longfei Huang
ECAI2
2025 SDMATTE: Grafting Diffusion Models for Interactive Matting
abstract
Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge regions. Diffusion models trained on billions of image-text pairs, demonstrate exceptional capability in modeling highly complex data distributions and synthesizing realistic texture details, while exhibiting robust text-driven interaction capabilities, making them an attractive solution for interactive matting. To this end, we propose SDMatte, a diffusion-driven interactive matting model, with three key contributions. First, we exploit the powerful priors of diffusion models and transform the text-driven interaction capability into visual prompt-driven interaction capability to enable interactive matting. Second, we integrate coordinate embeddings of visual prompts and opacity embeddings of target objects into U-Net, enhancing SDMatte's sensitivity to spatial position information and opacity information. Third, we propose a masked self-attention mechanism that enables the model to focus on areas specified by visual prompts, leading to better performance. Extensive experiments on multiple datasets demonstrate the superior performance of our method, validating its effectiveness in interactive matting. Our code and model are available at https://github.com/vivoCameraResearch/SDMatte.
Longfei Huang, Hao Zhang 0063, Jinwei Chen 0003, Lunde Chen, Peng-Tao Jiang
ICCV1
2025 Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion
abstract
Multimodal learning (MML) is significantly constrained by modality imbalance, leading to suboptimal performance in practice. While existing approaches primarily focus on balancing the learning of different modalities to address this issue, they fundamentally overlook the inherent disproportion in model classification ability, which serves as the primary cause of this phenomenon. In this paper, we propose a novel multimodal learning approach to dynamically balance the classification ability of weak and strong modalities by incorporating the principle of boosting. Concretely, we first propose a sustained boosting algorithm in multimodal learning by simultaneously optimizing the classification and residual errors. Subsequently, we introduce an adaptive classifier assignment strategy to dynamically facilitate the classification performance of the weak modality. Furthermore, we theoretically analyze the convergence property of the cross-modal gap function, ensuring the effectiveness of the proposed boosting scheme. To this end, the classification ability of strong and weak modalities is expected to be balanced, thereby mitigating the imbalance issue. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SOTA) multimodal learning baselines. The source code is available at https://github.com/njustkmg/NeurIPS25-AUG.
Qing-Yuan Jiang, Longfei Huang, Yang Yang 0074
NeurIPS2
2025 A lightweight intrusion detection system for connected autonomous vehicles based on ECANet and image encoding
Zhuoqun Xia, Longfei Huang, Jingjing Tan, Wei Hao 0002, Kejun Long
J. Inf. Secur. Appl.2
2025 DATI-IDS: Domain Adaptation and Time-Series Imaging-Based Intrusion Detection System for Connected Autonomous Vehicles
abstract
With the advancement of artificial intelligence, automobiles are progressively transitioning from traditional mechanization to Connected Autonomous Vehicles (CAVs), significantly enhancing driving comfort and safety. As the standard communication protocol in CAVs, the Controller Area Network (CAN) remains vulnerable to attacks due to the lack of robust security mechanisms. While existing deep learning-based vehicle network intrusion detection systems can effectively identify known attacks, their ability to detect unknown attacks is limited due to the same data distribution in the source and target domain. To address this issue, we propose a domain adaptation and time-series imaging-based intrusion detection system (DATI-IDS) to detect known and unknown attacks, where the deep domain adaptation method is used to solve the source and target domain data distribution difference problem by optimizing the multiple kernel maximum mean discrepancy (MK-MMD) between the source domain and target domain images and the classification loss, and the time-series imaging method is used to capture temporal dependencies and improve efficiency by transforming the CAN ID sequence into a two-dimensional gramian angular summation field (GASF) image. The effectiveness of the proposed model is evaluated across nine distinct unknown attack scenarios using the Car-Hacking dataset and the survival analysis dataset. Comparative analysis with previous studies demonstrates superior performance, faster inference times, and reduced model complexity.
Jingjing Tan, Longfei Huang, Zhuoqun Xia, Ke Gu 0002, Wei Hao 0002, Kejun Long, Lingxuan Zeng
IEEE Trans. Intell. Transp. Syst.2
2025 Conditional Data-Sharing Privacy-Preserving Scheme in Blockchain-Based Social Internet of Vehicles
abstract
Social Internet of Vehicles (SIoVs) is an important information exchange platform to provide comprehensive traffic services by sharing vehicle-aware data. However, traditional data sharing methods can not provide the security of decentralized data sharing, making it possible for some malicious third parties to initiate dishonest behaviors. Additionally, the lack of access control for data sharing in SIoVs easily leads to unauthorized data sharing, thus user privacy is threatened and the source of false data is difficult to be traced. In this paper, we propose a conditional data-sharing privacy-preserving scheme for blockchain-based social internet of vehicles. In our scheme, a lightweight ledger-based blockchain system is designed, which combines with the ciphertext-policy attribute-based encryption method to realize anonymous one-to-many sharing of data with fine-grained access management. Also, a collaborative identity tracing method is constructed to trace malicious users who provide false data. Our scheme can effectively prevent second-hand data sharing and safeguard user privacy. Moreover, related experimental results validate the efficiency of our scheme.
Zhuoqun Xia, Jiahuan Man, Ke Gu 0002, Xiong Li 0002, Longfei Huang
IEEE Trans. Sustain. Comput.5
2024 Refining Visual Perception for Decoration Display: A Self-Enhanced Deep Captioning Model
Longfei Huang, Weili Guo, Yang Yang 0074
ACML1
2024 Unlocking Versatile Locomotion: A Novel Quadrupedal Robot with 4-DoFs Legs for Roller Skating
abstract
Roller skating with passive wheels on a quadrupedal robot is more efficient than traditional walking. However, the typical mammalian quadruped robot with 3-DoFs legs can only perform one dynamic roller skating gait and has difficulty achieving turning motion. To address this limitation, we designed a novel quadrupedal robot with each leg having 4-DoFs to enable various roller skating locomotion including Swizzling, Stroking, and trot-like gaits while easily achieving turning motions. We considered the geometrical characteristics of the passive wheel and used the Levenberg-Marquardt method in robot kinematics to improve precision for both roller skating kinematics and contact point position for the dynamics controller. The position of the robot foot and the yaw angle of the passive wheel are decoupled for motion planning of all proposed gaits. Our proposed kinematics with wheeled geometry was verified through experiments to have higher precision, while the feasibility of all proposed roller-skating gaits was confirmed during straight motion and turning motion with a small radius on our prototype robot. Finally, we discussed the mobility efficiency of different roller skating gaits which were found to be more efficient than walking.
Ripeng Qin, Longfei Huang, Zongbo He, Kun Xu 0007, Xilun Ding
ICRA3
2024 TIDL-IDS: A Time-Series Imaging and Deep Learning-Based IDS for Connected Autonomous Vehicles
Zhuoqun Xia, Longfei Huang, Jingjing Tan, Faqun Jiang, Zhenzhen Hu 0002
ISC (2)2
2020 Semantic trajectory segmentation based on change-point detection and ontology
abstract
Trajectory segmentation is a fundamental issue in GPS trajectory analytics. The task of dividing a raw trajectory into reasonable sub-trajectories and annotating them based on moving subject’s intentions and application domains remains a challenge. This is due to the highly dynamic nature of individuals’ patterns of movement and the complex relationships between such patterns and surrounding points of interest. In this paper, we present a framework called SEMANTIC-SEG for automatic semantic segmentation of trajectories from GPS readings. For the decomposition component of SEMANTIC-SEG, a moving pattern change detection (MPCD) algorithm is proposed to divide the raw trajectory into segments that are homogeneous in their movement conditions. A generic ontology and a spatiotemporal probability model for segmentation are then introduced to implement a bottom-up ontology-based reasoning for semantic enrichment. The experimental results on three real-world datasets show that MPCD can more effectively identify the semantically significant change-points in a pattern of movement than four existing baseline methods. Moreover, experiments are conducted to demonstrate how the proposed SEMANTIC-SEG framework can be applied.
Yuan Gao 0045, Longfei Huang, Jun Feng 0003, Xin Wang 0004
Int. J. Geogr. Inf. Sci.2