EDBT 2026 Demo / reviewers in the wild / expert
Yuqing Huang
dblp:134/5853
· DBLP profile ↗
21ranked-venue papers
2as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Superimposed Noise Accumulation Problem in Sequential Knowledge Editing of Large Language ModelsabstractSequential knowledge editing techniques aim to continuously update knowledge in large language models at low cost, preventing models from generating outdated or incorrect information. However, existing sequential editing methods suffer from a significant decline in editing success rates after long-term editing. Through theoretical analysis and experiments, our findings reveal that as the number of edits increases, the model's output increasingly deviates from the desired target, leading to a drop in editing success rates. We refer to this issue as the superimposed noise accumulation problem. Our further analysis demonstrates that the problem is related to the erroneous activation of irrelevant knowledge and conflicts between activated knowledge. Based on this analysis, a method named DeltaEdit is proposed that reduces conflicts between knowledge through dynamic orthogonal constraint strategies. Experiments show that DeltaEdit significantly reduces superimposed noise, achieving a 16.8% improvement in editing performance over the strongest baseline. Ding Cao, Yuqing Huang, Xuesong He, Rongxi Guo, Guiquan Liu, Guangzhong Sun |
AAAI | 3 |
| 2026 | Steering into Weights: Persistent Behavioral Editing for Targeted Harm Categories
Ding Cao, Yuqing Huang, Guiquan Liu |
ICIC (24) | 4 |
| 2026 | A transformer fusion framework with intra-modal local and inter-modal global attention for image-text multimodal classification
Yuqing Huang, Wencheng Lin, Hong Zhao 0002, Yifeng Zheng 0004, Wenjie Zhang 0003 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | CMFS: CLIP-Guided Modality Interaction for Mitigating Noise in Multi-Modal Image Fusion and SegmentationabstractInfrared-visible image fusion and semantic segmentation are pivotal tasks for robust scene understanding under challenging conditions such as low light. However, existing methods often struggle with high noise, modality inconsistencies, and inefficient cross-modal interactions, limiting fusion quality and segmentation accuracy. To this end, we propose CMFS, a unified framework that leverages CLIP-guided modality interaction to mitigate noise in multi-modal image fusion and segmentation. Our approach features a region-aware Modal Interaction Alignment module that combines a VMamba-based encoder with an additional shuffle layer to obtain more robust features and a CLIP-guided, regionally constrained multi-modal feature interaction block to emphasize foreground targets while suppressing low-light noise. Additionally, a Frequency-Spatial Collaboration module uses selective scanning and integrates wavelet-, spatial-, and Fourier-domain features to achieve adaptive denoising and balanced feature allocation. Furthermore, we employ a low-rank mixture-of-experts with dynamic routing to improve region-specific fusion and enhance pixel-level accuracy. Extensive experiments on several benchmarks show that, compared with state-of-the-art methods, the proposed approach demonstrates effectiveness in both image fusion quality and semantic segmentation accuracy, especially in complex environments. The source code will be released at IJCAI2025-CMFS. Guilin Su, Yuqing Huang, Zhenyu He 0001 |
IJCAI | 2 |
| 2025 | From Indicators to Insights: Diversity-Optimized for Medical Series-Text Decoding via LLMsabstractMedical time-series analysis differs fundamentally from general ones by requiring specialized domain knowledge to interpret complex signals and clinical context.
Large language models (LLMs) hold great promise for augmenting medical time-series analysis by complementing raw series with rich contextual knowledge drawn from biomedical literature and clinical guidelines.
However, realizing this potential depends on precise and meaningful prompts that guide the LLM to key information.
Yet, determining what constitutes effective prompt content remains non-trivial—especially in medical settings where signal interpretation often hinges on subtle, expert-defined decision-making indicators.
To this end, we propose InDiGO, a knowledge-aware evolutionary learning framework that integrates clinical signals and decision-making indicators through iterative optimization.
Across four medical benchmarks, InDiGO consistently outperforms prior methods.
The code is available at: https://github.com/jinxyBJTU/InDiGO. Xiyuan Jin, Jing Wang 0060, Ziwei Lin, Qianru Jia, Yuqing Huang, Xiaojun Ning 0001, Zhonghua Shi, Youfang Lin |
NeurIPS | 5 |
| 2025 | LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream TrackersabstractTransformer-based algorithms, such as LoRAT, have significantly enhanced object-tracking performance. However, these approaches rely on a standard attention mechanism, which incurs quadratic token complexity, making real-time inference computationally expensive. In this paper, we introduce LoRATv2, a novel tracking framework that addresses these limitations with three main contributions.
First, LoRATv2 integrates frame-wise causal attention, which ensures full self-attention within each frame while enabling causal dependencies across frames, significantly reducing computational overhead. Moreover, key-value (KV) caching is employed to efficiently reuse past embeddings for further speedup.
Second, building on LoRAT's parameter-efficient fine-tuning, we propose Stream-Specific LoRA Adapters (SSLA). As frame-wise causal attention introduces asymmetry in how streams access temporal information, SSLA assigns dedicated LoRA modules to the template and each search stream, with the main ViT backbone remaining frozen. This allows specialized adaptation for each stream's role in temporal tracking.
Third, we introduce a two-phase progressive training strategy, which first trains a single-search-frame tracker and then gradually extends it to multi-search-frame inputs by introducing additional LoRA modules. This curriculum-based learning paradigm improves long-term tracking while maintaining training efficiency.
In extensive experiments on multiple benchmarks, LoRATv2 achieves state-of-the-art performance, substantially improved efficiency, and a superior performance-to-FLOPs ratio over state-of-the-art trackers.
The code is available at https://github.com/LitingLin/LoRATv2. Liting Lin, Heng Fan 0001, Yuqing Huang, Yaowei Wang 0001, Yong Xu 0007, Haibin Ling |
NeurIPS | 4 |
| 2025 | RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question AnsweringabstractIn real-world scenarios, providing user queries with visually enhanced responses can considerably benefit understanding and memory, underscoring the great value of interleaved image-text generation. Despite recent progress, like the visual autoregressive model that unifies text and image processing in a single transformer architecture, generating high-quality interleaved content remains challenging. Moreover, evaluations of these interleaved sequences largely remain underexplored, with existing benchmarks often limited by unimodal metrics that inadequately assess the intricacies of combined image-text outputs. To address these issues, we present RAG-IGBench, a thorough benchmark designed specifically to evaluate the task of Interleaved Generation based on Retrieval-Augmented Generation (RAG-IG) in open-domain question answering. RAG-IG integrates multimodal large language models (MLLMs) with retrieval mechanisms, enabling the models to access external image-text information for generating coherent multimodal content. Distinct from previous datasets, RAG-IGBench draws on the latest publicly available content from social platforms and introduces innovative evaluation metrics that measure the quality of text and images, as well as their consistency. Through extensive experiments with state-of-the-art MLLMs (both open-source and proprietary) on RAG-IGBench, we provide an in-depth analysis examining the capabilities and limitations of these models. Additionally, we validate our evaluation metrics by demonstrating their high correlation with human assessments. Models fine-tuned on RAG-IGBench's training set exhibit improved performance across multiple benchmarks, confirming both the quality and practical utility of our dataset. Our benchmark is available at https://github.com/zry13/RAG-IGBench. Rongyang Zhang, Yuqing Huang, Chengqiang Lu, Qimeng Wang, Yan Gao 0017, Yao Hu 0002, Hao Wang 0076, Enhong Chen |
NeurIPS | 2 |
| 2025 | Non-uniform sampling reconstruction for symmetrical NMR spectroscopy by exploiting inherent symmetry
Enping Lin, Ze Fang, Yuqing Huang, Yu Yang 0002, Zhong Chen 0005 |
Signal Process. | 3 |
| 2024 | RTracker: Recoverable Tracking via PN Tree Structured MemoryabstractExisting tracking methods mainly focus on learning better target representation or developing more robust prediction models to improve tracking performance. While tracking performance has significantly improved, the target loss issue occurs frequently due to tracking failures, complete occlusion, or out-of-view situations. However, con-siderably less attention is paid to the self-recovery issue of tracking methods, which is crucial for practical applications. To this end, we propose a recoverable tracking framework, RTracker, that uses a tree-structured memory to dynamically associate a tracker and a detector to enable self-recovery ability. Specifically, we propose a Positive-Negative Tree-structured memory to chronologically store and maintain positive and negative target samples. Upon the PN tree memory, we develop corresponding walking rules for determining the state of the target and define a set of control flows to unite the tracker and the detector in different tracking scenarios. Our core idea is to use the support samples of positive and negative target categories to establish a relative distance-based criterion for a reliable assessment of target loss. The favorable performance in comparison against the state-of-the-art methods on nu-merous challenging benchmarks demonstrates the effectiveness of the proposed algorithm. All the source code and trained models will be released at https://github.com/NorahGreen/RTracker. Yuqing Huang, Xin Li 0034, Zikun Zhou, Yaowei Wang 0001, Zhenyu He 0001, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2024 | Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingabstractThermal infrared(TIR) data exhibits higher tolerance to extreme environments, making it a valuable complement to RGB data in tracking tasks. RGBT tracking aims to leverage information from RGB and TIR images for stable and robust tracking. However, existing RGBT tracking methods face challenges due to significant modality differences and selective emphasis on interactive information, leading to inefficiencies in the cross-modal interaction. To address these issues, we propose a novel Integrating Interaction into Modality-shared Features with ViT(IIMF) framework, which is a simplified cross-modal interaction network including modality-shared, RGB modality-specific, and TIR modality-specific branches. The Modality-shared branch aggregates modality-shared information and implements inter-modal interaction. Specifically, our approach first extracts modality-shared features from RGB and TIR features with a cross-attention mechanism. Furthermore, we design a Cross-Attention-based Modality-shared Information Aggregation(CAMIA) module to further aggregate modality-shared information with modality-shared tokens. We evaluate our model on three widely-used benchmark datasets and extensive experiments demonstrate that our method achieves state-of-the-art performance. All the source code are released at https://github.com/Liqiu-Chen/IIMF. Liqiu Chen, Yuqing Huang, Zikun Zhou, Zhenyu He 0001 |
ACM Multimedia | 2 |
| 2024 | An Internet of Things Management System for Roadside Parking Space Based on Solar Power Supply and RF Energy TransmissionabstractCurrently, the automatic management of roadside parking space remains a significant challenge. Installing cameras next to each roadside parking space is an extremely expensive solution. Utilizing Bluetooth beacons to identify vehicles requires the installation of a Bluetooth device on each vehicle. This not only necessitates vehicle power supply but also poses installation difficulties. Additionally, the Bluetooth devices installed on parking space need to operate for extended periods, consuming a significant amount of electrical energy and requiring grid power supply. To reduce system costs, alleviate installation difficulties, and achieve complete autonomy in power supply, this article proposes a solar-powered and radio frequency (RF) energy transmission-based Internet of Things (IoT) management system for roadside parking space. We have designed both the mobile and fixed terminal of the system. The fixed terminal is powered by solar panels, enabling it to emit RF energy, retrieve vehicle information, automatically track time, and upload data. The mobile terminal is designed as a passive device capable of receiving RF energy and transmitting vehicle information. Through experimental testing, the system successfully achieves automatic retrieval and upload of vehicle information and parking duration, thereby validating the feasibility of the proposed system. Since the mobile terminal is a passive device and the fixed terminal is powered by solar panels, no external power supply is required. This project introduces for the first time the utilization of solar energy conversion into RF power supply, enabling the mobile terminal to function as a passive device without the need for external power or internal batteries. Ge Shi 0001, Zhebin Shi, Xiudeng Wang, Yinshui Xia, Shengyao Jia, Mang Shi, Yuqing Huang |
IEEE Internet Things J. | 8 |
| 2024 | Cross-Modality Proposal-Guided Feature Mining for Unregistered RGB-Thermal Pedestrian DetectionabstractRGB-Thermal (RGB-T) pedestrian detection aims to locate pedestrians in RGB-T image pairs to exploit the complementation between the two modalities for improving detection robustness in extreme conditions. Most existing algorithms assume that the RGB-T image pairs are well registered, while in the real world, they are not ideally aligned due to parallax or different field-of-view of the cameras. The pedestrians in misaligned image pairs may be located at different positions in two images, which results in two challenges: 1) how to achieve inter-modality complementation using spatially misaligned RGB-T pedestrian patches and 2) how to recognize unpaired pedestrians at the boundary. To address these issues, we propose a new paradigm for unregistered RGB-T pedestrian detection, which predicts two separate pedestrian locations in RGB and thermal images. Specifically, we propose a cross-modality proposal-guided feature mining (CPFM) mechanism to extract two precise fusion features for representing a pedestrian in the two modalities, even if the given RGB-T image pair is unaligned. It enables us to effectively exploit the complementation between the two modalities. With the CPFM mechanism, we build a two-stream dense detector that predicts two pedestrian locations in the two modalities based on the corresponding fusion features mined by the CPFM mechanism. In addition, we design a data augmentation method, named Homography, to simulate the discrepancy in scales and views between images. We also investigate two non-maximum suppression (NMS) methods for post-processing purposes. Favorable experimental results demonstrate the effectiveness and robustness of our method in addressing unregistered pedestrians with different shifts. Zikun Zhou, Yuqing Huang, Gaojun Li, Zhenyu He 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | CiteTracker: Correlating Image and Text for Visual TrackingabstractExisting visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method. The source code and trained models will be released at https://github.com/NorahGreen/CiteTracker. Xin Li 0034, Yuqing Huang, Zhenyu He 0001, Yaowei Wang 0001, Huchuan Lu, Ming-Hsuan Yang 0001 |
ICCV | 2 |
| 2023 | Hypercomplex Low Rank Reconstruction for NMR Spectroscopy
Jiaying Zhan, Zhangren Tu, Yirong Zhou, Jianfan Wu, Qing Hong, Yuqing Huang, Vladislav Orekhov, Xiaobo Qu 0001, Di Guo 0003 |
Signal Process. | 7 |
| 2022 | Accurate Summary-based Cardinality Estimation Through the Lens of Cardinality Estimation GraphsabstractThis paper is an experimental and analytical study of two classes of summary-based cardinality estimators that use statistics about input relations and small-size joins in the context of graph database management systems: (i) optimistic estimators that make uniformity and conditional independence assumptions; and (ii) the recent pessimistic estimators that use information theoretic linear programs (LPs). We begin by analyzing how optimistic estimators use pre-computed statistics to generate cardinality estimates. We show these estimators can be modeled as picking bottom-to-top paths in a cardinality estimation graph (CEG), which contains sub-queries as nodes and edges whose weights are average degree statistics. We show that existing optimistic estimators have either undefined or fixed choices for picking CEG paths as their estimates and ignore alternative choices. Instead, we outline a space of optimistic estimators to make an estimate on CEGs, which subsumes existing estimators. We show, using an extensive empirical analysis, that effective paths depend on the structure of the queries. While on acyclic queries and queries with small-size cycles, using the maximum-weight path is effective to address the well known underestimation problem, on queries with larger cycles these estimates tend to overestimate, which can be addressed by using minimum weight paths. We next show that optimistic estimators and seemingly disparate LP-based pessimistic estimators are in fact connected. Specifically, we show that CEGs can also model some recent pessimistic estimators. This connection allows us to adopt an optimization from pessimistic estimators to optimistic ones, and provide insights into the pessimistic estimators, such as showing that they have combinatorial solutions. Jeremy Chen, Yuqing Huang, Mushi Wang, Semih Salihoglu, Kenneth Salem |
Proc. VLDB Endow. | 2 |
| 2022 | A Novel Rapid-Flooding Approach With Real-Time Delay Compensation for Wireless-Sensor Network Time SynchronizationabstractOne-way-broadcast-based flooding time synchronization algorithms are commonly used in wireless-sensor networks (WSNs). However, the packet delay and clock drift pose a challenge to accuracy, as they entail serious by-hop error accumulation problems in the WSNs. To overcome this, a rapid-flooding multibroadcast time synchronization with real-time delay compensation (RDC-RMTS) is proposed in this article. By using a rapid-flooding protocol, flooding latency of the referenced time information is significantly reduced in the RDC-RMTS. In addition, a new joint clock skew-offset maximum-likelihood estimation (MLE) is developed to obtain the accurate clock parameter estimations and the real-time packet delay estimation. Moreover, an innovative implementation of the RDC-RMTS is designed with an adaptive clock offset estimation. The experimental results indicate that the RDC-RMTS can easily reduce the variable delay and significantly slow the growth of by-hop error accumulation. Thus, the proposed RDC-RMTS can achieve accurate time synchronization in large-scale complex WSNs. Fanrong Shi, Simon X. Yang, Xianguo Tuo, Lili Ran, Yuqing Huang |
IEEE Trans. Cybern. | 5 |
| 2021 | Diversity and consistency embedding learning for multi-view subspace clustering
Yong Mi, Zhenwen Ren, Mithun Mukherjee 0001, Yuqing Huang, Quan-Sen Sun, Liwan Chen |
Appl. Intell. | 4 |
| 2021 | Multiple kernel clustering with pure graph learning scheme
Xingfeng Li 0004, Zhenwen Ren, Haoyun Lei, Yuqing Huang, Quan-Sen Sun |
Neurocomputing | 4 |
| 2021 | Robust multi-view graph clustering in latent energy-preserving embedding space
Zhenwen Ren, Xingfeng Li 0004, Mithun Mukherjee 0001, Yuqing Huang, Quan-Sen Sun |
Inf. Sci. | 4 |
| 2020 | Robust energy preserving embedding for multi-view subspace clustering
Haoran Li 0009, Zhenwen Ren, Mithun Mukherjee 0001, Yuqing Huang, Quan-Sen Sun, Xingfeng Li 0004, Liwan Chen |
Knowl. Based Syst. | 4 |
| 2012 | Preliminary Search Engine for Open Protein IdentificationabstractProtein identification is the most important and basic problem for proteomics. Using tandem mass spectrometry and database search is one of the most widely used identification techniques. However, the improved sensitivity of mass spectrometers, rapid expansion of databases and more complex analysis, like post-translational modification and non-specific enzymatic digestion, have challenged current restricted protein identification search engines in scale and speed severely. In this paper, we proposed an open protein identification method relaxing enzyme, and presented our distributed design to support big protein database with non-specific digestion analysis based on pFind, a practical tandem mass spectra search engine developed in China. With classical bigger protein databases ipi. HUMAN and uniprot-sprot we got nearly linear speedup in a 20-blade cluster. By further analysis, we can expect real time identification to some extent. Hao Chi, Yuanzheng Lu, Yuqing Huang, Simin He 0001 |
PDCAT | 4 |