EDBT 2026 Demo / reviewers in the wild / expert
Honghao Fu
dblp:284/1243
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Theory of computation · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAGabstractScaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows.While retrievalaugmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence into a compact context, most existing methods (i) flatten videos into independent segments, breaking their inherent spatio-temporal structure, and (ii) depend on explicit semantic matching, which can miss cues that are implicitly relevant to the query's intent.To overcome these limitations, we propose VideoStir, a structured and intent-aware long-video RAG framework.It firstly structures a video as a spatiotemporal graph at clip level, and then performs multi-hop retrieval to aggregate evidence across distant yet contextually related events.Furthermore, it introduces an MLLM-backed intentrelevance scorer that retrieves frames based on their alignment with the query's reasoning intent.To support this capability, we curate IR-600K, a large-scale dataset tailored for learning frame-query intent alignment.Experiments show that VideoStir is competitive with stateof-the-art baselines without relying on auxiliary information, highlighting the promise of shifting long-video RAG from flattened semantic matching to structured, intent-aware reasoning.Codes and checkpoints are available at https: //github.com/RomGai/VideoStir. Honghao Fu, Yiwei Wang 0001, Dailing Zhang, Jun Liu 0036, Yujun Cai |
ACL (1) | 1 |
| 2026 | SDR-GAIN: A High Real-Time Occluded Pedestrian Pose Completion Method for Autonomous DrivingabstractWith the advancement of vision-based autonomous driving technology, pedestrian detection have become an important component for improving traffic safety and driving system robustness. Nevertheless, in complex traffic scenarios, conventional pose estimation approaches frequently fail to accurately reconstruct occluded keypoints, primarily due to obstructions caused by vehicles, vegetation, or architectural elements. To address this issue, we propose a novel real-time occluded pedestrian pose completion framework termed Separation and Dimensionality Reduction-based Generative Adversarial Imputation Nets (SDR-GAIN). Unlike previous approaches that train visual models to distinguish occlusion patterns, SDR-GAIN aims to learn human pose directly from the numerical distribution of keypoint coordinates and interpolate missing positions. It employs a self-supervised adversarial learning paradigm to train lightweight generators with residual structures for the imputation of missing pose keypoints. Additionally, it integrates multiple pose standardization techniques to alleviate the difficulty of the learning process. Experiments conducted on the COCO and JAAD datasets demonstrate that SDR-GAIN surpasses conventional machine learning and Transformer-based missing data interpolation algorithms in accurately recovering occluded pedestrian keypoints, while simultaneously achieving microsecond-level real-time inference. Honghao Fu, Yongli Gu, Yidong Yan, Yilang Shen, Libo Sun 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | A Cryptographic Perspective on the Verifiability of Quantum AdvantageabstractIn recent years, achieving verifiable quantum advantage on a NISQ device has emerged as an important open problem in quantum information. The sampling-based quantum advantages are not known to have efficient verification methods. This article investigates the verification of quantum advantage from a cryptographic perspective. We establish a strong connection between the verifiability of quantum advantage and cryptographic and complexity primitives, including efficiently samplable, statistically far but computationally indistinguishable pairs of (mixed) quantum states ( EFI ), pseudorandom states ( PRS ), and variants of minimum circuit size problems ( MCSP ). Specifically, we prove that a) a sampling-based quantum advantage is either verifiable or can be used to build EFI and even PRS and b) polynomial-time algorithms for a variant of MCSP would imply efficient verification of quantum advantages. Our work shows that the quest for verifiable quantum advantages may lead to applications of quantum cryptography, and the construction of quantum primitives can provide new insights into the verifiability of quantum advantages. Nai-Hui Chia, Honghao Fu, Fang Song 0001, Penghui Yao |
ACM Trans. Quantum Comput. | 2 |
| 2025 | VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for MinecraftabstractLarge language models (LLMs) have shown significant promise in embodied decisionmaking tasks within virtual open-world environments.Nonetheless, their performance is hindered by the absence of domain-specific knowledge.Methods that finetune on largescale domain-specific data entail prohibitive development costs.This paper introduces Vista-Wise, a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis.It reduces the requirement for domain-specific training data from millions of samples to a few hundred.VistaWise integrates visual information and textual dependencies into a cross-modal knowledge graph (KG), enabling a comprehensive and accurate understanding of multimodal environments.We also equip the agent with a retrieval-based pooling strategy to extract task-related information from the KG, and a desktop-level skill library to support direct operation of the Minecraft desktop client via mouse and keyboard inputs.Experimental results demonstrate that VistaWise achieves state-of-the-art performance across various open-world tasks, highlighting its effectiveness in reducing development costs while enhancing agent performance. Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai |
EMNLP | 1 |
| 2025 | BrainVis: Exploring the Bridge between Brain and Visual Signals via Image ReconstructionabstractAnalyzing and reconstructing visual stimuli from brain signals effectively advances our understanding of the human visual system. However, EEG signals are complex and contain significant noise, leading to substantial limitations in existing approaches of visual stimuli reconstruction from EEG. These limitations include difficulties in aligning EEG embeddings with fine-grained semantic information and a heavy reliance on additional large-scale datasets for training. To address these challenges, we propose a novel approach called BrainVis. This approach introduces a self-supervised paradigm to learn EEG time-domain features and incorporates frequency-domain features to enhance EEG representations. We also propose a multi-modal alignment method called semantic interpolation to achieve fine-grained semantic reconstruction. Additionally, we employ cascaded diffusion models to reconstruct images. Using only 9.1% of the training data required by previous mask modeling works, our proposed BrainVis outperforms state-of-the-art methods in both semantic fidelity reconstruction and generation quality. The code is available at https://github.com/RomGai/BrainVis. Honghao Fu, Hao Wang 0094, Jing Jih Chin, Zhiqi Shen 0001 |
ICASSP | 1 |
| 2025 | Corruption-Agnostic Sign Language TranslationabstractSign Language Translation (SLT) aims to translate sign languages into spoken languages, acting as a bridge between the hard-of-hearing community and the hearing world. However, existing SLT research predominantly depends on meticulously curated datasets (e.g., PHOENIX-2014T), whereas practical SLT systems may encounter various forms of unanticipated corruptions that can compromise input quality, such as weather effects, camera blurring, external noise, or intrusion of extraneous objects. Consequently, the challenges faced in real-world applications are not fully considered, nor is the robustness required for SLT systems operating in diverse and unpredictable environments adequately evaluated. In order to make existing SLT models applicable in real-world and improve their robustness against various unknown corruptions, we introduce a Corruption-Agnostic Robust Sign Language Translation (CAR-SLT) framework. CAR-SLT consists of two primary components: Decoupled Information Bottleneck(DIB) and Adversarial Perturbation Alignment(APA). DIB decouples SLT-relevant from SLT-irrelevant information within the input, ensuring that the features used for translation contain only necessary elements for SLT. By focusing on SLT-relevant information that is less susceptible to corruptions, DIB enhances the robustness of SLT model. APA employs adversarial strategies to generate gradient-perturbed visual features that simulate scenarios where inputs are corrupted. Subsequently, APA uses two alignment strategies to maintain consistent model performance regardless of input condition, thus improving robustness against unknown corruptions. To evaluate the robustness of SLT systems and simulate the corruptions that a SLT system might encounter in real-world applications, we introduce a challenging setting called conrruption-agnostic sign language translation, along with the PHOENIX-2014T-C dataset. In this setting, the SLT model is trained on clean data but tested on corrupted data, where both the types of corruptions and their locations within the videos are unknown. The PHOENIX-2014T-C incorporates various types of corruptions encountered in real-world scenarios. We evaluate previous state-of-the-art methods as well as our method in this setting. The experimental results demonstrate that previous SLT methods do not perform well under corruption-agnostic setting, while our method exhibits superior performance. Honghao Fu, Yidong Chen 0001 |
IJCNN | 1 |
| 2025 | The Computational Advantage of MIP* Vanishes in the Presence of NoiseabstractThe class MIP* of quantum multiprover interactive proof systems with entanglement is much more powerful than its classical counterpart MIP [ 8 , 31 , 32 ]: while MIP = NEXP, the quantum class MIP * is equal to RE, a class including the halting problem. This is because the provers in MIP * can share unbounded quantum entanglement. However, recent works [ 53 , 54 ] have shown that this advantage is significantly reduced if the provers’ shared state contains noise. This article attempts to exactly characterize the effect of noise on the computational power of quantum multiprover interactive proof systems. We investigate the quantum two-prover one-round interactive system MIP * [poly, O (1)], where the verifier sends polynomially many bits to the provers and the provers send back constantly many bits. We show that noise completely destroys the computational advantage given by shared entanglement in this model. Specifically, we show that if the provers are allowed to share arbitrarily many EPR states, where each EPR state is affected by an arbitrarily small constant amount of noise, the resulting complexity class is equivalent to NEXP = MIP. This improves significantly on the previous best-known bound of NEEEXP (nondeterministic triply exponential time) [ 53 ]. We also show that this collapse in power is due to noise, rather than the O (1) answer size, by showing that allowing for noiseless EPR states gives the class the full power of RE = MIP * [poly, poly]. Along the way, we develop two technical tools of independent interest. First, we give a new, deterministic tester for the positivity of an exponentially large matrix, provided that it has a low-degree Fourier decomposition in terms of Pauli matrices. Secondly, we develop a new invariance principle for smooth matrix functions having bounded third-order Fréchet derivatives or which are Lipschitz continuous. Yangjing Dong, Honghao Fu, Anand Natarajan 0001, Minglong Qin, Haochen Xu, Penghui Yao |
J. ACM | 2 |
| 2025 | HDTCNet: A hybrid-dimensional convolutional network for multivariate time series classification
Yongli Gu, Hanlin Qin, Naveed Akhtar, Shuai Yuan 0013, Honghao Fu, Shuowen Yang, Ajmal Mian |
Pattern Recognit. | 6 |
| 2025 | Data augmentation and debiasing for signers in signer-independent sign language translation
Honghao Fu, Yidong Chen 0001 |
J. Supercomput. | 1 |
| 2024 | The Computational Advantage of MIP^∗ Vanishes in the Presence of NoiseabstractQuantum multiprover interactive proof systems with entanglement MIP* are much more powerful than its classical counterpart MIP (Babai et al. '91, Ji et al. '20): while MIP = NEXP, the quantum class MIP* is equal to RE, a class including the halting problem. This is because the provers in MIP* can share unbounded quantum entanglement. However, recent works of Qin and Yao '21 and '23 have shown that this advantage is significantly reduced if the provers' shared state contains noise. This paper attempts to exactly characterize the effect of noise on the computational power of quantum multiprover interactive proof systems. We investigate the quantum two-prover one-round interactive system MIP*[poly, O(1)], where the verifier sends polynomially many bits to the provers and the provers send back constantly many bits. We show noise completely destroys the computational advantage given by shared entanglement in this model. Specifically, we show that if the provers are allowed to share arbitrarily many noisy EPR states, where each EPR state is affected by an arbitrarily small constant amount of noise, the resulting complexity class is equivalent to NEXP = MIP. This improves significantly on the previous best-known bound of NEEEXP (nondeterministic triply exponential time) by Qin and Yao '21. We also show that this collapse in power is due to the noise, rather than the O(1) answer size, by showing that allowing for noiseless EPR states gives the class the full power of RE = MIP*[poly, poly]. Along the way, we develop two technical tools of independent interest. First, we give a new, deterministic tester for the positivity of an exponentially large matrix, provided it has a low-degree Fourier decomposition in terms of Pauli matrices. Secondly, we develop a new invariance principle for smooth matrix functions having bounded third-order Fréchet derivatives or which are Lipschitz continous. Yangjing Dong, Honghao Fu, Anand Natarajan 0001, Minglong Qin, Haochen Xu, Penghui Yao |
CCC | 2 |
| 2023 | Parallel Self-Testing of EPR Pairs Under Computational AssumptionsabstractSelf-testing is a fundamental feature of quantum mechanics that allows a classical verifier to force untrusted quantum devices to prepare certain states and perform certain measurements on them. The standard approach assumes at least two spatially separated devices. Recently, Metger and Vidick [Quantum, 2021] showed that a single EPR pair of a single quantum device can be self-tested under computational assumptions. In this work, we generalize their results to give the first parallel self-test of $N$ EPR pairs and measurements on them in the single-device setting under the same computational assumptions. We show that our protocol can be passed with probability negligibly close to $1$ by an honest quantum device using poly$(N)$ resources. Moreover, we show that any quantum device that fails our protocol with probability at most $ε$ must be poly$(N,ε)$-close to being honest in the appropriate sense. In particular, our protocol can test any distribution over tensor products of computational or Hadamard basis measurements, making it suitable for applications such as device-independent quantum key distribution under computational assumptions. Moreover, a simplified version of our protocol is the first that can efficiently certify an arbitrary number of qubits of a single cloud quantum computer using only classical communication. Honghao Fu, Daochen Wang |
ICALP | 1 |