EDBT 2026 Demo / reviewers in the wild / expert
Haodong Jing
dblp:322/3946
· DBLP profile ↗
23ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-6643-7588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural RepresentationabstractVisual neural decoding is an important research topic at the intersection of cognitive neuroscience and machine learning. While recent progress has been made in EEG-based neural decoding, reconstructing dynamic visual content remains challenging. In the field of EEG decoding, current models either utilize pre-trained encoders for feature extraction or employ graph neural networks to represent the spatio-temporal information embedding, resulting in poor model representation and high complexity. We propose EVOKE -- an innovative framework for zero-shot decoding of high-fidelity videos from EEG signals. EVOKE employs Implicit Neural Representations to perform complete spatial modeling of EEG and continuously decouples information in the EEG-INR perceptual space. Additionally, we construct a Hierarchical-aware Attention Module (HAM) to decode EEG from three feature anchors: visual, semantic, motion, and progressively control task inference. The Motion Attention Flow (MAF) we developed overcomes the limitations of capturing motion features in dynamic stimuli, creating a more robust representation that enhances reconstruction consistency. Comprehensive experiments prove that SOTA performance of EVOKE (0.353 SSIM, 0.715 CLIP-pcc). We provide an effective method for converting brain activity into rich visual experiences and set a new benchmark for brain multimodal generation. Haodong Jing, Panqi Yang, Dongyao Jiang, Nanning Zheng 0001 |
AAAI | 1 |
| 2026 | UniHOI: Unified Human-Object Interaction Understanding via Unified Token SpaceabstractIn the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose UniHOI, which jointly models HOI detection and generation via a unified token space, thereby effectively promoting knowledge sharing and enhancing generalization. Specifically, we introduce a symmetric interaction-aware attention module and a unified semi-supervised learning paradigm, enabling effective bidirectional mapping between images and interaction semantics even under limited annotations. Extensive experiments demonstrate that UniHOI achieves state-of-the-art performance in both HOI detection and generation. Specifically, UniHOI improves accuracy by 4.9% on long-tailed HOI detection and boosts interaction metrics by 42.0% on open-vocabulary generation tasks. Panqi Yang, Haodong Jing, Nanning Zheng 0001 |
AAAI | 2 |
| 2026 | We Need a More Robust Classifier: Dual Causal Learning Empowers Domain-Incremental Time Series ClassificationabstractThe World Wide Web thrives on intelligent services that rely on accurate time series classification, which has recently witnessed significant progress driven by advances in deep learning. However, existing studies face challenges in domain incremental learning. In this paper, we propose a lightweight and robust dual-causal disentanglement framework (DualCD) to enhance the robustness of models under domain incremental scenarios, which can be seamlessly integrated into time series classification models. Specifically, DualCD first introduces a temporal feature disentanglement module to capture class-causal features and spurious features. The causal features can offer sufficient predictive power to support the classifier in domain incremental learning settings. To accurately capture these causal features, we further design a dual-causal intervention mechanism to eliminate the influence of both intra-class and inter-class confounding features. This mechanism constructs variant samples by combining the current class's causal features with intra-class spurious features and with causal features from other classes. The causal intervention loss encourages the model to accurately predict the labels of these variant samples based solely on the causal features. Extensive experiments on multiple datasets and models demonstrate that DualCD effectively improves performance in domain incremental scenarios. We summarize our rich experiments into a comprehensive benchmark to facilitate research in domain incremental time series classification. Peibo Duan, Haodong Jing, Mingyang Geng, Jialu Xu, Bin Zhang 0001, Binwu Wang |
WWW | 4 |
| 2026 | InstrucRobo: Object-centric multi-instruction decoupling model for explainable robotic manipulation
Panqi Yang, Haodong Jing, Nanning Zheng 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | UniBVR: Balancing visual and reasoning abilities in unified 3D scene understanding
Panqi Yang, Haodong Jing, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2026 | Mind-VAD: Brain-Inspired fMRI-to-Video Precise Reconstruction via Cross-Modal Autoregressive DiffusionabstractDecoding visual information from brain activity is important and challenging. Existing studies have successfully reconstructed static images from fMRI signals, but fMRI-based dynamic visual reconstruction still has limitations: the low temporal resolution of fMRI limits the coherence of the video, the lack of fine-grained alignment of brain with visual features at different scales, and the deviation of diffusion modeling from the logic of visual encoding. Research has shown that the visual cortex cognitive system streams perceptual and semantic information to different brain regions for processing. Inspired by this, we propose Mind-VAD, a novel dynamic visual reconstruction paradigm that redefines the brain decoding process as cross-modal generation guided by visual stream. Specifically, we design a time-sensitive encoder to accurately achieve temporal localization, and learn different dynamic visual representations of fMRI through an autoencoder with different scales. Then proceed with full-sequence, faster video autoregressive diffusion generation through a coarse-to-fine process. We include the latest AIGC Benchmark in evaluation. Extensive experiments show that Mind-VAD realized accurate and smooth reconstruction, achieves SOTA performance (87.3% classification, 27.5% semantic improvement) on several downstream tasks and more demanding metrics. Overall, Mind-VAD provides a neural interpretable and powerful framework for visual decoding and its potential applications. Haodong Jing, Wenjie Gao 0001, Dongyao Jiang, Shuai Huang 0002, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | DAMind: Zero-Shot Visual Cross-Domain Alignment and Representation for EEG DecodingabstractTo efficiently assist humans in various tasks, it is crucial to accurately decode and understand the rich information embedded in brain's visual cognition. Existing brain-driven research often fails to overcome the challenge of small target data domains, and the lack of explicit semantic, spatial, and other information constraints on feature extractors prevents brain decoding models from learning uniform cross-domain representations, leading to degradation of their performance in unseen domains. To overcome these limitations, we propose DAMind, a multimodal EEG-based model for robust visual cross-domain alignment and decoding. Our approach integrates VLM with brain-inspired cognitive mechanisms, leveraging the strong image-text representation abilities to learn both fine-grained primary visual features and high-level semantic concepts from neural signals, provide effective visual fine-tuning using the visual guidance mechanism. DAMind introduces a stepwise EEG encoding process aligned with visual processing, and employs an instruction-based learning strategy for effective cross-domain zero-shot transfer. Its robust architecture efficiently achieves good generalization performance, enabling the mapping of EEG signals from multiple domains to a unified learning domain. We construct a comprehensive EEG decoding benchmark EBench, DAMind achieves state-of-the-art results on several visual tasks, and outperforms the baseline in zero-shot setting. Haodong Jing, Panqi Yang, Shuai Huang 0002, Badong Chen, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRIabstractDeciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing shallow subject-specific adapters to map cross-subject fMRI data into unified representations. A shared deep decoding model then decodes these features into the target feature space. We use both visual and textual supervision for multi-modal brain decoding and integrate high-level perception decoding with pixel-wise reconstruction guided by high-level perceptions. Our extensive experiments reveal several interesting insights: 1) Training with cross-subject fMRI benefits both high-level and low-level decoding models; 2) Merging high-level and low-level information improves reconstruction performance at both levels; 3) Transfer learning is effective for new subjects with limited training data by training new adapters; 4) Decoders trained on visually-elicited brain activity can generalize to decode imagery-induced activity, though with reduced performance. Guibo Zhu, Haodong Jing, Nanning Zheng 0001 |
AAAI | 4 |
| 2025 | Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot LearningabstractHuman reasoning naturally combines concepts to identify unseen compositions, a capability that Compositional Zero-Shot Learning (CZSL) aims to replicate in machine learning models. However, we observe that focusing solely on typical image classification tasks in CZSL may limit models’ compositional generalization potential. To address this, we introduce C-EgoExo, a video-based benchmark, along with a compositional action recognition task to enable more comprehensive evaluations. Inspired by human reasoning processes, we propose a Dual-branch Hybrid Discrimination (DHD) framework, featuring two branches that decode visual inputs in distinct observation sequences. Through a cross-attention mechanism and a contextual dependency encoder, DHD effectively mitigates challenges posed by conditional variance. We further design a Copula-based orthogonal decoding loss to counteract contextual interference in primitive decoding. Our approach demonstrates outstanding performance across diverse CZSL tasks, excelling in both image-based and video-based modalities and in attribute-object and action-object compositions, setting a new benchmark for CZSL evaluation. Dongyao Jiang, Haodong Jing, Nanning Zheng 0001 |
CVPR | 2 |
| 2025 | Refiner: Fine-grained Cross-modal Concepts Refinement for Compositional Zero-Shot LearningabstractRecent Compositional Zero-Shot Learning (CZSL) methods increasingly adopt the pre-trained vision-language models to capture the contextual relations between image and text spaces. However, the single-class-token design from Transformer-based encoder inevitably captures contextual information from unrelated objects and background, thus hindering the modeling of fine-grained class-specific visual features. Suffering from cross-modal gap, prior methods also struggle to improve compositional recognition performance. To address these issues, we propose a fine-grained cross-modal concepts refinement framework, termed as Refiner, which comprises two pivotal components: (i) the fine-grained concepts refinement of image embeddings to capture state-object context within visual scenes, and (ii) the cross-modal information fusion to mitigate the modality gap. By leveraging learnable query vectors to capture region-specific semantic information pertinent to composition labels, our approach refines visual representations with fine-grained state-object context information. As for cross-modal information fusion, we construct a robust image-to-text mapping by aligning visual embeddings with states, objects, and compositions, respectively. Extensive experiments demonstrate that our Refiner achieves new state-of-the-art performance across all popular benchmarks in both closed- and open-world settings. Haodong Jing, Hui Chen 0036, Nanning Zheng 0001 |
ICASSP | 2 |
| 2025 | Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRI
Haodong Jing, Dongyao Jiang, Haibo Hua, Nanning Zheng 0001 |
ICCV | 1 |
| 2025 | PlaneHEC: Efficient Hand-Eye Calibration for Multi-View Robotic Arm via Any Point Cloud Plane DetectionabstractHand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic systems, usually rely on accurate geometric models or manual assistance, generalize poorly, and can be very complicated and inefficient. Therefore, in this study, we propose PlaneHEC, a generalized hand-eye calibration method that does not require complex models and can be accomplished using only depth cameras, which achieves the optimal and fastest calibration results using arbitrary planar surfaces like walls and tables. PlaneHEC introduces hand-eye calibration equations based on planar constraints, which makes it strongly interpretable and generalizable. PlaneHEC also uses a comprehensive solution that starts with a closed-form solution and improves it with iterative optimization, which greatly improves accuracy. We comprehensively evaluated the performance of PlaneHEC in both simulated and real-world environments and compared the results with other point-cloud-based calibration methods, proving its superiority. Our approach achieves universal and fast calibration with an innovative design of computational models, providing a strong contribution to the development of multi-agent systems and embodied intelligence. Haodong Jing, Yang Liao, Nanning Zheng 0001 |
ICRA | 2 |
| 2025 | SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and LightingabstractHigh-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies (domain gaps) of input images. To address this issue, we propose SAMap, a novel map learning framework that enhances domain generalization capabilities of existing models by reducing domain discrepancies in input images. SAMap innovatively introduces a Semantic Aligner, an image-to-image transformation module that aligns images from different domains into a unified domain space while preserving semantic consistency. To train this aligner, we leverage Vision-Language Models (VLMs) that have acquired image-text alignment capabilities. Specifically, we first train a Prompt Learner that combines handcrafted and learnable prompts to capture domain-invariant semantic information. We then train Semantic Aligner through dual supervision mechanisms: a content preservation loss that maintains feature consistency across transformations and a semantic alignment loss that leverages VLM’s encoders to align transformed images with domain-invariant textual representations. Adequate experiments on the NuScenes dataset demonstrate that when integrated with three existing HD map prediction methods, SAMap achieves a performance improvement of up to 11.6% on unseen domains (rain or night conditions), effectively validating its generalization capabilities across domains. Wenjie Gao 0001, Haodong Jing, Jiawei Fu 0001, Shi-tao Chen, Nanning Zheng 0001 |
IROS | 2 |
| 2025 | MINDEV: Multi-modal Integrated Diffusion Framework for Video Reconstruction from EEG SignalsabstractDespite recent progress in decoding static images from brain activity, reconstructing dynamic visual experiences from EEG signals remains challenging due to the complex temporal dynamics involved. Current approaches primarily rely on pre-trained video generation models while failing to fully leverage the rich temporal-spatial information embedded in EEG signals for video synthesis. This paper proposes MINDEV Multi-modal Integrated Neural DEcoding and Visualization), a framework that places EEG signal processing at the core of video reconstruction. We introduce three key technical contributions: (1) a dual-branch feature extractor that captures both temporal dynamics and spatial relationships in EEG signals, (2) an EEG-driven semantic bridge that uses neural patterns to guide language model interpretation, and (3) a multi-modal video synthesis pipeline where EEG features lead the generation process while semantic guidance provides refinement. Our framework prioritizes the millisecond-level temporal resolution of EEG signals, using them to drive both visual content generation and semantic understanding. Evaluated on the SEED-DV dataset, MINDEV demonstrates superior performance with a semantic classification accuracy of 93.2% and a structural similarity index (SSIM) of 0.4777, establishing a new state-of-the-art for EEG-based video reconstruction. Our code is publicly available at https://github.com/HHarr1son/MINDEV. Shuai Huang 0002, Yongxiong Wang, Haodong Jing, Chendong Qin, Jingqun Tang |
ACM Multimedia | 4 |
| 2025 | NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG SignalsabstractTranslating brain activity into meaningful visual content has long been recognized as a fundamental challenge in neuroscience and brain-computer interface research. Recent advances in EEG-based neural decoding have shown promise, yet two critical limitations remain in this area: poor generalization across subjects and constraints to specific visual tasks. We introduce NEED, the first unified framework achieving zero-shot cross-subject and cross-task generalization for EEG-based visual reconstruction. Our approach addresses three fundamental challenges: (1) cross-subject variability through an Individual Adaptation Module pretrained on multiple EEG datasets to normalize subject-specific patterns, (2) limited spatial resolution and complex temporal dynamics via a dual-pathway architecture capturing both low-level visual dynamics and high-level semantics, and (3) task specificity constraints through a unified inference mechanism adaptable to different visual domains. For video reconstruction, NEED achieves better performance than existing methods. Importantly, Our model maintains 93.7% of within-subject classification performance and 92.4% of visual reconstruction quality when generalizing to unseen subjects, while achieving an SSIM of 0.352 when transferring directly to static image reconstruction without fine-tuning, demonstrating how neural decoding can move beyond subject and task boundaries toward truly generalizable brain-computer interfaces. Shuai Huang 0002, Haodong Jing, Qixian Zhang, Litao Chang, Yating Feng, Chendong Qin, Shuwen Jia, Siyi Sun, Yongxiong Wang |
NeurIPS | 3 |
| 2025 | Exploring inter- and intra-modal relations in compositional zero-shot learning
Hui Chen 0036, Haodong Jing, Nanning Zheng 0001 |
Neurocomputing | 3 |
| 2025 | Pinpointing visual content: Disentangled features in multimodal model for EEG representation learning and decoding
Haodong Jing, Panqi Yang, Haibo Hua, Nanning Zheng 0001 |
Knowl. Based Syst. | 1 |
| 2025 | An Optimal Obstacle Avoidance Method Using Reinforcement Learning-Based Decision Parameterization for Autonomous VehiclesabstractDue to the complexity and uncertainty of traffic environment and the fragmentation of traditional planning and control strategy, conventional planning and control algorithms have the problems of low and unstable efficiency. In order to improve both planning and control efficiency while ensuring the validity of the planned trajectories for autonomous vehicles, a reinforcement learning and optimal control integrated method (RL-OCIM) for planning and control is proposed. The proposed method ensures the obstacle avoidance through incorporating lane change based on proximal policy optimization (PPO) with the optimal control method (OCM). The trained PPO agent takes the structured information of the ego vehicle and the surrounding obstacles as input, and generates the trajectory terminal state for the ego vehicle. Subsequently, the continuous trajectory solving problem is transformed into an optimal control problem (OCP), resulting in the generation of a feasible trajectory that includes vehicle steering and longitudinal acceleration based on the vehicle kinematics model. In addition, to improve the efficiency of agent training, a pre-trained neural network is employed to generate trajectory feature points by taking the trajectory terminal states as input. During the agent training process, the trajectory feature points are obtained directly from the neural network after selecting an action. Experiments are conducted on a hybrid Lincoln MKZ vehicle to demonstrate and validate the effectiveness and efficiency of the proposed RL-OCIM, which includes various driving scenarios such as static obstacle avoidance and dynamic obstacle avoidance. Zhaobo Qin, Haodong Jing, Guodong Xiong, Rongjun Ding |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Hidden States in LLMs Improve EEG Representation Learning and Visual DecodingabstractAnalyzing brain signals and reconstructing visual stimuli from the brain can facilitate further exploration on cognitive functions of the human brain, which have attracted strong interest in neuroscience and artificial intelligence. However, due to defects such as complex noises and the lack of alignment accuracy, efficient methods for extracting information from Electroencephalogram (EEG) signals are still very limited, making it difficult to perform EEG visual decoding tasks. Our study shows a way to handle the issues by proposing a new method for EEG representation learning and visual decoding, thus completing end-to-end image reconstruction tasks from EEG signals. We utilize the ability of semantic extraction and prediction of large language models (LLMs) to enhance the performance of EEG feature extraction. For semantic representation learning, we align EEG signals with target semantic embeddings, which are obtained from hidden states of Large Language Model Meta AI 2 (LLaMa-2) by inputting descriptions of images into the model. We also extract visual features from EEG signals to improve the quantity of the reconstructed images at low levels. Then we fuse semantic features and visual features by applying a pre-trained diffusion model and finally generate the corresponding images. We are the first to incorporate the LLM into EEG visual decoding tasks. Our method achieves the state-of-the-art result of EEG classification accuracy and the quality of reconstructed images on ImageNet-EEG datasets. In one word, our work is an important step forward in the field of exploiting the relationship between language models and human visual cognition. Our codes are available at https://github.com/lay-atsa/llm4eeg. Aoyang Liu, Haodong Jing, Nanning Zheng 0001 |
ECAI | 2 |
| 2024 | MRSP: Learn Multi-representations of Single Primitive for Compositional Zero-Shot Learning
Dongyao Jiang, Hui Chen 0036, Haodong Jing, Nanning Zheng 0001 |
ECCV (65) | 3 |
| 2024 | Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map ConstructionabstractHigh-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion by nearby vehicles, the performance of these methods significantly declines in complex scenarios and long-range detection tasks. In this paper, we explore a new perspective that boosts HD map construction through the use of satellite maps to complement onboard sensors. We initially generate the satellite map tiles for each sample in nuScenes and release a complementary dataset for further research. To enable better integration of satellite maps with existing methods, we propose a hierarchical fusion module, which includes feature-level fusion and BEV-level fusion. The feature-level fusion, composed of a mask generator and a masked cross-attention mechanism, is used to refine the features from onboard sensors. The BEV-level fusion mitigates the coordinate differences between features obtained from onboard sensors and satellite maps through an alignment module. The experimental results on the augmented nuScenes showcase the seamless integration of our module into three existing HD map construction methods. The satellite maps and our proposed module notably enhance their performance in both HD map semantic segmentation and instance detection tasks. Our code will be available at https://github.com/xjtu-csgao/SatforHDMap. Wenjie Gao 0001, Jiawei Fu 0001, Yanqing Shen, Haodong Jing, Shi-tao Chen, Nanning Zheng 0001 |
ICRA | 4 |
| 2024 | Hierarchical Bayesian Causality Network to Extract High-Level Semantic Information in Visual CortexabstractFunctional MRI (fMRI) is a brain signal with high spatial resolution, and visual cognitive processes and semantic information in the brain can be represented and obtained through fMRI. In this paper, we design single-graphic and matched/unmatched double-graphic visual stimulus experiments and collect 12 subjects' fMRI data to explore the brain's visual perception processes. In the double-graphic stimulus experiment, we focus on the high-level semantic information as "matching", and remove tail-to-tail conjunction by designing a model to screen the matching-related voxels. Then, we perform Bayesian causal learning between fMRI voxels based on the transfer entropy, establish a hierarchical Bayesian causal network (HBcausalNet) of the visual cortex, and use the model for visual stimulus image reconstruction. HBcausalNet achieves an average accuracy of 70.57% and 53.70% in single- and double-graphic stimulus image reconstruction tasks, respectively, higher than HcorrNet and HcasaulNet. The results show that the matching-related voxel screening and causality analysis method in this paper can extract the "matching" information in fMRI, obtain a direct causal relationship between matching information and fMRI, and explore the causal inference process in the brain. It suggests that our model can effectively extract high-level semantic information in brain signals and model effective connections and visual perception processes in the visual cortex of the brain. Ming Du 0001, Haodong Jing, Nanning Zheng 0001 |
Int. J. Neural Syst. | 4 |
| 2022 | Images Structure Reconstruction from fMRI by Unsupervised Learning Based on VAE
Haodong Jing, Jianji Wang 0001, Weihua Wu |
ICANN (3) | 2 |