EDBT 2026 Demo / reviewers in the wild / expert
Huanyu He
dblp:250/1199
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-4110-1919ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on brief video segments containing isolated events and basic causal relations, lacking comprehensive and structured causality analysis for videos with multiple interconnected events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relations between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD identifies the causal associations between these events to derive a comprehensive and structured event-level video causal graph explaining why and how the result event occurred. To address the challenges of MECD, we devise a novel framework inspired by the Granger Causality method, incorporating an efficient mask-based event prediction model to perform an Event Granger Test. It estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to mitigate challenges in MECD like causality confounding and illusory causality. Additionally, context chain reasoning is introduced to conduct more robust and generalized reasoning. Experiments validate the effectiveness of our framework in reasoning complete causal relations, outperforming GPT-4o and VideoChat2 by 5.77% and 2.70%, respectively. Further experiments demonstrate that causal relation graphs can also contribute to downstream video understanding tasks such as video question answering and video event prediction. Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Yihang Chen 0002, Tianyao He, Chaofan Gan, Huanyu He, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | DDSPR: Dynamic Domain Selection and Pseudo-label Refinement for Cross-Subject EEG-based Emotion Recognition
Qinyu Hai, Liying Yang 0001, Yumeng Ye, Jingtao Du, Huanyu He |
CogSci | 6 |
| 2025 | AC-CDCN: A Cross-Subject EEG Emotion Recognition Model with Anti-Collapse Domain Generalization
Yubin Sun, Liying Yang 0001, Huanyu He, Jingtao Du |
CogSci | 3 |
| 2025 | JMS2A: Joint Multi-source Domain and Two-step Alignment Strategy for Cross-subject EEG Emotion Recognition
Liying Yang 0001, Jingtao Du, Huanyu He |
CogSci | 4 |
| 2025 | Variational Box Representation with Normalizing Flows for Pedestrian Detection
Huanyu He |
PRCV (17) | 1 |
| 2025 | Toward Accurate and Robust Pedestrian Detection via Variational Inference
Huanyu He, Weiyao Lin, Tianyao He, Yuxi Li 0009 |
Int. J. Comput. Vis. | 1 |
| 2024 | Cross-Subject Emotion Classification with Residual Pseudo-Label Distance-aware Dual-ClassifierabstractEmotion plays a crucial role in information exchange and decision-making processes. Emotion recognition based on electroencephalography (EEG) captures and analyzes brain electrical activity, effectively reflecting emotional characteristics and providing unique advantages for constructing intelligent emotional systems. However, due to significant differences in feature distribution between the source and target domains, traditional models struggle with generalization in cross-subject tasks. To address this issue, this paper proposes a new model for EEG-based emotion analysis that employs the Pseudo-Label Distance-aware Dual-Classifier (PL-DDC) strategy, which relies on a Residual Temporal-Frequency Analysis (RTFA) structure. We name this model Residual Pseudo-Label Distance-aware Dual-Classifier (RPL-DDC). The RTFA module effectively captures mixed dependencies in complex time-frequency signals, while the PL-DDC strategy introduces a distance loss between the source and target domains. By combining classifier output discrepancies and a pseudo-label mechanism, it progressively reduces the feature distribution gap between the two domains, thereby enhancing the model’s classification performance. Experimental results on the SEED and SEED-IV datasets show that the proposed model achieves accuracy of 89.03% ± 5.03 on the SEED dataset and 75.49% ± 10.44 on the SEED-IV dataset. Compared to traditional baseline models, these results validate the effectiveness of the proposed model in cross-subject emotion recognition. Jingtao Du, Liying Yang 0001, Huanyu He, Jiulin Fu |
BIBM | 3 |
| 2024 | MDAC: EEG Emotion Recognition with Multi-Scale Dual Attention Capsule NetworkabstractIn recent years, deep learning has exhibited significant prowess in the field of EEG-based affective recognition. The attention mechanism has always been a focal point of interest within the domain of deep learning. However, existing EEG analysis techniques still face challenges in accurately pinpointing emotion-related signals across both spatial and temporal scales. We propose a novel model, named MDAC, which is based on a multi-scale dual attention and a capsule network tuned for EEG data, to address the aforementioned issue. Initially, the MDA module applies varying scales of perceptual fields to the data and integrates them to simultaneously obtain attention weights at the pixel level and the sampling rate level, achieving precise weighting of EEG signals in both spatial and temporal resolutions. Furthermore, we have increased the number of convolutional channels and the dimensionality of the primary capsules in CapsNet to better align with the characteristics of EEG data, proposing a structural configuration more apt for EEG-based affective recognition. We conducted subject-dependent and subject-independent experiments on the DEAP dataset to validate our model. In the subject-dependent experiments, the accuracy rates for both the valence and arousal dimensions were 99.58%. In the subject-independent experiments, the accuracy rates for the valence and arousal dimensions were 98.15% and 98.04%, respectively. The experimental results corroborate the efficacy of the method we proposed in this paper for the task of emotion recognition. Huanyu He, Liying Yang 0001, Jingtao Du, Qinyu Hai, Jiulin Fu |
BIBM | 1 |
| 2024 | Visibility-guided Human Body Reconstruction from Uncalibrated Multi-view CamerasabstractWe present a novel method for 3D human body reconstruction with multi-view images from calibration-free cameras by multi-view fusion with explicit visibility modelling. Existing multi-view methods usually establish geometric constraints by using accurate camera intrinsic and extrinsic parameters. Despite remarkable performances, multi-view camera calibration often requires complex operations and additional maintenance to fix camera positions and angles, which restrict its applicability to real-world scenarios. In contrast, we leverage vertex-wise visibility prediction as calibration cues to guide the multi-view human body aggregation, which eliminates the need for camera calibration. Specifically, we estimate the UV position map and the vertex-wise visibility map of human body in each camera view, which allows us to align and aggregate multi-view information in a hierarchical manner. To further improve the alignment between human body and vertex-wise visual features, we propose an Occlusion-aware UV-pixel Refinement (OUVR) module, which takes the previous result of coarse alignment as input. The visible vertices are disentangled from the UV map and are reprojected on the image to describe the misalignment of current body estimation and image features. The UV map representation is adopted throughout the refinement process to avoid the potential error propagation brought by parametric representation. The effectiveness of our approach is validated on 3D human body reconstruction, as it surpasses current leading multi-view fusion methods, and showing comparable performance to methods that require accurate multi-view camera calibration. Zhenyu Xie, Huanyu He, Gui Zou, Weiyao Lin |
ICMR | 2 |
| 2023 | BiCCT: A Compact Convolutional Transformer for EEG Emotion RecognitionabstractEmotion is a manifestation of human’s internal psychological and physiological reactions. Understanding and recognizing emotions is one of the important ways to understand human behavior and human-computer interaction. However, with the widespread application of deep learning in the field of EEG emotion recognition, the number of parameters and model size have increased accordingly. In this paper, we combined the Bi-hemisphere asymmetry theory and Compact Convolutional Transformer to propose a model named BiCCT to recognize emotions, which has fewer training parameters and can achieve higher recognition performance. We first constructed three different matrices of the recorded EEG information according to the international 10-20 system to preserve the temporal information and spatial information of the EEG signals. Next, we applied an improved Transformer architecture, which achieves fewer model parameters and a lightweight structure through the token pooling module and the Convolutional Tokenizer module. We conducted a set of subject-dependent and a set of subject-dependent shuffle experiments on the DEAP dataset. The first set of experiments used a subject-known movie to predict a completely unknown movie. In the second set of experiments, all the data of the subject were randomly divided into training set and test set. We obtained 67.42% for valence and 67.81% for arousal in the first set of experiments. We achieved 94.41% accuracy in the valence dimension and 95.15% accuracy in the arousal dimension in the second set of experiments. At the same time, our model parameters are only 0.17M, which is far lower than other models. It means that our model is lighter and faster in training speed, and has the ability to be deployed in some scenarios with limited computing resources potential. Liying Yang 0001, Chengchuang Tang, Qian Zhang 0074, Huanyu He |
BIBM | 5 |
| 2022 | Trace-Level Invisible Enhanced Network for 6D Pose EstimationabstractEstimating 6D pose of the object from a single image is es-sential for robotic manipulation. Many recent learning-based methods directly regress the pose from 2D-3D points corre-spondence. The problem is that, these methods only make use of visible information from the single-view image, resulting ambiguity for the network to solve pose from the limited cor-responding pairs. To overcome this problem, this paper intro-duces INVNet, integrating invisible information into the visi-ble 2D-3D correspondence to model geometry features of the 3D object. Instead of directly reconstruct the coordinate of in-visible points, we propose Trace-level Geometry Path, which estimates the trace-level depth of the object model for each image pixel. Specifically, our INVNet generates dense visible correspondence as well as Trace-level Geometry Path map, then learn to solve 6D pose from them. Meanwhile, each cam-era ray along with Trace-level Geometry Path is transformed to the object space by the predicted pose to compute invisi-ble correspondence loss from visible one, back to enhance its learning. Extensive experiments show that our approach out-performs state-of-the-art methods on the benchmark LM and LM-O datasets. Hanbo Sang, Zelin Ni, Huanyu He, Xuesong Gao, Qihao Sun, Sihai Zhang, Supavadee Aramvith, Weiyao Lin |
ICME | 3 |
| 2021 | Variational Pedestrian DetectionabstractPedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the current IoU-based ground truth assignment procedure in classical object detection methods. In this paper, we develop a unique perspective of pedestrian detection as a variational inference problem. We formulate a novel and efficient algorithm for pedestrian detection by modeling the dense proposals as a latent variable while proposing a customized Auto-Encoding Variational Bayes (AEVB) algorithm. Through the optimization of our proposed algorithm, a classical detector can be fashioned into a variational pedestrian detector. Experiments conducted on CrowdHuman and CityPersons datasets show that the proposed algorithm serves as an efficient solution to handle the dense pedestrian detection problem for the case of single-stage detectors. Our method can also be flexibly applied to two-stage detectors, achieving notable performance enhancement. Huanyu He, Yuxi Li 0009, John See, Weiyao Lin |
CVPR | 2 |