EDBT 2026 Demo / reviewers in the wild / expert
Haojian Huang
dblp:355/4169
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-0661-712XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Trustworthy machine learning · 21% Language models and text generation · 19% Vision and language · 14% | |
| Computer graphics and multimedia
2 papers |
Image and video coding · 88% Multimedia analysis and retrieval · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
2.6 | 3 | 2026 | Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026 Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025 CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation › neural network uncertainty
evidential deep learning |
1.9 | 3 | 2026 | Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025 CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024 Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
1.7 | 2 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
1.1 | 2 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025 |
Natural language and speech › Language models and text generation › LLM agents
agent memory |
1.0 | 1 | 2026 | EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer · AAAI 2026 |
Computer vision › Vision and language
moment retrieval |
1.0 | 1 | 2026 | Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
1.0 | 1 | 2026 | EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 |
Computer vision › 3D vision
physical plausibility |
0.9 | 1 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 |
Machine learning › Generative modeling › video generation
physics-informed video generation |
0.9 | 1 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 |
Computer vision › Video understanding and tracking › spatio-temporal understanding
spatio-temporal video understanding |
0.9 | 1 | 2025 | VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025 |
Computer vision › Video understanding and tracking › video analytics
sports video analysis |
0.9 | 1 | 2025 | FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning · ACM Multimedia 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025 |
Computer vision › Vision and language
video-language model |
0.9 | 1 | 2025 | VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025 |
Data mining › predictive modeling
classification |
0.9 | 1 | 2025 | Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025 |
Data mining › predictive modeling › classification › pattern classification
multi-view classification |
0.9 | 1 | 2025 | Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025 |
Image and video coding › image quality assessment
AI-generated image quality assessment |
0.9 | 1 | 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment · ACM Multimedia 2025 |
Image and video coding
image quality assessment |
0.9 | 1 | 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment · ACM Multimedia 2025 |
Computer vision › Face, body and person analysis › facial expression analysis › facial expression recognition
dynamic facial expression recognition |
0.8 | 1 | 2024 | FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024 |
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition |
0.8 | 1 | 2024 | FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
multi-modal zero-shot learning |
0.8 | 1 | 2024 | CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.8 | 1 | 2024 | CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.3 | 1 | 2025 | FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning · ACM Multimedia 2025 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2024 | FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024 |
Multimedia analysis and retrieval › multimodal learning
multimodal representation learning |
0.2 | 1 | 2024 | CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
evidential deep learning · 3.3markov random field · 1.7feature-neighborhood structure · 1.7subjective experience-based memory · 1.0query reconstruction · 1.0maze navigation · 1.0match-2 elimination · 1.0deep evidential regression · 1.0cross-modal fusion · 1.0direct preference optimization · 0.9cross-modal resonance · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalabstractIn the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval. Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He |
AAAI | 1 |
| 2026 | EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVerabstractMost existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce two dynamic spatial benchmarks—locally observable maze navigation and match-2 elimination—that systematically evaluate models' abilities in spatial understanding and adaptive planning when local perception, environment feedback, and global objectives are tightly coupled. Each action triggers structural changes in the environment, requiring continuous update of cognition and strategy. We further propose a subjective experience-based memory mechanism for cross-task experience transfer and validation. Experiments show that our benchmarks reveal key limitations of mainstream models in dynamic spatial reasoning and long-term memory, providing a comprehensive platform for future methodological advances. Pukun Zhao, Miaowei Wang, Fanqing Zhou, Haojian Huang |
AAAI | 6 |
| 2025 | Trusted Unified Feature-Neighborhood Dynamics for Multi-View ClassificationabstractMulti-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views. Haojian Huang, Chuanyu Qin, Zhe Liu 0041, Kaijing Ma, Han Fang 0002, Chao Ban, Hao Sun 0038, Zhongjiang He |
AAAI | 1 |
| 2025 | VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Haojian Huang, Shengqiong Wu, Meng Luo 0010, Jinlan Fu, Xinya Du, Hanwang Zhang, Hao Fei 0001 |
ICML | 1 |
| 2025 | FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
Haojian Huang, Xinxiang Yin, Dian Shao |
ACM Multimedia | 2 |
| 2025 | Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | Hierarchical Fine-grained Preference Optimization for Physically Plausible Video GenerationabstractRecent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose **PhysHPO**, a novel framework for Hierarchical Cross-Modal Direct Preference Optimization, to tackle this challenge by enabling fine-grained preference alignment for physically plausible video generation. PhysHPO optimizes video alignment across four hierarchical granularities: a) ***Instance Level***, aligning the overall video content with the input prompt; b) ***State Level***, ensuring temporal consistency using boundary frames as anchors; c) ***Motion Level***, modeling motion trajectories for realistic dynamics; and d) ***Semantic Level***, maintaining logical consistency between narrative and visuals. Recognizing that real-world videos are the best reflections of physical phenomena, we further introduce an automated data selection pipeline to efficiently identify and utilize *"good data"* from existing large-scale text-video datasets, thereby eliminating the need for costly and time-intensive dataset construction. Extensive experiments on both physics-focused and general capability benchmarks demonstrate that PhysHPO significantly improves physical plausibility and overall video generation quality of advanced models. To the best of our knowledge, this is the first work to explore fine-grained preference alignment and data selection for video generation, paving the way for more realistic and human-preferred video generation paradigms. Harold Haodong Chen, Haojian Huang, Qifeng Chen 0001, Harry Yang, Ser-Nam Lim |
NeurIPS | 2 |
| 2025 | L2-regularization based two-way weighted neutrosophic clustering with Manhattan and Euclidean distances
Haoye Qiu, Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan, Muhammet Deveci, Tapan Senapati |
Fuzzy Sets Syst. | 3 |
| 2024 | FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
Haojian Huang, Junhao Dong 0001, Mingzhe Zheng, Dian Shao |
ACM Multimedia | 2 |
| 2024 | CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning
Haojian Huang, Xiaozhen Qiao, Zhuo Chen 0007, Bingyu Li 0002, Zhe Sun 0007, Mulin Chen, Xuelong Li 0001 |
ACM Multimedia | 1 |
| 2024 | Adaptive weighted multi-view evidential clustering with feature preferenceabstractMulti-view clustering has attracted substantial attention thanks to its ability to integrate information from diverse views. However, the existing methods can only generate hard or fuzzy partitions, which cannot effectively represent the uncertainty and imprecision when facing objects in overlapping clusters, thus increasing the risk of error. To solve the above problems, in this paper, we propose an adaptive weighted multi-view evidential clustering (WMVEC) method based on the theory of belief functions to characterize the uncertainty and imprecision in cluster assignment. Technically, we integrate view weight assignments and credal partition between objects and cluster prototypes into a joint learning framework. The credal partition offers a more comprehensive insight into the data by enabling objects to be associated with not only singleton clusters but also subsets of these clusters (termed meta-clusters) and the empty set, which represents a noise cluster. To avoid the interference of irrelevant and redundant features, we further present a weighted multi-view evidential clustering with feature preference (WMVEC-FP) to learn the importance of each feature under different views. We suggest the objective functions of WMVEC and WMVEC-FP and design alternating optimization schemes to obtain the optimal solutions, respectively. Through an extensive array of experiments, it has been demonstrated that our proposed clustering methods outperform other related and state-of-the-art methods in terms of their advantages and overall effectiveness. Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan, Muhammet Deveci |
Knowl. Based Syst. | 2 |
| 2023 | Adaptive Weighted Multi-view Evidential Clustering
Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan |
ICANN (4) | 2 |
| 2023 | Comment on "New cosine similarity and distance measures for Fermatean fuzzy sets and TOPSIS approach"
Zhe Liu 0041, Haojian Huang |
Knowl. Inf. Syst. | 2 |