Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Haojian Huang

dblp:355/4169 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-0661-712XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 21% Language models and text generation · 19% Vision and language · 14%
Computer graphics and multimedia
2 papers
Image and video coding · 88% Multimedia analysis and retrieval · 12%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
uncertainty estimation
2.632026
Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026
Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024
Machine learning › Trustworthy machine learning › uncertainty estimation › neural network uncertainty
evidential deep learning
1.932026
Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024
Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
1.722025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
1.122025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Natural language and speech › Language models and text generation › LLM agents
agent memory
1.012026
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer · AAAI 2026
Computer vision › Vision and language
moment retrieval
1.012026
Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
1.012026
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.912025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
Computer vision › 3D vision
physical plausibility
0.912025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
Machine learning › Generative modeling › video generation
physics-informed video generation
0.912025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
Computer vision › Video understanding and tracking › spatio-temporal understanding
spatio-temporal video understanding
0.912025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Computer vision › Video understanding and tracking › video analytics
sports video analysis
0.912025
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning · ACM Multimedia 2025
Machine learning › Generative modeling
video generation
0.912025
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation · NeurIPS 2025
Computer vision › Vision and language
video-language model
0.912025
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models · ICML 2025
Data mining › predictive modeling
classification
0.912025
Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025
Data mining › predictive modeling › classification › pattern classification
multi-view classification
0.912025
Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification · AAAI 2025
Image and video coding › image quality assessment
AI-generated image quality assessment
0.912025
Text-Visual Semantic Constrained AI-Generated Image Quality Assessment · ACM Multimedia 2025
Image and video coding
image quality assessment
0.912025
Text-Visual Semantic Constrained AI-Generated Image Quality Assessment · ACM Multimedia 2025
Computer vision › Face, body and person analysis › facial expression analysis › facial expression recognition
dynamic facial expression recognition
0.812024
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.812024
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024
Machine learning › Transfer learning and domain adaptation › zero-shot learning
multi-modal zero-shot learning
0.812024
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.812024
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.312025
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning · ACM Multimedia 2025
Computer vision › Vision and language
vision-language model
0.212024
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs · ACM Multimedia 2024
Multimedia analysis and retrieval › multimodal learning
multimodal representation learning
0.212024
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

evidential deep learning · 3.3markov random field · 1.7feature-neighborhood structure · 1.7subjective experience-based memory · 1.0query reconstruction · 1.0maze navigation · 1.0match-2 elimination · 1.0deep evidential regression · 1.0cross-modal fusion · 1.0direct preference optimization · 0.9cross-modal resonance · 0.8
YearPublicationVenuePosition
2026 Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval
abstract
In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval.
Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He
AAAI1
2026 EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
abstract
Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce two dynamic spatial benchmarks—locally observable maze navigation and match-2 elimination—that systematically evaluate models' abilities in spatial understanding and adaptive planning when local perception, environment feedback, and global objectives are tightly coupled. Each action triggers structural changes in the environment, requiring continuous update of cognition and strategy. We further propose a subjective experience-based memory mechanism for cross-task experience transfer and validation. Experiments show that our benchmarks reveal key limitations of mainstream models in dynamic spatial reasoning and long-term memory, providing a comprehensive platform for future methodological advances.
Pukun Zhao, Miaowei Wang, Fanqing Zhou, Haojian Huang
AAAI6
2025 Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification
abstract
Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views.
Haojian Huang, Chuanyu Qin, Zhe Liu 0041, Kaijing Ma, Han Fang 0002, Chao Ban, Hao Sun 0038, Zhongjiang He
AAAI1
2025 VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Haojian Huang, Shengqiong Wu, Meng Luo 0010, Jinlan Fu, Xinya Du, Hanwang Zhang, Hao Fei 0001
ICML1
2025 FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
Haojian Huang, Xinxiang Yin, Dian Shao
ACM Multimedia2
2025 Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001
ACM Multimedia3
2025 Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
abstract
Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose **PhysHPO**, a novel framework for Hierarchical Cross-Modal Direct Preference Optimization, to tackle this challenge by enabling fine-grained preference alignment for physically plausible video generation. PhysHPO optimizes video alignment across four hierarchical granularities: a) ***Instance Level***, aligning the overall video content with the input prompt; b) ***State Level***, ensuring temporal consistency using boundary frames as anchors; c) ***Motion Level***, modeling motion trajectories for realistic dynamics; and d) ***Semantic Level***, maintaining logical consistency between narrative and visuals. Recognizing that real-world videos are the best reflections of physical phenomena, we further introduce an automated data selection pipeline to efficiently identify and utilize *"good data"* from existing large-scale text-video datasets, thereby eliminating the need for costly and time-intensive dataset construction. Extensive experiments on both physics-focused and general capability benchmarks demonstrate that PhysHPO significantly improves physical plausibility and overall video generation quality of advanced models. To the best of our knowledge, this is the first work to explore fine-grained preference alignment and data selection for video generation, paving the way for more realistic and human-preferred video generation paradigms.
Harold Haodong Chen, Haojian Huang, Qifeng Chen 0001, Harry Yang, Ser-Nam Lim
NeurIPS2
2025 L2-regularization based two-way weighted neutrosophic clustering with Manhattan and Euclidean distances
Haoye Qiu, Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan, Muhammet Deveci, Tapan Senapati
Fuzzy Sets Syst.3
2024 FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
Haojian Huang, Junhao Dong 0001, Mingzhe Zheng, Dian Shao
ACM Multimedia2
2024 CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning
Haojian Huang, Xiaozhen Qiao, Zhuo Chen 0007, Bingyu Li 0002, Zhe Sun 0007, Mulin Chen, Xuelong Li 0001
ACM Multimedia1
2024 Adaptive weighted multi-view evidential clustering with feature preference
abstract
Multi-view clustering has attracted substantial attention thanks to its ability to integrate information from diverse views. However, the existing methods can only generate hard or fuzzy partitions, which cannot effectively represent the uncertainty and imprecision when facing objects in overlapping clusters, thus increasing the risk of error. To solve the above problems, in this paper, we propose an adaptive weighted multi-view evidential clustering (WMVEC) method based on the theory of belief functions to characterize the uncertainty and imprecision in cluster assignment. Technically, we integrate view weight assignments and credal partition between objects and cluster prototypes into a joint learning framework. The credal partition offers a more comprehensive insight into the data by enabling objects to be associated with not only singleton clusters but also subsets of these clusters (termed meta-clusters) and the empty set, which represents a noise cluster. To avoid the interference of irrelevant and redundant features, we further present a weighted multi-view evidential clustering with feature preference (WMVEC-FP) to learn the importance of each feature under different views. We suggest the objective functions of WMVEC and WMVEC-FP and design alternating optimization schemes to obtain the optimal solutions, respectively. Through an extensive array of experiments, it has been demonstrated that our proposed clustering methods outperform other related and state-of-the-art methods in terms of their advantages and overall effectiveness.
Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan, Muhammet Deveci
Knowl. Based Syst.2
2023 Adaptive Weighted Multi-view Evidential Clustering
Zhe Liu 0041, Haojian Huang, Sukumar Letchmunan
ICANN (4)2
2023 Comment on "New cosine similarity and distance measures for Fermatean fuzzy sets and TOPSIS approach"
Zhe Liu 0041, Haojian Huang
Knowl. Inf. Syst.2