VLDB 2026 Research / reviewers in the wild / expert
Wenze Huang
dblp:354/2800
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0007-0344-2233ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 40% Face, body and person analysis · 23% Efficient and distributed learning · 23% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition |
1.0 | 1 | 2026 | HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition · AAAI 2026 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
1.0 | 1 | 2026 | HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition · AAAI 2026 |
Computer vision › Video understanding and tracking
action segmentation |
0.9 | 1 | 2025 | Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025 |
Computer vision › Video understanding and tracking › action segmentation › human action segmentation
skeleton-based action segmentation |
0.9 | 1 | 2025 | Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model |
0.3 | 1 | 2026 | HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition · AAAI 2026 |
Machine learning › Graph learning › graph structure learning
relation graph learning |
0.3 | 1 | 2025 | Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation · IEEE Trans. Image Process. 2025 |
Methods — techniques the papers use, named apart from their topics
multi-scale adapters · 1.0kronecker product · 1.0dual-branch interactive router · 1.0spatio-temporal modeling · 0.9large language model · 0.9graph neural network · 0.9contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression RecognitionabstractFacial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters. Yu Gao 0010, Haoyu Ji 0001, Zhiyong Wang 0009, Wenze Huang, Xueting Liu 0009, Weihong Ren, Honghai Liu 0001 |
AAAI | 4 |
| 2026 | Topology-Motion Decoupling Framework With Textual Regularization for Skeleton-Based Temporal Action SegmentationabstractSkeleton-based temporal action segmentation aims to capture key information in long skeleton motion sequences to temporally segment and identify actions at a fine-grained level. Existing approaches have achieved promising results by improving the modeling of topological spatial relationships and long-term temporal dependencies. However, current methods often overlook the distinct nature of motion and topological information, applying a monolithic modeling paradigm to both. This approach fails to fully exploit their differential contributions to precise boundary localization and effective class discrimination. To address these limitations, we propose a novel Topology-Motion Decoupling Framework (TMD). Our framework incorporates three key designs. First, an auxiliary Differential Motion Perception Branch explicitly models the temporal gradients of skeletal trajectory to decouple boundary-sensitive motion features. Second, we introduce two effective fusion modules that integrate the complementary features from both branches for mutual enhancement. Finally, a Boundary-Aware Textual Regularization scheme leverages a dual set of semantic prompts for boundary/non-boundary to differentially guide the feature learning process. By design, TMD explicitly mitigates semantic and temporal confusion between actions, thereby enhancing inter-class discriminability and boundary awareness. Extensive experiments on five challenging public datasets demonstrate that our TMD achieves state-of-the-art performance. Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action SegmentationabstractSkeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results. Haoyu Ji 0001, Bowen Chen 0004, Weihong Ren, Wenze Huang, Zhiyong Wang 0009, Honghai Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action SegmentationabstractThe skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks (GCNs) for intraframe spatial modeling and temporal convolutional networks (TCNs) for interframe temporal modeling to discern motion patterns. However, these approaches often overlook the distinctive nature of essential action elements across various actions, including engaged core body parts and key subactions. This oversight limits the ability to distinguish different actions within a given sequence. To address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along joints and sequences to identify core joints and key subactions at different scales. Specifically, in terms of the spatial domain, the intrasnippet cross-joint attention (CJA) module divides the sequence into distinct snippets and computes attention to establish intricate joint semantic relationships, emphasizing the identification of core motion joints. In terms of the temporal domain, in the encoder, the intrasnippet cross-frame attention (CFA) module segments the sequence in a blockwise expansion manner and establishes interframe relationships to highlight the most discriminative frames. In the decoder, clip-level representations at various temporal scales are initially generated through an hourglass-like sampling process, followed by the intrasnippet cross-scale attention (CSA) module to integrate the key clip information across different time scales. The performance evaluation on five public datasets demonstrates that ME-ST achieves state-of-the-art (SOTA) performance. Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |