Xiangdong Huang 0002

dblp:154/0507-2 · also Xiang-Dong Huang 0002 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
14since 2021 · last 2027
0000-0003-1217-3809ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 13 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Neural network-based design of variable fractional-delay FIR filters with controllable cutoff frequency
Xinmiao Sun, Chaofang Hu, Jun Tang 0012, Xiangdong Huang 0002
Signal Process.5
2026 DRFusionRec: Enhancing Rationality and Diversity in Garment Recommendations
abstract
Fashion recommendation is crucial for consumers to express their self-image and personal style. To boost recommendation accuracy and rationality, researchers have explored state-of-the-art (SOTA) methods that incorporate items’ visual and textual information, along with their pairing records. However, these methods fail to fully explore how latent semantic-stylistic correlations between items influence recommendations, and inadequately tackle exposure bias, which results in skewed, restricted recommendations that prioritize popular or frequently observed outfits. To address these limitations, we propose DRFusionRec, a fashion recommender that enhances recommendation rationality and diversity by fully leveraging item semantic-stylistic correlations while mitigating exposure bias. Specifically, to harness rich item correlations, we introduce a novel Multi-Factor Relationship Measurement (MRM) matrix. It integrates semantic and stylistic features by mining the synergistic interaction probabilities across semantically and stylistically adjacent items, capturing the latent compatibility patterns. This matrix is then used to refine item features for richer details. To address homogeneity from exposure bias, we propose an Adaptive Propensity Score (APS) strategy. By dynamically weighting item popularity (direct influence) and popularity of style-similar neighbors (indirect contextual influence), we model exposure confounders to derive item propensity scores. Integrating these scores into item features effectively mitigates the bias. Lastly, MRM and APS synergistically optimize item representations for rationale-based and diverse recommendations. Experimental results validate that the DRFusionRec outperforms SOTA methods in capturing item compatibility, ensuring diverse recommendations, and maintaining reasonable complexity.
Wenxin Ding, Xiangdong Huang 0002, Shufang Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2025 KA-MIN: Knowledge-Aware Multimodal Interaction Network for Emotion Recognition in Conversation
abstract
Emotion recognition in conversations (ERC) has garnered significant attention for its critical role in human-computer interaction systems. ERC benefits from multimodal data, which offers diverse perspectives on emotional states, and commonsense knowledge (CSK), which enriches the context by incorporating real-world understanding of human behavior. However, existing ERC studies have not fully exploited the potential of multimodal-CSK interactions for complementary information learning from these sources. To address this, we innovatively propose a Knowledge-Aware Multimodal Interaction Network (KA-MIN). KA-MIN is designed to capture complementary emotional information from CSK-multimodal interactions, thereby facilitating the ERC task. To achieve this, KA-MIN begins by combining six relation types of CSK, leveraging their differences between multimodal emotional information. The fused CSK features are then refined to incorporate context and emotional information using multimodal contextual guidance. Subsequently, we construct a novel knowledge-aware multimodal graph structure that allows the CSK information to interact with multimodal information, leading to more comprehensive multimodal and context modeling. During the graph learning process, the CSK-multimodal interactions capture the complementary emotional information between CSK and multimodal features. Finally, we dynamically fuse the multimodal emotional information with the informative CSK and textual guidance to obtain the final utterance representations, which encompass effective emotional information from both multimodal and CSK features. Extensive experiments on two popular multimodal ERC datasets demonstrate the superiority and effectiveness of the proposed KA-MIN framework.
Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Zan Gao 0001, Yuting Su 0001, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2024 Robust adaptive beamforming algorithm based on coprime array with sensor gain-phase error
Xiangdong Huang 0002, Nian Hu
Signal Process.1
2023 Multi-loop graph convolutional network for multimodal conversational emotion recognition
Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Jing Liu 0002
J. Vis. Commun. Image Represent.2
2023 MALN: Multimodal Adversarial Learning Network for Conversational Emotion Recognition
abstract
Multimodal emotion recognition in conversations (ERC) aims to identify the emotional state of constituent utterances expressed by multiple speakers in dialogue from multimodal data. Existing multimodal ERC approaches focus on modeling the global context of the dialogue and neglect to mine the characteristic information from the corresponding utterances expressed by the same speaker. Additionally, information from different modalities exhibits commonality and diversity for emotional expression. The commonality and diversity of multimodal information are compensated for each other but not effectively exploited in previous multimodal ERC works. To tackle these issues, we propose a novel Multimodal Adversarial Learning Network (MALN). MALN first mines the speaker’s characteristics from context sequences and then incorporate them with the unimodal features. Afterward, we design a novel adversarial module AMDM to exploit both commonality and diversity from the unimodal features. Finally, AMDM fuses different modalities to generate refined utterance representations for emotion classification. Extensive experiments are conducted on two public multimodal ERC datasets, IEMOCAP and MELD. Through the experiments, MALN shows its superiority over the state-of-the-art methods.
Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Ming Liu 0004, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Cross-Domain Image-Object Retrieval Based on Weighted Optimal Transport
abstract
Given a 2D image query and a pool of 3D objects, the goal of image-object retrieval is to rank the 3D objects according to how well their content fits the query. Previous methods usually project 2D images and 3D objects into a joint embedding space and minimize the distance metric to complete the retrieval task. Since 2D images and 3D objects come from two different domains with large discrepancy, even when 3D objects and 2D images are mapped to a shared space, the gap in feature distribution remains significant, which always leads to domain misalignment. In this work, we propose a novel image-object retrieval method by leveraging optimal transport theory. Specifically, to tackle the dimensionality gap between 2D images and 3D objects, we first represent a 3D object via a sequence of its 2D projections. We then design a Cross-Domain View Attention module (CDVA) to automatically compute the optimal combination of 3D object projections given a 2D query image. Next, we exploit Weighted Optimal Transport (WOT)-based distance to depict the discrepancy between 2D images and 3D objects, and reduce the discrepancy to achieve instance-level alignment. Through this scheme, the transported 2D images and 3D objects with the same label are enforced to follow similar distributions. Finally, we design an explicit Category Centroid Alignment module (CCA) to achieve class-level alignment to improve the retrieval performance. Extensive experiments show that our method can achieve competitive performance on the MI3DOR and MI3DOR-2 benchmarks.
Nian Hu, Xiangdong Huang 0002, Wenhui Li 0001, Xuanya Li, Anan Liu
IEEE Trans. Multim.2
2022 Robust Multitarget Device-Free Localization and Tracking via Visible Light Sensing
abstract
Device-free localization (DFL) is a promising technique that can locate and track targets carrying no devices via detecting shadowed links caused by targets. There are two main challenging problems needed to be solved for the multitarget DFL. One is that pseudotargets can be produced by the intersection points of the shadowed links caused by different targets in the localization stage. The other is that targets’ identities may be wrongfully swapped when targets come close to each other in the tracking stage. To solve the two problems, a robust multitarget DFL framework using visible light sensing is proposed. The proposed framework introduces a positioning label to verify the authenticity of the estimated target. Moreover, a successive cancellation algorithm is utilized to obtain coarse target areas, which are used as the constraint information for calculating locations and features of targets. With the prior information related to the locations and features of targets, a multitarget tracking method based on the Hungarian algorithm and the spatial filter is proposed to solve the target assignment problem. Extensive simulation results show that the proposed multitarget DFL framework achieves outstanding performance both in localization and tracking accuracy.
Shuai Zhang 0017, Xiangdong Huang 0002
IEEE Internet Things J.4
2022 Collaborative Distribution Alignment for 2D image-based 3D shape retrieval
Nian Hu, Heyu Zhou, Anan Liu, Xiangdong Huang 0002, Shenyuan Zhang, Guoqing Jin, Junbo Guo, Xuanya Li
J. Vis. Commun. Image Represent.4
2022 Tracking by dynamic template: Dual update mechanism
Jing Liu 0002, Xiangdong Huang 0002, Yuting Su 0001
J. Vis. Commun. Image Represent.3
2022 Joint carrier and DOA estimation for multi-band sources based on sub-Nyquist sampling coprime array with large time lags
Xiangdong Huang 0002, Xuecheng Zhao, Jinying Ma
Signal Process.1
2022 A Feature Transformation Framework With Selective Pseudo-Labeling for 2D Image-Based 3D Shape Retrieval
abstract
2D image-based 3D shape retrieval (2D-to-3D) aims at searching the corresponding 3D shapes (unlabeled) when given a 2D image (labeled), which is a fundamental task in computer vision and has gained a surge of attention in recent years. However, extensive prior works are limited by two settings, 1) reducing domain discrepancy while ignoring the 3D shape style, 2) 3D shapes are simply and brutally pseudo-annotated by the 2D image-supervised classifier, neglecting the structure information underlying the 3D shape domain. To remedy these issues, we propose a feature transformation framework with selective pseudo-labeling (FTSPL) for 2D-to-3D task. Specifically, we first employ CNNs to produce both 2D image and 3D shape (described as multiple views) features, then we force the inter-domain centroid alignment class-wisely to reduce the overall domain discrepancy. In addition to this, we further exploit the intra-category attribute variation (covariance) of 3D shape features to transform the 2D image features. By doing so, we can equip 2D features with 3D shape style. Since the centroid and covariance estimation of 3D shape features require accurate label predictions, we put forward a selective pseudo-labeling module, which can assign reliable pseudo-labels for 3D shapes via nearest category centroid and cluster analysis, respectively, while preserving the structure information of 3D shapes. Comprehensive experiments validate that our model surpasses the state-of-the-arts on standard 2D-to-3D benchmarks (MI3DOR and MI3DOR-2).
Nian Hu, Heyu Zhou, Xiangdong Huang 0002, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2022 LR-GCN: Latent Relation-Aware Graph Convolutional Network for Conversational Emotion Recognition
abstract
As an intersection of artificial intelligence and human communication analysis, Emotion Recognition in Conversation (ERC) has attracted much research attention in recent years. Existing studies, however, are limited in adequately exploiting latent relations among the constituent utterances. In this paper, we address this issue by proposing a novel approach named Latent Relation-Aware Graph Convolutional Network (LR-GCN), where both speaker dependency of the interlocutors is leveraged and latent correlations among the utterances are captured for ERC. Specifically, we first establish a graph model to incorporate the context information and speaker dependency of the conversation. Afterward, the multi-head attention mechanism is introduced to explore the latent correlations among the utterances and generate a set of all-linked graphs. Here, aiming to simultaneously exploit the original modeled speaker dependency and the explored correlation information, we introduce a dense connection layer to capture more structural information of the generated graphs. Through a multi-branch graph network, we achieve a unified representation of each utterance for final prediction. Detailed evaluations on two benchmark datasets demonstrate LR-GCN outperforms the state-of-the-art approaches.
Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Dan Song 0006, Weizhi Nie
IEEE Trans. Multim.2
2021 Interactive Multimodal Attention Network for Emotion Recognition in Conversation
abstract
In this letter, we propose a novel Interactive Multimodal Attention Network (IMAN) for emotion recognition in conversations. IMAN introduces a cross-modal attention fusion module to capture cross-modal interactions of multimodal information, and employs a conversational modeling module to explore the context information and speaker dependency of the whole conversation. Concretely, the cross-modal attention fusion module captures the cross-modal interactions and complementary information among the pre-extracted unimodal features from textual, visual, acoustic modalities based on the cross-modal attention block. Afterward, the updated features from each modality are fused to concentrate more on the informative modality and achieve a refined feature for each constituent utterance. The conversational modeling module defines three different gated recurrent units (GRUs) with respect to the context information, the speaker dependency, and the emotional state of utterances. In this way, we exploit the speaker dependency and contextual information to obtain the emotional state of utterances for emotion classification. Empirical evaluations on the multimodal benchmark IEMOCAP dataset demonstrate that our IMAN achieves competitive performance compared to the state-of-the-art approaches.
Minjie Ren, Xiangdong Huang 0002, Xiaoqi Shi, Weizhi Nie
IEEE Signal Process. Lett.2
2019 Fast Palette Mode Decision Methods for Coding Game Videos With HEVC-SCC
abstract
The live broadcasting of game playing is becoming more and more popular. The game video is one of the typical video data which should be coded with high efficiency video coding-screen content coding (HEVC-SCC). In HEVC-SCC, the palette mode is an important technique which can effectively enhance the coding performance. However, this technique also demands high-coding computation. In this paper, we propose a fast palette mode pre-decision method based on the analysis of the color complexity of the video data. Considering only the coding computation of palette mode of HEVC-SCC, our algorithm saves about 74.24% computation in the experiments of game videos. Whereas, the BDBR only increases by about 0.36% and the BDPSNR decreases by about 0.03 dB on average.
Yu Liu 0004, Jinglin Sun, Xiangdong Huang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2018 Effective pattern recognition and find-density-peaks clustering based blind identification for underdetermined speech mixing systems
Xiangdong Huang 0002, Runan Song, Wei Lu 0026
Multim. Tools Appl.1
2018 Harmonics extraction based speech recovery for underdetermined mixing systems
Xiangdong Huang 0002
Multim. Tools Appl.2
2018 A Novel Multiconnected Convolutional Network for Super-Resolution
abstract
Convolutional neural networks exhibit superior performance for single image super-resolution (SISR) tasks. However, as the network grows deeper, features from the earlier layers are impeded or less used in later layers. In SISR, the earlier layers are mainly composed of local features that are essential to the task. In this letter, we present a novel multiconnected convolutional network for SISR tasks by enhancing the combination of both low- and high-level features. We design a structure built on multiconnected blocks to extract diversified and complicated features via the concatenation of low-level features to high-level features. In addition to stacking multiconnected blocks, a long skip-connection is implemented to further aggregate features of the first layer and a specific later layer. Furthermore, we employ a flexible two-parameter loss function to optimize the training process. The proposed method yields state-of-the-art performance both in terms of quantitative metrics and visual quality. The method also outperforms existing methods on datasets via unknown degrading operators, indicating an excellent generalization ability.
Jinghui Chu, Wei Lu 0026, Xiangdong Huang 0002
IEEE Signal Process. Lett.4