Kaixiang Huang

dblp:140/0880 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-1256-153XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 71% Image and video processing · 29%
Artificial intelligence
3 papers
3D vision · 51% Vision and language · 22% Trustworthy machine learning · 16%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
3d scene graph
1.012026
GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › 3D vision
3d scene understanding
1.012026
GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › Vision and language
cross-modal matching
1.012026
Purified Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2026
Multimedia analysis and retrieval
image retrieval
1.012026
Purified Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2026
Multimedia analysis and retrieval › image retrieval
sketch-based image retrieval
1.012026
Purified Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2026
Multimedia analysis and retrieval › image retrieval › sketch-based image retrieval
zero-shot sketch-based image retrieval
1.012026
Purified Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2026
Computer vision › 3D vision › 3d scene understanding
3d visual grounding
0.912025
Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding · NeurIPS 2025
Image and video processing
document image analysis
0.912025
Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch Grounding · ACM Multimedia 2025
Image and video processing › document image analysis › graphics recognition
handwritten mathematical expression recognition
0.912025
Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch Grounding · ACM Multimedia 2025
Multimedia analysis and retrieval › multimodal emotion recognition
emotion recognition in conversation
0.312018
Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder · ACM Multimedia 2018
Multimedia analysis and retrieval › multimedia analysis
multimodal conversation analysis
0.312018
Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder · ACM Multimedia 2018
Multimedia analysis and retrieval › affective computing
sentiment analysis
0.312018
Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder · ACM Multimedia 2018
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.312026
Purified Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2026
Computer vision › Segmentation and scene understanding
scene graph
0.312026
GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

visual-cross-linguistic sampler · 2.0transformer decoder · 2.0masked matching · 2.0volumetric pooling · 1.0granularity transformer · 1.0multimodal sketch grounding · 0.9multimodal reasoning · 0.9chain-of-thought · 0.9modality attention fusion · 0.3hierarchical encoder-decoder · 0.3attention mechanism · 0.3
YearPublicationVenuePosition
2026 A collision-free motion planning method for cable-drive redundant manipulators with deep reinforcement learning-based expert guidance and long short-term memory
Biyi Cheng, Xinde Zhang, Kaixiang Huang, Chiliang Zhong, Yingyuan Guan, Xueming Yin, Yuyuan Qiu
Expert Syst. Appl.5
2026 GasSeg: A lightweight real-time infrared gas segmentation network for edge devices
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Guodong Lu, Shengfeng He
Pattern Recognit.4
2026 An Intelligent Multitask Framework for Industrial Gas Leak Detection and Analysis With Infrared Optical Gas Imaging
abstract
Infrared (IR) optical gas imaging (OGI) is widely adopted in industrial environments for detecting fugitive gas emissions. However, conventional IR OGI systems rely heavily on manual inspection, lacking capabilities for active leak localization and in-depth analysis, which increases labor costs and risks of human error. To address these challenges, we present LeakHunter, an intelligent multitask framework designed for industrial gas leak monitoring and decision support. LeakHunter integrates seamlessly with IR cameras and can be deployed on edge computing devices, enabling real-time, on-site leak detection in harsh industrial settings. At the core of LeakHunter is a novel keypoint detection paradigm tailored for IR OGI, capable of localizing both leak sources and diffusion endpoints to enable effective spatiotemporal trend analysis. The framework also estimates critical leak attributes, including plume morphology and flow rate, supporting rapid and informed response. To further enhance detection accuracy, we introduce a biomimetic attention module that improves gas-background separation under complex thermal conditions, and a collaborative multitask head for efficient cross-task feature sharing. In addition, two benchmark datasets are proposed, one of which is a field-test set collected in real industrial scenarios. Experiments demonstrate that LeakHunter achieves state-of-the-art performance across multiple tasks, with an F2 score of 94.8% for gas segmentation and 97.9% for leak keypoint localization, while running at 28.4 FPS on a portable IR OGI device. These results highlight its potential as a deployable, intelligent solution for enhancing industrial safety and automation.
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Zaixing He, Guodong Lu
IEEE Trans. Ind. Informatics4
2026 Purified Zero-Shot Sketch-Based Image Retrieval
abstract
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler that generates linguistic masks based on semantic labels to obscure semantically irrelevant image features. Our novel contribution is the concept of purified masked matching (PMM), which comprises two processes: (1)reconstruction, which compels the image encoder to reconstruct the masked image feature, and (2)interaction, which involves a transformer decoder that processes both sketch and masked image features to investigate cross-domain relationships for effective matching. Evaluated on the TU-Berlin, Sketchy, and QuickDraw datasets, PuXIM sets new benchmarks in terms of performance. Importantly, the distraction-agnostic nature of the matching process renders PuXIM more conducive to training, enabling efficient adaptation to zero-shot scenarios with reduced data requirements and low data quality.
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Shengfeng He
IEEE Trans. Multim.4
2026 GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction
abstract
Predicting 3D Semantic Scene Graphs (3DSSG) is vital for understanding complex scenes by constructing structured representations. Current methods struggle with significant granularity discrepancies among instances, often relying on features at a single scale, which hampers their ability to perceive and interact with differently sized instances. To tackle this challenge, we introduce GranSSG, a novel approach that integrates volumetric granular awareness into 3DSSG prediction. Central to GranSSG is the Volumetric Pooling block, which aggregates features from multiple instance volumes, enhancing the representation of instance patterns across different granularities. Complementing this, the Granularity Transformer block dynamically directs attention to instance features across various network layers, ensuring precise perception of instances regardless of their granularity. Furthermore, the Cross-Granularity Correlation Transformer block mitigates performance degradation in instance pair relationship prediction by adaptively fusing hybrid features from different granularities, providing a comprehensive representation of instance pairs. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that GranSSG significantly enhances prediction performance, setting a new state-of-the-art in 3DSSG prediction.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Jiao Yi, Guodong Lu, Shengfeng He
IEEE Trans. Vis. Comput. Graph.1
2025 Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch Grounding
Jin Wang 0015, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
ACM Multimedia4
2025 Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding
abstract
3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights using a Corroborative Refinement strategy, which adaptively reorganizes information to correct ambiguities arising from biased or incomplete observations. Through this two-stage deliberation, Refer-Judge significantly enhances the reliability of data judgments. Extensive experiments demonstrate that our framework not only achieves human-level discrimination at the scene level but also improves the performance of baseline algorithms via data purification. Code is available at https://github.com/Hermione-HKX/Refer_Judge.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Huan Yu 0002, Guodong Lu, Shengfeng He
NeurIPS1
2024 Cross-Modal Pixel-and-Stroke representation aligning networks for free-hand sketch recognition
Jin Wang 0015, Jingru Yang, Ping Ni, Guodong Lu, Heming Fang, Huan Yu 0002, Kaixiang Huang
Expert Syst. Appl.9
2024 IEFM and IDS: Enhancing 3D environment perception via information encoding in indoor point cloud semantic segmentation
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Guodong Lu, Huan Yu 0002
Neurocomputing1
2024 MsVFE and V-SIAM: Attention-based multi-scale feature interaction and fusion for outdoor LiDAR semantic segmentation
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Huan Yu 0002, Wenming Zou
Neurocomputing3
2024 Granular3D: Delving into multi-granularity 3D scene graph prediction
abstract
This paper addresses the significant challenges in 3D Semantic Scene Graph (3DSSG) prediction, essential for understanding complex 3D environments. Traditional approaches, primarily using PointNet and Graph Convolutional Networks , struggle with effectively extracting multi-grained features from intricate 3D scenes , largely due to a focus on global scene processing and single-scale feature extraction. To overcome these limitations, we introduce Granular3D, a novel approach that shifts the focus towards multi-granularity analysis by predicting relation triplets from specific sub-scenes. One key is the Adaptive Instance Enveloping Method (AIEM), which establishes an approximate envelope structure around irregular instances, providing shape-adaptive local point cloud sampling, thereby comprehensively covering the contextual environments of instances. Moreover, Granular3D incorporates a Hierarchical Dual-Stage Network (HDSN), which differentiates and processes features of instances and their pairs at varying scales, leading to a targeted prediction of instance categories and their relationships. To advance the perception of sub-scene in HDSN, we design a Gather Point Transformer structure (GaPT) that enables the combinatorial interaction of local information from multiple point cloud sets, achieving a more comprehensive local contextual feature extraction. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that our methods provide substantial improvements, establishing a new state-of-the-art in 3DSSG prediction, boosting the top-50 triplet accuracy by +2.8%.
Kaixiang Huang, Jingru Yang, Jin Wang 0015, Shengfeng He, Zhan Wang 0002, Haiyan He, Guodong Lu
Pattern Recognit.1
2018 Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder
abstract
Human conversation analysis is challenging because the meaning can be expressed through words, intonation, or even body language and facial expression. We introduce a hierarchical encoder-decoder structure with attention mechanism for conversation analysis. The hierarchical encoder learns word-level features from video, audio, and text data that are then formulated into conversation-level features. The corresponding hierarchical decoder is able to predict different attributes at given time instances. To integrate multiple sensory inputs, we introduce a novel fusion strategy with modality attention. We evaluated our system on published emotion recognition, sentiment analysis, and speaker trait analysis datasets. Our system outperformed previous state-of-the-art approaches in both classification and regressions tasks on three datasets. We also outperformed previous approaches in generalization tests on two commonly used datasets. We achieved comparable performance in predicting co-existing labels using the proposed model instead of multiple individual models. In addition, the easily-visualized modality and temporal attention demonstrated that the proposed attention mechanism helps feature selection and improves model interpretability.
Xinyu Li 0003, Kaixiang Huang, Shiyu Fu, Kangning Yang, Shuhong Chen, Moliang Zhou, Ivan Marsic
ACM Multimedia3
2013 One Wide-Sense Circuit Tree per Traffic Class Based Inter-domain Multicast
abstract
Traditional multicast protocol forms multicast trees rooted at different sources to forward packets. If the multicast sources and receivers are in different domains, these trees will produce a great number of multicast states in the backbone, resulting in poor scalability. Therefore, we propose a one Wide-Sense Circuit Tree per Traffic Class based inter-domain multicast (WSCT-TC), in which a Wide-Sense Circuit Tree (WSCT) is established for a class of multicast traffic. The WSCT is established in the backbone, along which multicast packets are forwarded by label switching. The spec of WSCT can be reconfigured according to the QoS (Quality of Service) requirement of multicast applications, to provide preferable QoS. Simulating experiment shows that WSCT-TC behaves better scalability.
Chaoling Li, Kaixiang Huang
NAS3