Hongchen Luo

dblp:295/9297 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
20since 2021 · last 2027
0000-0003-0744-2074ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2027 START: Structural-semantic collaborative perception transfer for incremental 3D affordance grounding
Pengyi Zhao, Hongchen Luo, Jiao Wang 0005
Expert Syst. Appl.2
2026 I²B-LPO: Latent Policy Optimization via Iterative Information Bottleneck
abstract
Huilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming LI, Chuyang Zhao, Jihai Zhang, MengChang Wang, Yang Cao, Yu Kang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Huilin Deng, Hongchen Luo, Zhuoyue Chen, Xinghao Zhao, Chuyang Zhao, Mengchang Wang, Yang Cao 0010, Yu Kang 0001
ACL (1)2
2026 Visual-Geometric Collaborative Guidance for Affordance Learning
Hongchen Luo, Wei Zhai, Jiao Wang 0002, Yang Cao 0010, Zhengjun Zha
Int. J. Comput. Vis.1
2026 Integrating causal-context-aware teammate modeling with risk-constrained online planning for meta reinforcement learning in Ad Hoc teamwork
Jiao Wang 0002, Hongchen Luo
Knowl. Based Syst.4
2026 VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) enables the inspection of unseen objects by bridging textual prompts and visual features, showing great potential in flexible manufacturing. While existing ZSAD methods rely on predefined prompts and struggle with unseen defects, Multimodal Large Language Models (MLLMs) offer promising solutions through their generative and interpretative capabilities. However, adapting MLLMs to Industrial Anomaly Detection (IAD) remains challenging due to fine-grained anomaly patterns and subtle visual distinctions. We propose VMAD (Visual-enhanced MLLM Anomaly Detection), a framework that enriches MLLM with visual IAD knowledge through two key components: a Defect-Sensitive Structure Learning scheme that transfers patch-similarities for improved discrimination, and a Locality-enhanced Token Compression that leverages multi-level local features for fine-grained detection. We also introduce RIAD, a comprehensive IAD dataset with detailed anomaly annotations. Extensive experiments on MVTec-AD, Visa, WFDD, and RIAD demonstrate VMAD’s superior performance. The dataset and code will be publicly available at https://github.com/denghuilin-cyber/VMAD.
Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001
IEEE Trans Autom. Sci. Eng.2
2026 Corrections to "VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection"
abstract
In the above article [1], an earlier draft of Fig. 6 was inadvertently included. The correct Fig. 6 is presented on the next page.Fig. 6.Zero-shot anomaly segmentation on MVTec-AD, WFDD, and ViSA datasets.
Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001
IEEE Trans Autom. Sci. Eng.2
2026 GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification
abstract
Glaucoma is a leading cause of irreversible blindness worldwide, with asymptomatic early stages often delaying diagnosis and treatment. Early and accurate diagnosis requires integrating complementary information from multiple ocular imaging modalities. However, most existing studies rely on single- or dual-modality imaging, such as fundus and optical coherence tomography (OCT), for coarse binary classification, thereby restricting the exploitation of complementary information and hindering both early diagnosis and stage-specific treatment. To address these limitations, we propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT images, and visual field pattern deviation maps, annotated with four disease stages, enabling effective exploitation of multimodal complementary information and facilitating accurate diagnosis and treatment across disease stages. To effectively integrate cross-modal information, we propose hierarchical attentive masked modeling (HAMM) for multimodal glaucoma classification. Our framework employs hierarchical attentive encoders and light decoders to focus cross-modal representation learning on the encoder. The attention module, named multimodal-channel graph attention (MCGA), boosts glaucoma classification performance by emulating two key clinical reasoning steps: first, it uses a multi-head modality gating mechanism to replicate ophthalmologists' confidence scoring of fundus, OCT, and VF modalities; then, MCGA leverages a relational graph attention network to cross-examine structural-functional consistencies of weighted modalities. The experiments on GLEAM demonstrate that tri-modal fusion significantly outperforms single-modal and dual-modal configurations. Moreover, our proposed HAMM achieves superior performance compared with state-of-the-art multimodal learning methods. The dataset and code are publicly available via https://github.com/microewing/HAMM.
Jiao Wang 0005, Hongchen Luo, Zhifen Guo, Ruiting Zhou, Man Tang
IEEE Trans. Medical Imaging4
2025 GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding
abstract
Open-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages that depict interactions with 3D geometries to introduce external interaction priors. However, they are still vulnerable to a limited semantic space by failing to leverage implied invariant geometries and potential interaction intentions. Normally, humans address complex tasks through multi-step reasoning and respond to diverse situations by leveraging associative and analogical thinking. In light of this, we propose GREAT (GeometRy-intEntion collAboraTive inference) for Open-Vocabulary 3D Object Affordance Grounding, a novel framework that mines the object invariant geometry attributes and performs analogically reason in potential interaction scenarios to form affordance knowledge, fully combining the knowledge with both geometries and visual contents to ground 3D object affordance. Besides, we introduce the Point Image Affordance Dataset v2 (PIADv2), the largest 3D object affordance dataset at present to support the task. Extensive experiments demonstrate the effectiveness and superiority of GREAT. The code and dataset are available at https://yawen-shao.github.io/GREAT/.
Yawen Shao, Wei Zhai, Yuhang Yang 0002, Hongchen Luo, Yang Cao 0010, Zhengjun Zha
CVPR4
2025 Multi-agent reinforcement learning for cooperative search under aperiodically intermittent communication
Longyue Fu, Jiao Wang 0002, Hongchen Luo
Expert Syst. Appl.3
2025 Temporal-spectral-spatial synchronization attention-based network for EEG emotion recognition
Zhifen Guo, Jiao Wang 0002, Hongchen Luo, Fengbin Ma
Knowl. Based Syst.3
2024 LEMON: Learning 3D Human-Object Interaction Relation from 2D Images
abstract
Learning 3D human-object interaction relation is piv-otal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements, e.g., human contact, object affordance, and human-object spatial relation, primarily from the perspective of either the human or the object. Which underexploit certain correlations between the interaction counterparts (human and object), and struggle to address the uncertainty in interactions. Actually, objects' functionalities potentially affect humans' interaction intentions, which reveals what the interaction is. Mean-while, the interacting humans and objects exhibit matching geometric structures, which presents how to interact. In light of this, we propose harnessing these inherent correlations between interaction counterparts to mitigate the uncertainty and jointly anticipate the above interaction el-ements in 3D space. To achieve this, we present LEMON (LEarning 3D huMan-Object iNteraction relation), a unified model that mines interaction intentions of the counter-parts and employs curvatures to guide the extraction of ge-ometric correlations, combining them to anticipate the interaction elements. Besides, the 3D Interaction Relation dataset (3DIR) is collected to serve as the test bed for training and evaluation. Extensive experiments demonstrate the superiority of LEMON over methods estimating each element in isolation. The code and dataset are available at https://yyvhang.github.io/LEMON.
Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Zhengjun Zha
CVPR3
2024 Bidirectional Progressive Transformer for Interaction Intention Anticipation
Zichen Zhang 0022, Hongchen Luo, Wei Zhai, Yang Cao 0010, Yu Kang 0001
ECCV (59)2
2024 Grounded Affordance from Exocentric View
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
Int. J. Comput. Vis.1
2024 Learning Visual Affordance Grounding From Demonstration Videos
abstract
Visual affordance grounding aims to segment all possible interaction regions between people and objects from an image/video, which benefits many applications, such as robot grasping and action recognition. Prevailing methods predominantly depend on the appearance feature of the objects to segment each region of the image, which encounters the following two problems: 1) there are multiple possible regions in an object that people interact with and 2) there are multiple possible human interactions in the same object region. To address these problems, we propose a hand-aided affordance grounding network (HAG-Net) that leverages the aided clues provided by the position and action of the hand in demonstration videos to eliminate the multiple possibilities and better locate the interaction regions in the object. Specifically, HAG-Net adopts a dual-branch structure to process the demonstration video and object image data. For the video branch, we introduce hand-aided attention to enhance the region around the hand in each video frame and then use the long short-term memory (LSTM) network to aggregate the action features. For the object branch, we introduce a semantic enhancement module (SEM) to make the network focus on different parts of the object according to the action classes and utilize a distillation loss to align the output features of the object branch with that of the video branch and transfer the knowledge in the video branch to the object branch. Quantitative and qualitative evaluations on two challenging datasets show that our method has achieved state-of-the-art results for affordance grounding. The source code is available at: https://github.com/lhc1224/HAG-Net.
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2023 Leverage Interactive Affinity for Affordance Learning
abstract
Perceiving potential “action possibilities” (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assignment paradigm and presume that there is a unique relationship between functional region and affordance label, yielding poor performance when adapting to unseen environments with large appearance variations. In this paper, we propose to leverage interactive affinity for affordance learning, i.e. extracting interactive affinity from human-object interaction and transferring it to non-interactive objects. Interactive affinity, which represents the contacts between different parts of the human body and local regions of the target object, can provide inherent cues of interconnectivity between humans and objects, thereby reducing the ambiguity of the perceived action possibilities. Specifically, we propose a pose-aided interactive affinity learning framework that exploits human pose to guide the network to learn the interactive affinity from human-object interactions. Particularly, a keypoint heuristic perception (KHP) scheme is devised to exploit the keypoint association of human pose to alleviate the uncertainties due to interaction diversities and contact occlusions. Besides, a contact-driven affordance learning (CAL) dataset is constructed by collecting and labeling over 5, 000 images. Experimental results demonstrate that our method outperforms the representative models regarding objective metrics and visual quality. Code and dataset: github.com/lhc1224/PIAL-Net.
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
CVPR1
2023 Grounding 3D Object Affordance from 2D Interactions in Images
abstract
Grounding 3D object affordance seeks to locate objects’ "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotations to declare interactive regions of interest on the object and establishing a mapping between the regions and affordances. However, the essence of learning object affordance is to understand how to use it, and the manner that detaches interactions is limited in generalization. Normally, humans possess the ability to perceive object affordances in the physical world through demonstration images or videos. Motivated by this, we introduce a novel task setting: grounding 3D object affordance from 2D interactions in images, which faces the challenge of anticipating affordance through interactions of different sources. To address this problem, we devise a novel Interaction-driven 3D Affordance Grounding Network (IAG), which aligns the region feature of objects from different sources and models the interactive contexts for 3D object affordance grounding. Besides, we collect a Point-Image Affordance Dataset (PIAD) to support the proposed task. Comprehensive experiments on PIAD demonstrate the reliability of the proposed task and the superiority of our method. The project is available at https://github.com/yyvhang/IAGNet.
Yuhang Yang 0002, Wei Zhai, Hongchen Luo, Yang Cao 0010, Jiebo Luo 0001, Zhengjun Zha
ICCV3
2022 Learning Affordance Grounding from Exocentric Images
abstract
Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability that transform the various exocentric interactions to invariant egocentric affordance so as to counter the impact of interactive diversity. To empower an agent with such ability, this paper proposes a task of affordance grounding from exocentric view, i.e., given exocentric human-object interaction and egocentric object images, learning the affordance knowledge of the object and transferring it to the egocentric image using only the affordance label as supervision. To this end, we devise a cross-view knowledge transfer framework that extracts affordance-specific features from exocentric interactions and enhances the perception of affordance regions by preserving affordance correlation. Specifically, an Affordance Invariance Mining module is devised to extract specific clues by minimizing the intra-class differences originated from interaction habits in exocentric images. Besides, an Affordance Co-relation Preserving strategy is presented to perceive and localize affordance by aligning the co-relation matrix of predicted results between the two views. Particularly, an affordance grounding dataset named AGD20K is constructed by collecting and labeling over 20K images from 36 affordance categories. Experimental results demonstrate that our method outperforms the representative models in terms of objective metrics and visual quality. Code: github.com/lhc1224/Cross-View-AG.
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
CVPR1
2022 Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation
Jinchi Huang, Bowen Cai 0001, Huan Fu, Mingming Gong, Chaohui Wang, Hongchen Luo, Rongfei Jia, Binqiang Zhao
ECCV (15)8
2022 One-Shot Object Affordance Detection in the Wild
Wei Zhai, Hongchen Luo, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
Int. J. Comput. Vis.2
2021 One-Shot Affordance Detection
abstract
Affordance detection refers to identifying the potential action possibilities of objects in an image, which is an important ability for robot perception and manipulation. To empower robots with this ability in unseen scenarios, we consider the challenging one-shot affordance detection problem in this paper, i.e., given a support image that depicts the action purpose, all objects in a scene with the common affordance should be detected. To this end, we devise a One-Shot Affordance Detection (OS-AD) network that firstly estimates the purpose and then transfers it to help detect the common affordance from all candidate images. Through collaboration learning, OS-AD can capture the common characteristics between objects having the same underlying affordance and learn a good adaptation capability for perceiving unseen affordances. Besides, we build a Purpose-driven Affordance Dataset (PAD) by collecting and labeling 4k images from 31 affordance and 72 object categories. Experimental results demonstrate the superiority of our model over previous representative ones in terms of both objective metrics and visual quality. The benchmark suite is at ProjectPage.
Hongchen Luo, Wei Zhai, Jing Zhang 0037, Yang Cao 0010, Dacheng Tao
IJCAI1