Haoyang Luo

dblp:308/6966 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedRS: Federated Learning Under Reliable Supervision for Multi-Organ Segmentation With Inconsistent Labels
abstract
Existing multi-organ segmentation methods usually rely on large and fully labeled datasets for training. However, medical image datasets are typically decentralized by privacy constraints and partially labeled due to the high costs of full annotation in clinical practice, resulting in label inconsistency across medical centers. Federated learning offers privacy-preserving decentralized training, but the label inconsistency leads to significant divergence in local model parameters across medical centers, thereby hindering the achievement of the global optimum. To resolve this issue, an effective and communication-efficient Federated Learning under Reliable Supervision (FedRS) is proposed, which ensures: i) the local models are trained with reliable supervisory information through the proposed Less-Forgetting and Less-Constraint loss functions, thereby reducing the divergence in local model parameters; and ii) the global model is aggregated based on the consistency of predictions between each local model (after local training) and the global model (received before training), thereby enhancing the reliability of the global model. Extensive experimental results on nine publicly available 3D abdominal CT image datasets show that our FedRS outperforms localized, centralized, and state-of-the-art federated learning methods on both in-federation and out-of-federation datasets, demonstrating its effectiveness and strong generalization capability. In particular, our FedRS only utilizes a model with only 4.1M parameters as its backbone, thereby significantly reducing its communication cost. The source code is publicly available at https://github.com/luohy812/FedRS.
Jie Du 0001, Haoyang Luo, Wenbing Chen, Peng Liu 0070, Tianfu Wang 0001
IEEE Trans. Medical Imaging2
2025 SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
abstract
3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or explicit instruction strictly corresponds to a specific affordance region and are unable to handle long-horizon tasks. Such a paradigm cannot actively reason about complex user intentions that often imply sequential affordances. In this paper, we introduce the Sequential 3D Affordance Reasoning task, which extends the traditional paradigm by reasoning from cumbersome user intentions and then decomposing them into a series of segmentation maps. Toward this, we construct the first instruction-based affordance segmentation benchmark that includes reasoning over both single and sequential affordances, comprising 180K instruction-point cloud pairs. Based on the benchmark, we propose our model, SeqAfford, to unlock the 3D multi-modal large language model with additional affordance segmentation abilities, which ensures reasoning with world knowledge and fine-grained affordance grounding in a cohesive framework. We further introduce a multi-granular language-point integration module to endow 3D dense prediction. Extensive experimental evaluations show that our model excels over well-established methods and exhibits open-world generalization with sequential reasoning abilities. Project page: https://seq-afford.github.io/.
Chunlin Yu, Hanqing Wang 0007, Ye Shi 0001, Haoyang Luo, Sibei Yang, Jingyi Yu 0001, Jingya Wang 0001
CVPR4
2025 SEGA-DCIM: Design Space Exploration-Guided Automatic Digital CIM Compiler with Multiple Precision Support
abstract
Digital computing-in-memory (DCIM) has been a popular solution for addressing the memory wall problem in recent years. However, the DCIM design still heavily relies on manual efforts, and the optimization of DCIM is often based on human experience. These disadvantages limit the time to market while increasing the design difficulty of DCIMs. This work proposes a design space exploration-guided automatic DCIM compiler (SEGA-DCIM) with multiple precision support, including integer and floating-point data precision operations. SEGA-DCIM can automatically generate netlists and layouts of DCIM designs by leveraging a template-based method. With a multi-objective genetic algorithm (MOGA)-based design space explorer, SEGA-DCIM can easily select appropriate DCIM designs for a specific application considering the trade-offs among area, power, and delay. As demonstrated by the experimental results, SEGA-DCIM offers solutions with wide design space, including integer and floating-point precision designs, while maintaining competitive performance compared to state-of-the-art (SOTA) DCIMs.
Haikang Diao, Haoyi Zhang, Haoyang Luo, Yibo Lin, Runsheng Wang, Yuan Wang 0001, Xiyuan Tang
DATE4
2025 Adder-DCIM: A Parallel Bit-Flexible Digital CIM Accelerator Joint Model Compression Framework for AdderNet Inference
abstract
Heavily constrained edge-side tasks necessitate AI chips with low power consumption, low latency, and low cost. In recent years, digital compute-in-memory (CIM) has emerged as a promising solution to enhance energy efficiency and throughput density. However, digital CIM still faces various challenges: significant power and area overhead from multiplication, difficulty in exploiting fine-grained sparsity, and throughput degradation associated with bit-serial architecture. In this work, we propose Adder-DCIM: an efficient parallel bit-flexible DCIM accelerator joint model compression framework for AdderNet inference, in which the key contributions are: 1) a CIM-friendly model compression framework that includes operator decomposition, lossless fine-grained sparsity, and Kullback-Leibler-divergence(KLD)-based Cin-wise mixed-precision quantization; 2) a synchronous parallel DCIM architecture for throughput improvement with mix-precision quantization; 3) a bit-flexible minimal selector circuit for efficient mixed-precision computation. The experimental results demonstrate that under a 28-nm process, the proposed Adder-DCIM achieves a peak energy efficiency of 134 TOPS/W and a peak throughput density of 6.49 TOPS/mm2at INT8. When running ResNet20 on CIFAR10 and ResNet50 on ImageNet, the proposed Adder-DCIM achieves 255 TOPS/[email protected] and 243 TOPS/[email protected] with only a slight decrease in accuracy by 0.34% and 0.7%, respectively. Compared to multiply-based DCIM, Adder-DCIM improves energy efficiency × throughput density metrics by 20.7× for ResNet50 inference.
Haikang Diao, Chuyue Tang, Bocheng Xu, Haoyang Luo, Meng Li 0004, Yuan Wang 0001, Xiyuan Tang
ICCAD4
2025 Beyond One-Hot Labels: Semantic Mixing for Model Calibration
abstract
Model calibration seeks to ensure that models produce confidence scores that accurately reflect the true likelihood of their predictions being correct. However, existing calibration approaches are fundamentally tied to datasets of one-hot labels implicitly assuming full certainty in all the annotations. Such datasets are effective for classification but provides insufficient knowledge of uncertainty for model calibration, necessitating the curation of datasets with numerically rich ground-truth confidence values. However, due to the scarcity of uncertain visual examples, such samples are not easily available as real datasets. In this paper, we introduce calibration-aware data augmentation to create synthetic datasets of diverse samples and their ground-truth uncertainty. Specifically, we present **Calibration-aware Semantic Mixing (CSM)**, a novel framework that generates training samples with mixed class characteristics and annotates them with distinct confidence scores via diffusion models. Based on this framework, we propose calibrated reannotation to tackle the misalignment between the annotated confidence score and the mixing ratio during the diffusion reverse process. Besides, we explore the loss functions that better fit the new data representation paradigm. Experimental results demonstrate that CSM achieves superior calibration compared to the state-of-the-art calibration approaches. Our code is [available here](https://github.com/E-Galois/CSM).
Haoyang Luo, Linwei Tao, Minjing Dong, Chang Xu 0002
ICML1
2025 GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
abstract
Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks like PPO remains underexplored. This gap is particularly significant given the widespread use of large-scale parallel GPU-accelerated simulators, such as IsaacLab, which are optimized for on-policy RL algorithms and enable rapid training of complex robotic tasks. A key challenge lies in computing state-action log-likelihoods under diffusion policies, which is straightforward for Gaussian policies but intractable for flow-based models due to irreversible forward-reverse processes and discretization errors (e.g., Euler-Maruyama approximations). To bridge this gap, we propose GenPO, a generative policy optimization framework that leverages exact diffusion inversion to construct invertible action mappings. GenPO introduces a novel doubled dummy action mechanism that enables invertibility via alternating updates, resolving log-likelihood computation barriers. Furthermore, we also use the action log-likelihood for unbiased entropy and KL divergence estimation, enabling KL-adaptive learning rates and entropy regularization in on-policy updates. Extensive experiments on eight IsaacLab benchmarks, including legged locomotion (Ant, Humanoid, Anymal-D, Unitree H1, Go2), dexterous manipulation (Shadow Hand), aerial control (Quadcopter), and robotic arm tasks (Franka), demonstrate GenPO’s superiority over existing RL baselines. Notably, GenPO is the first method to successfully integrate diffusion policies into on-policy RL, unlocking their potential for large-scale parallelized training and real-world robotic deployment.
Shutong Ding, Haoyang Luo, Weinan Zhang 0001, Jingya Wang 0001, Ye Shi 0001
NeurIPS4
2024 Exploiting Descriptive Completeness Prior for Cross Modal Hashing with Incomplete Labels
abstract
In this paper, we tackle the challenge of generating high-quality hash codes for cross-modal retrieval in the presence of incomplete labels, which creates uncertainty in distinguishing between positive and negative pairs. Vision-language models such as CLIP offer a potential solution by providing generic knowledge for missing label recovery, yet their zero-shot performance remains insufficient. To address this, we propose a novel Prompt Contrastive Recovery approach, \textbf{PCRIL}, which progressively identifies promising positive classes from unknown label sets and recursively searches for other relevant labels. Identifying unknowns is nontrivial due to the fixed and long-tailed patterns of positive label sets in training data, which hampers the discovery of new label combinations. Therefore, we consider each subset of positive labels and construct three types of negative prompts through deletion, addition, and replacement for prompt learning. The augmented supervision guides the model to measure the completeness of label sets, thus facilitating the subsequent greedy tree search for label completion. We also address extreme cases of significant unknown labels and lack of negative pairwise supervision by deriving two augmentation strategies: seeking unknown-complementary samples for mixup and random flipping for negative labels. Extensive experiments reveal the vulnerability of current methods and demonstrate the effectiveness of PCRIL, achieving an average 12\% mAP improvement to the current SOTA across all datasets. Our code is available at https://github.com/E-Galois/PCRIL.
Haoyang Luo, Zheng Zhang 0006, Yadan Luo
NeurIPS1
2024 CASCADE: A Framework for CNN Accelerator Synthesis With Concatenation and Refreshing Dataflow
abstract
Layer Pipeline (LP) represents an innovative architecture for neural network accelerators, which implements task-level pipelining at the granularity of layers. Despite improvements in throughput, LP architectures face challenges due to complicated dataflow design, intricate design space and high resource requirements. In this paper, we introduce an accelerator synthesis framework, CASCADE. CASCADE leverages a novel dataflow, CARD, to efficiently manage convolutional operations’ irregular memory access patterns using simplified logic and minimal buffers. It also employs advanced design space exploration methods to optimize unrolling parallelism and FIFO depth settings automatically for each layer. Finally, to further enhance resource efficiency, CASCADE leverages Lookup Table-based multiplication and accumulation units. With extensive experimental results, we demonstrate that CASCADE significantly outperforms existing works, achieving a$3\times $improvement in resource efficiency and a$4\times $improvement in power efficiency. It achieves over$1.5\times 10^{4}$frames per second throughput and 71.9% accuracy on ImageNet.
Qingyu Guo, Haoyang Luo, Meng Li 0004, Xiyuan Tang, Yuan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Contrastive Incomplete Cross-Modal Hashing
abstract
The success of current deep cross-modal hashing admits a default assumption of thefully-observedcross-modal data. However, such a rigorous common policy is hardly guaranteed for practical large-scale cases, which directly disable the training of prevalent cross-modal retrieval methods with incomplete cross-modal instances and unpaired relations. The main challenges come from the collapsed semantic- and modality-level similarity learning as well as uncertain cross-modal correspondence. In this paper, we propose a Contrastive Incomplete Cross-modal Hashing (CICH) network, which simultaneously determines the cross-modal semantic coordination, unbalanced similarity calibration, and contextual correspondence alignment. Specifically, we design a prototypical semantic similarity coordination module to globally rebuild partially-observed cross-modal similarities under an asymmetric learning scheme. Meanwhile, a semantic-aware contrastive hashing module is established to adaptively perceive and remedy the unbalanced similarities across different modalities with the semantic transition for generating discriminative hash codes. Additionally, a contextual correspondence alignment module is conceived to maximally capture shared knowledge across modalities and eliminate the correspondence uncertainty via a dual contextual information bottleneck formula. To the best of our knowledge, this isthe first successful attemptof enabling contrastive learning to incomplete deep cross-modal hashing. Extensive experiments validate the superiority of our CICH against state-of-the-art methods.
Haoyang Luo, Zheng Zhang 0006, Liqiang Nie
IEEE Trans. Knowl. Data Eng.1
2023 SAGERoute: Synergistic Analog Routing Considering Geometric and Electrical Constraints with Manual Design Compatibility
abstract
Routing is critical to the post-layout performance of analog circuits. As modern analog layouts need to consider both geometric constraints (e.g., design rules and low bending constraints) and electrical constraints (e.g., electromigration (EM), IR drop, symmetry, etc.), it becomes increasingly challenging to investigate the complicated design space. Most previous work has focused only on geometric constraints or basic electrical constraints, lacking holistic and systematic investigation. Such an approach is far from typical manual design practice and can not guarantee post-layout performance on real-world designs. In this work, we propose SAGERoute, a synergistic routing framework taking both geometric and electrical constraints into consideration. Through Steiner tree based wire sizing and guided detailed routing, the framework can generate high-quality routing solutions efficiently under versatile constraints on real-world analog designs.
Haoyi Zhang, Xiaohan Gao, Haoyang Luo, Xiyuan Tang, Junhua Liu 0001, Yibo Lin, Runsheng Wang, Ru Huang 0001
DATE3
2023 Modality-Invariant Asymmetric Networks for Cross-Modal Hashing
abstract
Cross-modal hashing has garnered considerable attention and gained great success in many cross-media similarity search applications due to its prominent computational efficiency and low storage overhead. However, it still remains challenging how to effectively take multilevel advantages of semantics on the entire database to jointly bridge the semantic and heterogeneity gaps across different modalities. In this paper, we propose a novel Modality-Invariant Asymmetric Networks (MIAN) architecture, which explores the asymmetric intra- and inter-modal similarity preservation under a probabilistic modality alignment framework. Specifically, an intra-modal asymmetric network is conceived to capture the query-vs-all internal pairwise similarities for each modality in a probabilistic asymmetric learning manner. Moreover, an inter-modal asymmetric network is deployed to fully harness the cross-modal semantic similarities supported by the maximum inner product search formula between two distinct hash embeddings. Particularly, the pairwise, piecewise and transformed semantics are jointly considered into one unified semantic-preserving hash codes learning scheme. Furthermore, we construct a modality alignment network to distill the redundancy-free visual features and maximize the conditional bottleneck information between different modalities. Such a network could close the heterogeneity and domain shift across different modalities. Extensive experiments evidence that our MIAN approach can outperform the state-of-the-art cross-modal hashing methods.
Zheng Zhang 0006, Haoyang Luo, Lei Zhu 0002, Guangming Lu 0002, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.2
2022 SHREC'22 track: Sketch-based 3D shape retrieval in the wild
Jie Qin 0004, Shuaihang Yuan, Jiaxin Chen 0002, Boulbaba Ben Amor, Yi Fang 0006, Nhat Hoang-Xuan, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Thien-Tri Cao, Nhat-Khang Ngô, Tuan-Luc Huynh, Hai-Dang Nguyen, Minh-Triet Tran, Haoyang Luo, Jianning Wang, Zheng Zhang 0006, Zihao Xin, Yang Wang 0023, Haiqin Chen, Qunying Zhou
Comput. Graph.14
2022 Cognitive multi-modal consistent hashing with flexible semantic transformation
Junfeng An, Haoyang Luo, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002
Inf. Process. Manag.2