EDBT 2026 Demo / reviewers in the wild / expert
Ronghua Luo
dblp:60/7089
· DBLP profile ↗
27ranked-venue papers
2as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OW-DAR: Dual-Granularity Adaptive Reconstruction-Error Modeling for Open-World Object DetectionabstractOpen-world object detection (OWOD) aims to detect known and unknown objects in dynamic environments. However, only known classes are labeled during training, making it challenging for detectors to recognize unknown objects during inference. Existing methods typically rely on supervision from known categories, leading models to overconfidently misclassify visually similar unknowns as known, and dissimilar ones as background. This known-class prior bias limits the model’s ability to detect unknown objects. In this paper, we propose a novel method, OW-DAR, which enhances foreground-background separability through collaborative fine-grained and coarse-grained modeling. At the fine-grained level, we propose Fine-grained Masked Reconstruction (FMR), which randomly masks regions of the feature map to guide the reconstruction toward semantic structures, rather than memorizing low-level patterns. At the coarse-grained level, we propose Adaptive Region-based Error Aggregation (AREA), which operates on object proposals to aggregate reconstruction errors. This enables the model to attend to semantically ambiguous foreground-background boundaries while suppressing the influence of local outliers during optimization. Finally, we leverage robust reconstruction errors to perform unsupervised foreground-background modeling, enabling probabilistic estimation for potential unknown objects. We validate the effectiveness of OW-DAR on standard OWOD benchmark. Experimental results demonstrate that OW-DAR consistently outperforms existing state-of-the-art methods, achieving a +18.8 improvement in unknown object recall (U-Recall). Linhua Ye, Xing Xi, Ronghua Luo |
AAAI | 3 |
| 2026 | SALLM: Open world object detection empowered by self adaptive learning and large model
Yangyang Huang, Linhua Ye, Ronghua Luo |
Expert Syst. Appl. | 3 |
| 2026 | Enhancing open world object detection via learning class-agnostic foundational attributes with large language models
Linhua Ye, Yifei Jiang, Ronghua Luo |
Expert Syst. Appl. | 3 |
| 2025 | OW-OVD: Unified Open World and Open Vocabulary Object DetectionabstractOpen world perception expands traditional closed-set frameworks, which assume a predefined set of known categories, to encompass dynamic real-world environments. Open World Object Detection (OWOD) and Open Vocabulary Object Detection (OVD) are two main research directions, each addressing unique challenges in dynamic environments. However, existing studies often focus on only one of these tasks, leaving the combined challenges of OWOD and OVD largely underexplored. In this paper, we propose a novel detector, OW-OVD, which inherits the zero-shot generalization capability of OVD detectors while incorporating the ability to actively detect unknown objects and progressively optimize performance through incremental learning, as seen in OWOD detectors. To achieve this, we start with a standard OVD detector and adapt it for OWOD tasks. For attribute selection, we propose the Visual Similarity Attribute Selection (VSAS) method, which identifies the most generalizable attributes by computing similarity distributions across annotated and unannotated regions. Additionally, to ensure the diversity of attributes, we incorporate a similarity constraint in the iterative process. Finally, to preserve the standard inference process of OVD, we propose the Hybrid Attribute-Uncertainty Fusion (HAUF) method. This method combines attribute similarity with known class uncertainty to infer the likelihood of an object belonging to an unknown class. We validated the effectiveness of OW-OVD through evaluations on two OWOD benchmarks, M-OWODB and S-OWODB. The results demonstrate that OW-OVD outperforms existing state-of-the-art models, achieving a +15.3 improvement in unknown object recall (U-Recall) and a +15.5 increase in unknown class average precision (U-mAP). Our code is available at: https://github.com/xxyzll/OW_OVD. Xing Xi, Yangyang Huang, Ronghua Luo |
CVPR | 3 |
| 2025 | Decoupled Modeling of Foreground and Background for Open-World Object Detection
Linhua Ye, Xing Xi, Yangyang Huang, Ronghua Luo |
ICIC (2) | 4 |
| 2025 | DLLM: Enhancing Open-World Object Detection with Dynamic Learning and Large ModelsabstractOpen World Object Detection (OWOD) often uses objectness scores as a key metric for identifying unknown objects, which leads to label bias issues. To address this problem, we applied Unsupervised Domain Adaptation (UDA) to OWOD and built an unbiased foreground/background predictor. However, current self-training methods in UDA often rely on fixed thresholds to calculate unsupervised loss, which fails to adapt to the learning difficulties of different categories. Thus, we propose a Dynamic Learning-driven Unsupervised Domain Adaptation (DLDUA) strategy to meet this challenge. Furthermore, we generate initial pseudo-labels for unknown objects using large models to recall more unknown objects. Experimental results show that our method significantly outperforms existing state-of-the-art (SOTA) methods in unknown object recall (improving by 15.2 U-Recall) and leads in inference speed by over 10 FPS compared to based on deformable DETR SOTA methods. Additionally, our method accelerates model convergence. Yangyang Huang, Xing Xi, Ronghua Luo |
ICME | 3 |
| 2025 | OW-VAP: Visual Attribute Parsing for Open World Object DetectionabstractOpen World Object Detection (OWOD) requires the detector to continuously identify and learn new categories. Existing methods rely on the large language model (LLM) to describe the visual attributes of known categories and use these attributes to mark potential objects. The performance of such methods is influenced by the accuracy of LLM descriptions, and selecting appropriate attributes during incremental learning remains a challenge. In this paper, we propose a novel OWOD framework, termed OW-VAP, which operates independently of LLM and requires only minimal object descriptions to detect unknown objects. Specifically, we propose a Visual Attribute Parser (VAP) that parses the attributes of visual regions and assesses object potential based on the similarity between these attributes and the object descriptions. To enable the VAP to recognize objects in unlabeled areas, we exploit potential objects within background regions. Finally, we propose Probabilistic Soft Label Assignment (PSLA) to prevent optimization conflicts from misidentifying background as foreground. Comparative results on the OWOD benchmark demonstrate that our approach surpasses existing state-of-the-art methods with a +13 improvement in U-Recall and a +8 increase in U-AP for unknown detection capabilities. Furthermore, OW-VAP approaches the unknown recall upper limit of the detector. Xing Xi, Weiqiang Wang 0002, Ronghua Luo |
ICML | 4 |
| 2025 | Collaborative Diffusion Models for RecommendationabstractRecently, recommendation models based on collaborative filtering have increasingly leveraged not only primary user-item interactions but also auxiliary information such as implicit relational structures (e.g., user-user or item-item graphs) and multimodal content (e.g., images and textual descriptions) to enhance recommendation performance. A key challenge in this context lies in effectively integrating auxiliary features derived from semantic structures or modality representations into user-item modeling, in a way that enhance performance without incurring detrimental effects. Furthermore, since these features often originate from heterogeneous semantic or modal spaces, they may include redundant or task-irrelevant information that can hinder the learning process. To address these issues, we propose the Collaborative Diffusion Models for Recommendation (CoDMR). CoDMR employs diffusion models in latent feature spaces to filter out task-irrelevant noise embedded in auxiliary features. It introduces task-relevant collaborative signals as conditional guidance during the denoising process, facilitating the generation of auxiliary representations aligned with the recommendation task. These refined features are then incorporated into user-item interaction modeling, resulting in enhanced representations for both users and items. Extensive experiments on three public datasets consistently demonstrate that our CoDMR method outperforms various competitive baselines. The source code of the model implementation is available at the link https://github.com/cmr123456/CoDMR. Mengru Chen, Lianghao Xia, Yong Xu 0007, Ronghua Luo |
SIGIR | 4 |
| 2025 | FMDL: Enhancing Open-World Object Detection with foundation models and dynamic learning
Yangyang Huang, Ronghua Luo |
Expert Syst. Appl. | 3 |
| 2025 | DDMCB: Open-world object detection empowered by Denoising Diffusion Models and Calibration Balance
Yangyang Huang, Xing Xi, Ronghua Luo |
Image Vis. Comput. | 3 |
| 2024 | M2SD: Multiple Mixing Self-Distillation for Few-Shot Class-Incremental LearningabstractFew-shot Class-incremental learning (FSCIL) is a challenging task in machine learning that aims to recognize new classes from a limited number of instances while preserving the ability to classify previously learned classes without retraining the entire model. This presents challenges in updating the model with new classes using limited training data, particularly in balancing acquiring new knowledge while retaining the old. We propose a novel method named Multiple Mxing Self-Distillation (M2SD) during the training phase to address these issues. Specifically, we propose a dual-branch structure that facilitates the expansion of the entire feature space to accommodate new classes. Furthermore, we introduce a feature enhancement component that can pass additional enhanced information back to the base network by self-distillation, resulting in improved classification performance upon adding new classes. After training, we discard both structures, leaving only the primary network to classify new class instances. Extensive experiments demonstrate that our approach achieves superior performance over previous state-of-the-art methods. Jinhao Lin, Weifeng Lin, Jun Huang 0004, Ronghua Luo |
AAAI | 5 |
| 2024 | LVMUM: Toward Open-World Object Detection with Large Vision Models and Unsupervised Modeling
Yangyang Huang, Xing Xi, Weiye Wu, Ronghua Luo |
ICIC (7) | 4 |
| 2024 | End-to-End Object Detection with YOLOF
Xing Xi, Yangyang Huang, Weiye Wu, Ronghua Luo |
ICIC (7) | 4 |
| 2024 | KTCN: Enhancing Open-World Object Detection with Knowledge Transfer and Class-Awareness Neutralization
Xing Xi, Yangyang Huang, Jinhao Lin, Ronghua Luo |
IJCAI | 4 |
| 2024 | Novel Category Discovery Across Domains with Contrastive Learning and Adaptive ClassifierabstractUnsupervised domain adaptation (UDA) has been developed to transfer knowledge from a fully labeled source domain to an unlabeled target domain, mitigating domain shift. The presence of category shift has prompted the research in open-oet domain adaptation (ODA) and universal domain adaptation (UNDA). However, current ODA and UNDA models often treat all novel classes as a unified unknown class or set a priori number for novel classes, which fall short in the novel category discovery in domain adaptation (NCD in DA) problem. Accurately detecting the actual number of novel classes and learning their discriminative features in the DA process are two critical issues underexplored in the problem of NCD in DA. To address these issues, we propose a two-stage framework named Contrastive Adaptive Network (CAN) for the NCD in ODA and UNDA tasks. In Stage 1, we employ two types of contrastive learning for representation learning followed by k-means clustering with various values of k to search for the most probable novel class number. In Stage 2, we propose a known-unknown mapping protocol and group all the classes into several "known-unknown" pairs. Based on them, we design a novel Adaptive Classifier, the size of which is set according to the novel class number predicted in Stage 1. Then we design two tailored loss functions for the training of Adaptive Classifier to learn the discriminative features of novel classes while mitigating the domain shift of common classes, with source and target domain, respectively. Extensive experimental results demonstrate that CAN outperforms the performance of existing SOTA methods on ODA and UNDA problems and showcases significant improvements in the discrimination of novel class features. Implementation is available at https://github.com/yushengyuann/CAN. Shengyuan Yu, Yangyang Huang, Tianwen Yang, Jinhao Lin, Ronghua Luo |
IJCNN | 5 |
| 2024 | UMB: Understanding Model Behavior for Open-World Object DetectionabstractOpen-World Object Detection (OWOD) is a challenging task that requires the detector to identify unlabeled objects and continuously demands the detector to learn new knowledge based on existing ones. Existing methods primarily focus on recalling unknown objects, neglecting to explore the reasons behind them. This paper aims to understand the model's behavior in predicting the unknown category. First, we model the text attribute and the positive sample probability, obtaining their empirical probability, which can be seen as the detector's estimation of the likelihood of the target with certain known attributes being predicted as the foreground. Then, we jointly decide whether the current object should be categorized in the unknown category based on the empirical, the in-distribution, and the out-of-distribution probability. Finally, based on the decision-making process, we can infer the similarity of an unknown object to known classes and identify the attribute with the most significant impact on the decision-making process. This additional information can help us understand the behavior of the model's prediction in the unknown class. The evaluation results on the Real-World Object Detection (RWD) benchmark, which consists of five real-world application datasets, show that we surpassed the previous state-of-the-art (SOTA) with an absolute gain of 5.3 mAP for unknown classes, reaching 20.5 mAP. Our code is available at https://github.com/xxyzll/UMB. Xing Xi, Yangyang Huang, Ronghua Luo |
NeurIPS | 4 |
| 2023 | Data-Free Class-Incremental Learning with Implicit Representation of PrototypesabstractClass-incremental learning (CIL) has attracted much attention in deep learning due to the challenge problem of catastrophic forgetting. Various methods have been proposed for CIL, including exemplar-based class-incremental learning (EBCIL), non-exemplar class-incremental learning (NECIL) and data-free class-incremental learning (DFCIL). Without storing any information (such as examples and prototypes) about the old classes, DFCIL is obviously the most challenging one. To address the problem of lacking information in DFCIL and with the assumption that the learned representations are not linearly separable, we propose a method called IRP. We use the L2-similarity classifier instead of the FC classifier, where each weight vector represents a prototype that implicitly records information about the classes. We use representation-prototype distance minimization (RPDM) to solve the problem of loose representation caused by overfitting. To alleviate the excessive deviation of old prototypes under long-term CIL, we add prototype changing limitation (PCL) and prototype momentum updating (PMU) in incremental stages. In addition, we design a method for resampling around old prototypes (RAOP) to maintain the decision boundary of the old classes. Numerous experiments on three benchmarks have shown that IRP is significantly superior to other DFCIL methods and performs comparably to NECIL and partial EBCIL methods. Tianwen Yang, Leixiong Huang, Ronghua Luo |
ECAI | 3 |
| 2023 | Heterogeneous Graph Contrastive Learning for RecommendationabstractGraph Neural Networks (GNNs) have become powerful tools in modeling graph-structured data in recommender systems. However, real-life recommendation scenarios usually involve heterogeneous relationships (e.g., social-aware user influence, knowledge-aware item dependency) which contains fruitful information to enhance the user preference learning. In this paper, we study the problem of heterogeneous graph-enhanced relational learning for recommendation. Recently, contrastive self-supervised learning has become successful in recommendation. In light of this, we propose a Heterogeneous Graph Contrastive Learning (HGCL), which is able to incorporate heterogeneous relational semantics into the user-item interaction modeling with contrastive learning-enhanced knowledge transfer across different views. However, the influence of heterogeneous side information on interactions may vary by users and items. To move this idea forward, we enhance our heterogeneous graph contrastive learning with meta networks to allow the personalized knowledge transformer with adaptive contrastive augmentation. The experimental results on three real-world datasets demonstrate the superiority of HGCL over state-of-the-art recommendation methods. Through ablation study, key components in HGCL method are validated to benefit the recommendation performance improvement. The source code of the model implementation is available at the link https://github.com/HKUDS/HGCL. Mengru Chen, Chao Huang 0001, Lianghao Xia, Wei Wei 0027, Yong Xu 0007, Ronghua Luo |
WSDM | 6 |
| 2022 | Class Incremental Learning based on Local Structure Constraints in Feature SpaceabstractWhen the training data is provided in the streamed or staged manner, the traditional deep learning methods have severe performance degradation on previous tasks or classes due to the updating of the parameters that have been well learned on old data. Incremental learning aims to avoid the catastrophic forgetting problem which the traditional deep learning framework have suffered from, and enable the model to learn in a continuous way. However, how to make a good trade-off between plasticity and stability, so as to perform well both for the new and the old classes is still an open problem for incremental learning. In this paper, we proposed a new incremental learning method called local structure constrained network(LSC-Net), which employs local relation constraints rather than strong global similarity constraints when updating the model. With the attentively designed inter-class and intra-class restriction, the proposed LSC-Net can make samples of the same class be more compact in the feature space, while expand the distance between samples of different classes. And LSC-Net adopts HM-SoftMax loss function rather than cross-entropy loss function in the classifier, so it can alleviate the effect of data imbalance. Experimental results on CIFAR-100 and ImageNet-100 show that LSC-Net can effectively overcame catastrophic forgetting and obtain competitive results that surpass state-of-the-art methods. Ronghua Luo, Zhenming Huang |
ICPR | 2 |
| 2021 | Compare Network with Task-dependent Embeddings for Few-shot LearningabstractFew-shot Learning, a recent research hotspot in computer vision, attempts to distinguish unseen classes with a few labeled samples. Although significant progress has been made in this field, difficulties still exist in promoting classification accuracy. To tackle this issue, we proposed an effective and novel approach called compare network with task-dependent embeddings for few-shot learning, which includes three modules: embedding module, self-attention module and compare module. With the features extracted by the embedding module, the self-attention module, a use of attention mechanisms as component, can fully capture the internal relevance of features in the context of a certain task so as to learn more discriminative and adaptive task-dependent embeddings by taking a view of the entire task. And the compare module, the direct use of attention mechanisms as metrics, can efficiently evaluate the similarity score between query and discriminative task-dependent embeddings. The effectiveness of our architecture has been verified on two standard classification benchmark, namely the miniImageNet and Omniglot. Experimental results show that our method achieved very competitive results when compared to the state-of-the-art method in terms of few-shot tasks and realized consistent improvements over baselines. Lingwei Chen, Ronghua Luo |
IJCNN | 2 |
| 2021 | Joint offloading and energy optimization for wireless powered mobile edge computing under nonlinear EH Model
Ronghua Luo, Guan Gui 0001 |
Peer-to-Peer Netw. Appl. | 2 |
| 2020 | Salient object detection based on backbone enhanced network
Ronghua Luo, Huailin Huang, WeiZeng Wu |
Image Vis. Comput. | 1 |
| 2017 | A Unified Framework for Metric Transfer LearningabstractTransfer learning has been proven to be effective for the problems where training data from a source domain and test data from a target domain are drawn from different distributions. To reduce the distribution divergence between the source domain and the target domain, many previous studies have been focused on designing and optimizing objective functions with the Euclidean distance to measure dissimilarity between instances. However, in some real-world applications, the Euclidean distance may be inappropriate to capture the intrinsic similarity or dissimilarity between instances. To deal with this issue, in this paper, we propose a metric transfer learning framework (MTLF) to encode metric learning in transfer learning. In MTLF, instance weights are learned and exploited to bridge the distributions of different domains, while Mahalanobis distance is learned simultaneously to maximize the intra-class distances and minimize the inter-class distances for the target domain. Unlike previous work where instance weights and Mahalanobis distance are trained in a pipelined framework that potentially leads to error propagation across different components, MTLF attempts to learn instance weights and a Mahalanobis distance in a parallel framework to make knowledge transfer across domains more effective. Furthermore, we develop general solutions to both classification and regression problems on top of MTLF, respectively. We conduct extensive experiments on several real-world datasets on object recognition, handwriting recognition, and WiFi location to verify the effectiveness of MTLF compared with a number of state-of-the-art methods. Sinno Jialin Pan, Hui Xiong 0001, Qingyao Wu, Ronghua Luo, Huaqing Min, Hengjie Song |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Laplacian regularized locality-constrained coding for image classification
Huaqing Min, Mingjie Liang, Ronghua Luo, Jin-Hui Zhu |
Neurocomputing | 3 |
| 2015 | Simultaneous Recognition and Modeling for Learning 3-D Object Models From Everyday ScenesabstractObject recognition and modeling have classically been studied separately, but practically, they are two closely correlated aspects. In this paper, by exploring the interrelations, we propose a framework to address these two problems at the same time, which we call simultaneous recognition and modeling. Differing from traditional recognition process which consists of off-line object model learning and on-line recognition procedures, our method is solely online. Starting with an empty object database, we incrementally build up object models while at the same time using these models to identify newly observed object views. In the proposed framework, objects are modeled as view graphs and a probabilistic observation model is presented. Both the appearance and the spatial structure of the object are examined, and a formulation based on maximum likelihood estimation is developed. Joint object recognition and modeling are achieved by solving the optimization problem. To evaluate the framework, we have developed a method for simultaneously learning multiple 3-D object models directly from the cluttered indoor environment and tested it using several everyday scenes. Experimental results demonstrate that the framework can cope with the recognition and modeling problem together nicely. Mingjie Liang, Huaqing Min, Ronghua Luo, Jin-Hui Zhu |
IEEE Trans. Cybern. | 3 |
| 2012 | Local Consistent Alignment for 3D modeling with an RGB-D cameraabstract3D model building is an essential topic in both computer vision and robotic community. Usually this is achieved with alignment methods which try to place the data into a common reference frame by finding out the transformations between them. Traditional local alignment approaches, such as ICP and its variants, mainly concern about pairwise constraints and leave out the transformation consistency requirements implied by data sequence. However, these consistencies are important clues to increase the precision of alignment. To use this information better, we present Local Transformation Consistent Alignment (LTCA), which takes triple of frames as input and simultaneously optimize the transformations between them so that local consistencies can be maintained. To evaluate the proposed method, we apply it to the 3D modeling problem with RGB-D image data as input. Comparisons with the state-of-the-art pairwise alignment- Generalized ICP are also made. Experimental results show that more precise alignment results can be obtained when consistency constraints are kept. Mingjie Liang, Huaqing Min, Ronghua Luo, Jin-Hui Zhu, Chang'an Yi |
ICARCV | 3 |
| 2011 | Simultaneous place and object recognition with mobile robot using pose encoded contextual informationabstractPlace and object recognition are two fundamental problems for mobile robot to understand its surroundings. In the field of computer vision it has been acknowledged that context plays an important role in image parsing, but in most of the researches contextual information is only used in one direction and little attention is paid to the relative pose context between objects and local features. We observe, however, place and object can serve as context to each other, that is the recognition of one facilitates the recognition of the other. In this paper, a new hierarchical random field which can encode multiple kinds of context including co-occurrence context, temporal context and relative pose context is proposed for simultaneous place and object recognition with a mobile platform. And a new kind of relative pose context, which is scale and rotation invariant, is defined to improve the stability of pose-encoded context. Experimental results with a mobile robot prove that the proposed method significantly improve the precision of the place and object recognition in familiar and unfamiliar environments. Ronghua Luo, Huaqing Min |
ICRA | 1 |