Botao Ren

dblp:158/1362 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0001-1646-9183ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 56% Segmentation and scene understanding · 9% Generative modeling · 7%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
oriented object detection
1.722025
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection · ICLR 2025
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances · CVPR 2025
Computer vision › Image recognition and object detection › object detection › weakly supervised object detection
point-supervised object detection
1.722025
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection · ICLR 2025
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances · CVPR 2025
Machine learning › Generative modeling › diffusion model
3d shape generation
0.912025
ArtFormer: Controllable Generation of Diverse 3D Articulated Objects · CVPR 2025
Computer vision › Image recognition and object detection › object detection
aerial object detection
0.912025
Feedback RoI Features Improve Aerial Object Detection · ICRA 2025
Computer vision › 3D vision › 3d shape modeling
articulated object generation
0.912025
ArtFormer: Controllable Generation of Diverse 3D Articulated Objects · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning
feature extraction
0.912025
Feedback RoI Features Improve Aerial Object Detection · ICRA 2025
Machine learning › Deep learning architectures and training › multi-scale representation
multi-scale feature extraction
0.912025
Feedback RoI Features Improve Aerial Object Detection · ICRA 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Feedback RoI Features Improve Aerial Object Detection · ICRA 2025
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection
0.912025
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances · CVPR 2025
Visual content generation and editing › 3d content generation
controllable 3d generation
0.912025
ArtFormer: Controllable Generation of Diverse 3D Articulated Objects · CVPR 2025
Computer vision › Segmentation and scene understanding
scene graph generation
0.712023
Improving Scene Graph Generation with Superpixel-Based Interaction Learning · ACM Multimedia 2023
Computer vision › Vision and language › visual relationship understanding
visual relation learning
0.712023
Improving Scene Graph Generation with Superpixel-Based Interaction Learning · ACM Multimedia 2023
Computer vision › Image recognition and object detection › object detection
robust object detection
0.312025
Feedback RoI Features Improve Aerial Object Detection · ICRA 2025
Computer vision › Segmentation and scene understanding › image segmentation › region-based segmentation
superpixel segmentation
0.212023
Improving Scene Graph Generation with Superpixel-Based Interaction Learning · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

transformer · 1.7signed distance function · 1.7shape prior · 1.7voronoi watershed loss · 0.9pseudo rotated box generation · 0.9principal component analysis · 0.9gaussian overlap loss · 0.9feedback feature selection · 0.9consistency loss · 0.9class probability map · 0.9
YearPublicationVenuePosition
2025 Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
abstract
With the rapidly increasing demand for oriented object detection (OOD), recent research involving weakly-supervised detectors for learning OOD from point annotations has gained great attention. In this paper, we rethink this challenging task setting with the layout among instances and present Point2RBox-v2. At the core are three principles: 1) Gaussian overlap loss. It learns an upper bound for each instance by treating objects as 2D Gaussian distributions and minimizing their overlap. 2) Voronoi watershed loss. It learns a lower bound for each instance through watershed on Voronoi tessellation. 3) Consistency loss. It learns the size/rotation variation between two output sets with respect to an input image and its augmented view. Supplemented by a few devised techniques, e.g. edge loss and copy-paste, the detector is further enhanced. To our best knowledge, Point2RBox-v2 is the first approach to explore the spatial layout among instances for learning point-supervised OOD. Our solution is elegant and lightweight, yet it is expected to give a competitive performance especially in densely packed scenes: 62.61%/86.15%/34.71% on DOTA/HRSC/FAIR1M.
Yi Yu 0010, Botao Ren, Peiyuan Zhang, Shaofeng Zhang, Feipeng Da, Junchi Yan, Xue Yang 0005
CVPR2
2025 ArtFormer: Controllable Generation of Diverse 3D Articulated Objects
abstract
This paper presents a novel framework for modeling and conditional generation of 3D articulated objects. Troubled by flexibility-quality tradeoffs, existing methods are often limited to using predefined structures or retrieving shapes from static datasets. To address these challenges, we parameterize an articulated object as a tree of tokens and employ a transformer to generate both the object’s high-level geometry code and its kinematic relations. Subsequently, each sub-part’s geometry is further decoded using a signed-distance-function (SDF) shape prior, facilitating the synthesis of high-quality 3D shapes. Our approach enables the generation of diverse objects with high-quality geometry and varying number of parts. Comprehensive experiments on conditional generation from text descriptions demonstrate the effectiveness and flexibility of our method.
Jiayi Su, Youhe Feng, Jinhua Song, Yangfan He, Botao Ren, Botian Xu
CVPR6
2025 PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
abstract
Single point supervised oriented object detection has gained attention and made initial progress within the community. Diverse from those approaches relying on one-shot samples or powerful pretrained models (e.g. SAM), PointOBB has shown promise due to its prior-free feature. In this paper, we propose PointOBB-v2, a simpler, faster, and stronger method to generate pseudo rotated boxes from points without relying on any other prior. Specifically, we first generate a Class Probability Map (CPM) by training the network with non-uniform positive and negative sampling. We show that the CPM is able to learn the approximate object regions and their contours. Then, Principal Component Analysis (PCA) is applied to accurately estimate the orientation and the boundary of objects. By further incorporating a separation mechanism, we resolve the confusion caused by the overlapping on the CPM, enabling its operation in high-density scenarios. Extensive comparisons demonstrate that our method achieves a training speed 15.58$\times$ faster and an accuracy improvement of 11.60\%/25.15\%/21.19\% on the DOTA-v1.0/v1.5/v2.0 datasets compared to the previous state-of-the-art, PointOBB. This significantly advances the cutting edge of single point supervised oriented detection in the modular track. Code and models will be released.
Botao Ren, Xue Yang 0005, Yi Yu 0010, Zhidong Deng
ICLR1
2025 Feedback RoI Features Improve Aerial Object Detection
abstract
Research in visual perception has shown that the human visual system utilizes high-level feedback information to guide lower-level processing, enabling adaptation to signals of varying characteristics. Inspired by this, we propose the Feedback multi-Level feature Extractor (Flex) to dynamically adjust feature selection in object detection based on image-wise and instance-level feedback information. This is particularly beneficial for applications such as aerial object detection, UAV-based target recognition and autonomous vehicle navigation, where global image quality issues like sensor degradation, foggy, or rainy conditions can impact detection performance. Flex adapts to variations in image quality, refining the feature extraction process to improve robustness against these challenges. Experimental results demonstrate that Flex consistently enhances a range of state-of-the-art methods on challenging aerial object detection datasets, including DOTA-v1.0, DOTA-v1.5, and HRSC2016. Furthermore, additional experiments on MS COCO confirm the module's effectiveness in general object detection tasks. Our quantitative and qualitative analyses reveal that the improvements are strongly correlated with image quality, aligning with our original motivation to address global image quality issues in real-world scenarios.
Botao Ren, Botian Xu, Hanwei Gao, Qiankun Yu, Zhidong Deng
ICRA1
2023 Improving Scene Graph Generation with Superpixel-Based Interaction Learning
abstract
Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooked problem in SGG is the coarse-grained interactions between boxes, which inadequately capture contextual semantics for relationship modeling, practically limiting the development of the field. In this paper, we take the initiative to explore and propose a generic paradigm termed Superpixel-based Interaction Learning (SIL) to remedy coarse-grained interactions at the box level. It allows us to model fine-grained interactions at the superpixel level in SGG. Specifically, (i) we treat a scene as a set of points and cluster them into superpixels representing sub-regions of the scene. (ii) We explore intra-entity and cross-entity interactions among the superpixels to enrich fine-grained interactions between entities at an earlier stage. Extensive experiments on two challenging benchmarks (Visual Genome and Open Image V6) prove that our SIL enables fine-grained interaction at the superpixel level above previous box-level methods, and significantly outperforms previous state-of-the-art methods across all metrics. More encouragingly, the proposed method can be applied to boost the performance of existing box-level approaches in a plug-and-play fashion. In particular, SIL brings an average improvement of 2.0% mR (even up to 3.4%) of baselines for the PredCls task on Visual Genome, which facilitates its integration into any existing box-level method.
Can Zhang 0001, Jinfa Huang, Botao Ren, Zhidong Deng
ACM Multimedia4