VLDB 2026 Research / reviewers in the wild / expert
Rufeng Zhang
dblp:248/9047
· DBLP profile ↗
13ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Privacy-Preserving Cooperative Optimal Operation for Reconfigurable Multimicrogrid Distribution Systems: A Noniterative Distributed Optimization ApproachabstractWith the proliferation of distributed generators (DGs), multimicrogrid distribution systems (MMDSs) have emerged, where cooperative optimal operation plays a crucial role in improving operational flexibility and economy. On the other hand, when considering distribution network reconfiguration (DNR), traditional iterative distributed algorithms can protect privacy but face challenges, such as high computational complexity and poor convergence. To address these issues, this article proposes a cooperative operation method for MMDSs based on an equivalent projection (EP) theory. Specifically, a bi-level cooperative method for MMDSs is first proposed. Therein, the flexible resources within the MMDS are identified, while DNR and soft open point (SOP) technologies are considered. Second, to protect privacy of the lower-level microgrids (MGs), the improved Fourier–Motzkin elimination (IFME) and Gaussian elimination methods are used to derive the low-dimensional projection feasible region (PFR) of the lower-level MGs. Third, a noniterative solution method based on the EP theory is developed to efficiently solve the bilevel cooperation model, while addressing the privacy concerns from participants. Numerical results show that the third-order approximate PFR model can achieve a lossless transformation of the bilevel model, effectively balancing computational efficiency and accuracy. Compared to traditional distributed algorithms, the developed noniterative solution method exhibits superior performance in solving mixed-integer programming (MIP) problems. Moreover, the proposed cooperative approach improves the operational performance of the MMDS, achieving a 7.22% reduction in operating costs. Rufeng Zhang, Yuyue Zhang, Yunyang Zou, Tao Jiang 0036, Xue Li 0031 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | 3D Assembly CompletionabstractAutomatic assembly is a promising research topic in 3D computer vision and robotics. Existing works focus on generating assembly (e.g., IKEA furniture) from scratch with a set of parts, namely 3D part assembly. In practice, there are higher demands for the robot to take over and finish an incomplete assembly (e.g., a half-assembled IKEA furniture) with an off-the-shelf toolkit, especially in human-robot and multi-agent collaborations. Compared to 3D part assembly, it is more complicated in nature and remains unexplored yet. The robot must understand the incomplete structure, infer what parts are missing, single out the correct parts from the toolkit and finally, assemble them with appropriate poses to finish the incomplete assembly. Geometrically similar parts in the toolkit can interfere, and this problem will be exacerbated with more missing parts. To tackle this issue, we propose a novel task called 3D assembly completion. Given an incomplete assembly, it aims to find its missing parts from a toolkit and predict the 6-DoF poses to make the assembly complete. To this end, we propose FiT, a framework for Finishing the incomplete 3D assembly with Transformer. We employ the encoder to model the incomplete assembly into memories. Candidate parts interact with memories in a memory-query paradigm for final candidate classification and pose prediction. Bipartite part matching and symmetric transformation consistency are embedded to refine the completion. For reasonable evaluation and further reference, we design two standard toolkits of different difficulty, containing different compositions of candidate parts. We conduct extensive comparisons with several baseline methods and ablation studies, demonstrating the effectiveness of the proposed method. Rufeng Zhang, Mingyu You, Hongjun Zhou, Bin He 0003 |
AAAI | 2 |
| 2023 | Sparse R-CNN: An End-to-End Framework for Object DetectionabstractObject detection serves as one of most fundamental computer vision tasks. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of an image feature map of size H×W. In this paper, we present Sparse R-CNN, a very simple and sparse method for object detection in images. In our method, a fixed sparse set of learned object proposals ( N in total) are provided to the object recognition head to perform classification and localization. By replacing HWk (up to hundreds of thousands) hand-designed object candidates with N (e.g., 100) learnable proposals, Sparse R-CNN makes all efforts related to object candidates design and one-to-many label assignment completely obsolete. More importantly, Sparse R-CNN directly outputs predictions without the non-maximum suppression (NMS) post-processing procedure. Thus, it establishes an end-to-end object detection framework. Sparse R-CNN demonstrates highly competitive accuracy, run-time and training convergence performance with the well-established detector baselines on the challenging COCO dataset and CrowdHuman dataset. We hope that our work can inspire re-thinking the convention of dense prior in object detectors and designing new high-performance detectors. Peize Sun, Rufeng Zhang, Yi Jiang 0009, Tao Kong, Chenfeng Xu, Masayoshi Tomizuka, Zehuan Yuan, Ping Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Self-Supervised Learning by Estimating Twin Class DistributionabstractWe present Twist, a simple and theoretically explainable self-supervised representation learning method by classifying large-scale unlabeled datasets in an end-to-end way. We employ a siamese network terminated by a softmax operation to produce twin class distributions of two augmented images. Without supervision, we enforce the class distributions of different augmentations to be consistent. However, simply minimizing the divergence between augmentations will generate collapsed solutions, i.e., outputting the same class distribution for all images. In this case, little information about the input images is preserved. To solve this problem, we propose to maximize the mutual information between the input image and the output class predictions. Specifically, we minimize the entropy of the distribution for each sample to make the class prediction assertive, and maximize the entropy of the mean distribution to make the predictions of different samples diverse. In this way, Twist can naturally avoid the collapsed solutions without specific designs such as asymmetric network, stop-gradient operation, or momentum encoder. As a result, Twist outperforms previous state-of-the-art methods on a wide range of tasks. Specifically on the semi-supervised classification task, Twist achieves 61.2% top-1 accuracy with 1% ImageNet labels using a ResNet-50 as backbone, surpassing previous best results by an improvement of 6.2%. Codes and pre-trained models are available at https://github.com/bytedance/TWIST. Feng Wang 0034, Tao Kong, Rufeng Zhang, Huaping Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | DenseCL: A simple framework for self-supervised dense visual pre-trainingabstractSelf-supervised learning aims to learn a universal feature representation without labels. To date, most existing self-supervised learning methods are designed and optimized for image classification. These pre-trained models can be sub-optimal for dense prediction tasks due to the discrepancy between image-level prediction and pixel-level prediction. To fill this gap, we aim to design an effective, dense self-supervised learning framework that directly works at the level of pixels (or local features) by taking into account the correspondence between local features. Specifically, we present dense contrastive learning (DenseCL), which implements self-supervised learning by optimizing a pairwise contrastive (dis)similarity loss at the pixel level between two views of input images. Compared to the supervised ImageNet pre-training and other self-supervised learning methods, our self-supervised DenseCL pre-training demonstrates consistently superior performance when transferring to downstream dense prediction tasks including object detection, semantic segmentation and instance segmentation. Specifically, our approach significantly outperforms the strong MoCo-v2 by 2.0% AP on PASCAL VOC object detection, 1.1% AP on COCO object detection, 0.9% AP on COCO instance segmentation, 3.0% mIoU on PASCAL VOC semantic segmentation and 1.8% mIoU on Cityscapes semantic segmentation. The improvements are up to 3.5% AP and 8.8% mIoU over MoCo-v2, and 6.1% AP and 6.1% mIoU over supervised counterpart with frozen-backbone evaluation protocol. Code and models are available at: https://git.io/DenseCL Rufeng Zhang, Chunhua Shen, Tao Kong |
Vis. Informatics | 2 |
| 2022 | SOLO: A Simple Framework for Instance SegmentationabstractCompared to many other dense prediction tasks, e.g., semantic segmentation, it is the arbitrary number of instances that has made instance segmentation much more challenging. In order to predict a mask for each instance, mainstream approaches either follow the "detect-then-segment" strategy (e.g., Mask R-CNN), or predict embedding vectors first then cluster pixels into individual instances. In this paper, we view the task of instance segmentation from a completely new perspective by introducing the notion of "instance categories", which assigns categories to each pixel within an instance according to the instance's location. With this notion, we propose segmenting objects by locations (SOLO), a simple, direct, and fast framework for instance segmentation with strong performance. We derive a few SOLO variants (e.g., Vanilla SOLO, Decoupled SOLO, Dynamic SOLO) following the basic principle. Our method directly maps a raw input image to the desired object categories and instance masks, eliminating the need for the grouping post-processing or the bounding box detection. Our approach achieves state-of-the-art results for instance segmentation in terms of both speed and accuracy, while being considerably simpler than the existing methods. Besides instance segmentation, our method yields state-of-the-art results in object detection (from our mask byproduct) and panoptic segmentation. We further demonstrate the flexibility and high-quality segmentation of SOLO by extending it to perform one-stage instance-level image matting. Code is available at: https://git.io/AdelaiDet. Rufeng Zhang, Chunhua Shen, Tao Kong, Lei Li 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Mask encoding: A general instance mask representation for object segmentation
Rufeng Zhang, Tao Kong, Mingyu You |
Pattern Recognit. | 1 |
| 2022 | Part-Guided Attention Learning for Vehicle Instance RetrievalabstractVehicle instance retrieval (IR) often requires one to recognize the fine-grained visual differences between vehicles. Besides the holistic appearance of vehicles which is easily affected by the viewpoint variation and distortion, vehicle parts also provide crucial cues to differentiate near-identical vehicles. Motivated by these observations, we introduce aPart-Guided Attention Network(PGAN) to pinpoint the prominent part regions and effectively combine the global and local information for discriminative feature learning. PGAN first detects the locations of different part components and salient regions regardless of the vehicle identity, which serves as thebottom-up attentionto narrow down the possible searching regions. To estimate the importance of detected parts, we propose aPart Attention Module(PAM) to adaptively locate the most discriminative regions with high-attention weights and suppress the distraction of irrelevant parts with relatively low weights. The PAM is guided by the identification loss and therefore providestop-down attentionthat enables attention to be calculated at the level of car parts and other salient regions. Finally, we aggregate the global appearance and local features together to improve the feature performance further. The PGAN combines part-guided bottom-up and top-down attention, global and local visual features in an end-to-end framework. Extensive experiments demonstrate that the proposed method achieves new state-of-the-art vehicle IR performance on four large-scale benchmark datasets.1 Xinyu Zhang 0015, Rufeng Zhang, Jiewei Cao, Dong Gong, Mingyu You, Chunhua Shen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsabstractWe present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size H × W. In our method, however, a fixed sparse set of learned object proposals, total length of N, are provided to object recognition head to perform classification and location. By eliminating HWk (up to hundreds of thousands) hand-designed object candidates to N (e.g. 100) learnable proposals, Sparse R-CNN completely avoids all efforts related to object candidates design and many-to-one label assignment. More importantly, final predictions are directly output without non-maximum suppression post-procedure. Sparse R-CNN demonstrates accuracy, run-time and training convergence performance on par with the well-established detector baselines on the challenging COCO dataset, e.g., achieving 45.0 AP in standard 3× training schedule and running at 22 fps using ResNet-50 FPN model. We hope our work could inspire re-thinking the convention of dense prior in object detectors. The code is available at: https://github.com/PeizeSun/SparseR-CNN. Peize Sun, Rufeng Zhang, Yi Jiang 0009, Tao Kong, Chenfeng Xu, Masayoshi Tomizuka, Lei Li 0005, Zehuan Yuan, Changhu Wang, Ping Luo 0002 |
CVPR | 2 |
| 2021 | Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingabstractTo date, most existing self-supervised learning methods are designed and optimized for image classification. These pre-trained models can be sub-optimal for dense prediction tasks due to the discrepancy between image-level prediction and pixel-level prediction. To fill this gap, we aim to design an effective, dense self-supervised learning method that directly works at the level of pixels (or local features) by taking into account the correspondence between local features. We present dense contrastive learning (DenseCL), which implements self-supervised learning by optimizing a pairwise contrastive (dis)similarity loss at the pixel level between two views of input images.Compared to the baseline method MoCo-v2, our method introduces negligible computation overhead (only <1% slower), but demonstrates consistently superior performance when transferring to downstream dense prediction tasks including object detection, semantic segmentation and instance segmentation; and outperforms the state-of-the-art methods by a large margin. Specifically, over the strong MoCo-v2 baseline, our method achieves significant improvements of 2.0% AP on PASCAL VOC object detection, 1.1% AP on COCO object detection, 0.9% AP on COCO instance segmentation, 3.0% mIoU on PASCAL VOC semantic segmentation and 1.8% mIoU on Cityscapes semantic segmentation.Code and models are available at: https://git.io/DenseCL Rufeng Zhang, Chunhua Shen, Tao Kong, Lei Li 0005 |
CVPR | 2 |
| 2021 | Stochastic Optimal Energy Management and Pricing for Load Serving Entity With Aggregated TCLs of Smart Buildings: A Stackelberg Game ApproachabstractWith the development of demand-side management in the smart grid, load-serving entity (LSE) plays a more important role for consumers, which purchases energy from the electricity market and sells it to consumers. Moreover, aggregated thermostatically controlled loads (TCLs) in smart buildings can provide additional demand response capacities and require efficient energy management methods. This article proposes a stochastic optimal energy management and pricing model for LSE with aggregated TCLs and energy storage based on the Stackelberg game and stochastic programming. The energy management and pricing problem are formulated as a bilevel optimization model. The upper level model aims to maximize LSE's expected profit under market price uncertainties and determines the offering prices to consumers. According to the offering price from upper level model, the lower level model optimizes the power purchasing pattern for consumers of two types of buildings with TCLs: factory and office buildings. The nonlinear bilevel model is reformulated and converted into mixed-integer linear programming using a strong duality theory. The proposed model is validated by numerical studies based on real market prices from the PJM electricity market. In addition, the impacts of energy storage, number of buildings, comfortable indoor temperature limits, and offering price limits on LSE's profit are analyzed. Rufeng Zhang, Tao Jiang 0036, Xue Li 0031, Houhe Chen |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Mask Encoding for Single Shot Instance SegmentationabstractTo date, instance segmentation is dominated by two-stage methods, as pioneered by Mask R-CNN. In contrast, one-stage alternatives cannot compete with Mask R-CNN in mask AP, mainly due to the difficulty of compactly representing masks, making the design of one-stage methods very challenging. In this work, we propose a simple single-shot instance segmentation framework, termed mask encoding based instance segmentation (MEInst). Instead of predicting the two-dimensional mask directly, MEInst distills it into a compact and fixed-dimensional representation vector, which allows the instance segmentation task to be incorporated into one-stage bounding-box detectors and results in a simple yet efficient instance segmentation framework. The proposed one-stage MEInst achieves 36.4% in mask AP with single-model (ResNeXt-101-FPN backbone) and single-scale testing on the MS-COCO benchmark. We show that the much simpler and flexible one-stage instance segmentation method, can also achieve competitive performance. This framework can be easily adapted for other instance-level recognition tasks. Code is available at: git.io/AdelaiDet Rufeng Zhang, Zhi Tian, Chunhua Shen, Mingyu You, Youliang Yan |
CVPR | 1 |
| 2020 | SOLOv2: Dynamic and Fast Instance SegmentationabstractIn this work, we design a simple, direct, and fast framework for instance segmentation with strong performance. To this end, we propose a novel and effective approach, termed SOLOv2, following the principle of the SOLO method [32]. First, our new framework is empowered by an efficient and holistic instance mask representation scheme, which dynamically segments each instance in the image, without resorting to bounding box detection. Specifically, the object mask generation is decoupled into a mask kernel prediction and mask feature learning, which are responsible for generating convolution kernels and the feature maps to be convolved with, respectively. Second, SOLOv2 significantly reduces inference overhead with our novel matrix non-maximum suppression (NMS) technique. Our Matrix NMS performs NMS with parallel matrix operations in one shot, and yields better results. We demonstrate that the proposed SOLOv2 achieves the state-of-the- art performance with high efficiency, making it suitable for both mobile and cloud applications. A light-weight version of SOLOv2 executes at 31.3 FPS and yields 37.1% AP on COCO test-dev. Moreover, our state-of-the-art results in object detection (from our mask byproduct) and panoptic segmentation show the potential of SOLOv2 to serve as a new strong baseline for many instance-level recognition tasks. Code is available at https://git.io/AdelaiDet Rufeng Zhang, Tao Kong, Lei Li 0005, Chunhua Shen |
NeurIPS | 2 |