Shuai Bai

dblp:208/8033 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
abstract
Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a comprehensive benchmark to effectively measure the literate capabilities of LMMs. Existing benchmarks are often limited by narrow scenarios and specified tasks. To this end, we introduce CC-OCR, a comprehensive benchmark that possesses a diverse range of scenarios, tasks, and challenges. CC-OCR comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, and released for the first time. We evaluate nine prominent LMMs and reveal both the strengths and weaknesses of these models, particularly in text grounding, multi-orientation, and hallucination of repetition. CC-OCR aims to comprehensively evaluate the capabilities of LMMs on OCR-centered tasks, facilitating continued progress in this crucial area.
Zhibo Yang 0003, Jun Tang 0008, Zhaohai Li, Jianqiang Wan, Humen Zhong, Xuejing Liu, Peng Wang 0028, Shuai Bai, Junyang Lin
ICCV10
2025 GD-NeRF: Generative Detail Compensation for One-shot Generalizable Neural Radiance Fields
abstract
In this article, we focus on the one-shot novel view synthesis task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous One-shot Generalizable Neural Radiance Field (OG-NeRF) methods solve this task in a finetuning-free manner, yet suffer from the blurry issue due to the encoder-only architecture that highly relies on the limited reference image. On the other hand, recent diffusion-based image-to-3D methods show vivid plausible results via distilling pre-trained 2D diffusion models, yet require tedious per-scene optimization. Targeting these issues, we propose GD-NeRF, a generative detail compensation framework that is both capable of producing vivid plausible details and is finetuning-free. Following a coarse-to-fine strategy, it is mainly composed of a One-stage Parallel Pipeline (OPP) and a Diffusion-based 3D-consistent Enhancer (Diff3DE). At the coarse stage, OPP first efficiently integrates the GAN model into the existing OG-NeRF pipeline for injecting primary in-distribution details. Then, at the fine stage, Diff3DE further leverages the pre-trained diffusion models to complement rich out-distribution details while maintaining decent 3D consistency. Extensive experiments on both the synthetic and real-world datasets show that GD-NeRF noticeably improves the vivid details while eliminating the need for per-scene finetuning.
Xiao Pan 0001, Zongxin Yang, Shuai Bai, Yi Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2024 An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Liang Chen 0024, Haozhe Zhao, Tianyu Liu 0001, Shuai Bai, Junyang Lin, Chang Zhou 0005, Baobao Chang
ECCV (81)4
2022 Single Stage Virtual Try-On Via Deformable Attention Flows
Shuai Bai, Huiling Zhou, Hongxia Yang
ECCV (15)1
2022 OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
abstract
In this work, we pursue a unified paradigm for multimodal pretraining to break the shackles of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.
Peng Wang 0028, An Yang, Rui Men, Junyang Lin, Shuai Bai, Chang Zhou 0005, Jingren Zhou 0001, Hongxia Yang
ICML5
2021 Dense Relation Distillation With Context-Aware Aggregation for Few-Shot Object Detection
abstract
Conventional deep learning based methods for object detection require a large amount of bounding box annotations for training, which is expensive to obtain such high quality annotated data. Few-shot object detection, which learns to adapt to novel classes with only a few annotated examples, is very challenging since the fine-grained feature of novel object can be easily overlooked with only a few data available. In this work, aiming to fully exploit features of annotated novel object and capture fine-grained features of query object, we propose Dense Relation Distillation with Context-aware Aggregation (DCNet) to tackle the few-shot detection problem. Built on the meta-learning based framework, Dense Relation Distillation module targets at fully exploiting support features, where support features and query feature are densely matched, covering all spatial locations in a feed-forward fashion. The abundant usage of the guidance information endows model the capability to handle common challenges such as appearance changes and occlusions. Moreover, to better capture scale-aware features, Context-aware Aggregation module adaptively harnesses features from different scales for a more comprehensive feature representation. Extensive experiments illustrate that our proposed approach achieves state-of-the-art results on PASCAL VOC and MS COCO datasets. Code will be made available at https://github.com/hzhupku/DCNet.
Hanzhe Hu, Shuai Bai, Aoxue Li, Jinshi Cui, Liwei Wang 0001
CVPR2
2020 Adaptive Dilated Network With Self-Correction Supervision for Counting
abstract
The counting problem aims to estimate the number of objects in images. Due to large scale variation and labeling deviations, it remains a challenging task. The static density map supervised learning framework is widely used in existing methods, which uses the Gaussian kernel to generate a density map as the learning target and utilizes the Euclidean distance to optimize the model. However, the framework is intolerable to the labeling deviations and can not reflect the scale variation. In this paper, we propose an adaptive dilated convolution and a novel supervised learning framework named self-correction (SC) supervision. In the supervision level, the SC supervision utilizes the outputs of the model to iteratively correct the annotations and employs the SC loss to simultaneously optimize the model from both the whole and the individuals. In the feature level, the proposed adaptive dilated convolution predicts a continuous value as the specific dilation rate for each location, which adapts the scale variation better than a discrete and static dilation rate. Extensive experiments illustrate that our approach has achieved a consistent improvement on four challenging benchmarks. Especially, our approach achieves better performance than the state-of-the-art methods on all benchmark datasets.
Shuai Bai, Zhiqun He, Yu Qiao 0001, Hanzhe Hu, Wei Wu 0021
CVPR1
2020 Class-Wise Dynamic Graph Convolution for Semantic Segmentation
Hanzhe Hu, Deyi Ji, Weihao Gan, Shuai Bai, Wei Wu 0021
ECCV (17)4
2020 Multi-Hierarchical Independent Correlation Filters For Visual Tracking
abstract
For visual object tracking, most of the traditional correlation filters (CF) based methods suffer from the bottleneck of feature redundancy and lack of motion information. In this paper, we design a novel tracking framework, called multi-hierarchical independent correlation filters (MHIT). The framework consists of hierarchical features selection, independent group CF online learning, adaptive multi-branch CF fusion and motion estimation module. Specifically, the multi-hierarchical deep features of CNN representing different semantic information can be fully employed to track multi-scale objects. To fully learn redundant deep features, each hierarchical feature is independently fed into a single branch to implement the online learning of parameters. Finally, an adaptive weight scheme is integrated into the framework to fuse these independent multi-branch CFs for robust visual object tracking. Furthermore, the motion estimation module is introduced to capture motion information, which effectively alleviates the problem of fast motion. Extensive experiments on OTB and VOT datasets show that the proposed MHIT tracker can significantly improve the tracking performance.
Shuai Bai, Zhiqun He, Hongliang Bai
ICME1
2017 A Deep Learning Method to Detect Web Attacks Using a Specially Designed CNN
Ming Zhang 0021, Boyi Xu, Shuai Bai, Shuaibing Lu, Zhechao Lin
ICONIP (5)3