Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhimin Bao

dblp:207/1909 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0005-9876-0423ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 35% Image recognition and object detection · 13% Information extraction and text analysis · 13%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 77% Distributed and cloud data management · 23%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual reasoning
1.722025
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025
V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025
Computer vision › Vision and language
multimodal reasoning
0.912025
V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025
Computer vision › Vision and language
visual question answering
0.912025
V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025
Computer vision › Video understanding and tracking › activity recognition
human activity recognition
0.812024
HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors · AAAI 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.712023
Span-level Aspect-based Sentiment Analysis via Table Filling · ACL (1) 2023
Natural language and speech › Information extraction and text analysis › document AI
table filling
0.712023
Span-level Aspect-based Sentiment Analysis via Table Filling · ACL (1) 2023
Data integration and cleaning › data quality
label error detection
0.712023
ENLD: Efficient Noisy Label Detection for Incremental Datasets in Data Lake · ICDE 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition · AAAI 2022
Machine learning › Representation and self-supervised learning › contrastive learning
hierarchical contrastive learning
0.612022
Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition · AAAI 2022
Computer vision › Image recognition and object detection
image classification
0.612022
NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition · CVPR 2022
Computer vision › Image recognition and object detection
scene text recognition
0.612022
Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition · AAAI 2022
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.612022
NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition · CVPR 2022
Computer vision › Video understanding and tracking
object tracking
0.312017
ReGLe: Spatially Regularized Graph Learning for Visual Tracking · ACM Multimedia 2017
Wearable and physiological sensing › camera-based sensing
event camera
0.212024
HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors · AAAI 2024
Distributed and cloud data management
data lake
0.212023
ENLD: Efficient Noisy Label Detection for Incremental Datasets in Data Lake · ICDE 2023
Computer vision › Image recognition and object detection
object detection
0.212022
NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition · CVPR 2022
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212022
NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition · CVPR 2022

Methods — techniques the papers use, named apart from their topics

large multimodal model · 1.7transformer · 1.5stemnet · 1.5spatial-temporal fusion · 1.5oracle alignment tuning · 0.9data augmentation · 0.9table decoding · 0.7table aggregation · 0.7sentiment consistency regularizer · 0.7representation-based clean sample selection · 0.7pre-trained model confidence · 0.7BERT · 0.7
YearPublicationVenuePosition
2025 V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me
abstract
Oracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMMs) for interpreting OBS. V-Oracle applies principles of pictographic character formation and frames the task as a visual question-answering (VQA) problem, establishing a multi-step reasoning chain. It proposes a multi-dimensional data augmentation for synthesizing high-quality OBS samples, and also implements a multi-phase oracle alignment tuning to improve LMMs’ visual reasoning capabilities. Moreover, to bridge the evaluation gap in the OBS field, we further introduce Oracle-Bench, a comprehensive benchmark that emphasizes process-oriented assessment and incorporates both standard and out-of-distribution setups for realistic evaluation. Extensive experimental results can demonstrate the effectiveness of our method in providing quantitative analyses and superior deciphering capability.
Runqi Qiao, Qiuna Tan, Guanting Dong 0001, MinhuiWu MinhuiWu, Jiapeng Wang 0005, Zhuoma Gongque, Yadong Xue, Zhimin Bao, Lan Yang 0014, Chen Li 0031, Honggang Zhang 0002
ACL (1)12
2025 We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
abstract
Runqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Runqi Qiao, Qiuna Tan, Guanting Dong 0001, Minhui Wu, Xiaoshuai Song, Jiapeng Wang 0005, Zhuoma Gongque, Shanglin Lei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Peiqing Yang 0003, Zhimin Bao, Muxi Diao, Chen Li 0031, Honggang Zhang 0002
ACL (1)17
2025 2-distance coloring of planar graphs without 4, 6-cycles
Yuehua Bu, Zhimin Bao, Hongguo Zhu
Discret. Appl. Math.2
2024 HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors
abstract
The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS.
Xiao Wang 0014, Zongzhen Wu, Bo Jiang 0002, Zhimin Bao, Lin Zhu 0012, Guoqi Li 0002, Yaowei Wang 0001, Yonghong Tian 0001
AAAI4
2023 Span-level Aspect-based Sentiment Analysis via Table Filling
abstract
In this paper, we propose a novel span-level model for Aspect-Based Sentiment Analysis (ABSA), which aims at identifying the sentiment polarity of the given aspect.In contrast to conventional ABSA models that focus on modeling the word-level dependencies between an aspect and its corresponding opinion expressions, in this paper, we propose Table Filling BERT (TF-BERT), which considers the consistency of multi-word opinion expressions at the span-level.Specially, we learn the span representations with a table filling method, by constructing an upper triangular table for each sentiment polarity, of which the elements represent the sentiment intensities of the specific sentiment polarity for all spans in the sentence.Two methods are then proposed, including tabledecoding and table-aggregation, to filter out target spans or aggregate each table for sentiment polarity classification.In addition, we design a sentiment consistency regularizer to guarantee the sentiment consistency of each span for different sentiment polarities.Experimental results on three benchmarks demonstrate the effectiveness of our proposed model.
Mao Zhang 0002, Yongxin Zhu 0003, Zhimin Bao, Xing Sun 0001, Linli Xu 0002
ACL (1)4
2023 ENLD: Efficient Noisy Label Detection for Incremental Datasets in Data Lake
abstract
Due to the difficulty of obtaining high-quality data in real-world scenarios, datasets inevitably contain noisy labeled data, leading to inefficient data usage and poor model performance. Thus, noisy label detection is an important research topic. Previous efforts mainly focus on noisy label detection on specific datasets that have been collected. Some works select clean samples based on relations between representations during the training process; some works utilize confidence outputs of a pre-trained model for noisy label detection. However, how to perform efficient and fine-grained noisy label detection on constantly arriving datasets in a data lake with a large amount of inventory data has not been explored. The rapidly growing volume and changing distribution of data make conventional methods either incur large computation overhead due to repeated training or become increasingly ineffective on newly arriving data. To address these challenges, in this work, we propose a novel approach ENLD to perform efficient and accurate noisy label detection on incremental datasets. Our extensive experiments demonstrate that ENLD outperforms the next best method in both efficiency and accuracy, which achieves 3.65 ×-4.97× detection speedup and higher average f1 scores with various noise rate settings.
Xuanke You, Lan Zhang 0002, Junyang Wang 0004, Zhimin Bao, Shuaishuai Dong
ICDE4
2022 Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition
abstract
We introduce Perceiving Stroke-Semantic Context (PerSec), a new approach to self-supervised representation learning tailored for Scene Text Recognition (STR) task. Considering scene text images carry both visual and semantic properties, we equip our PerSec with dual context perceivers which can contrast and learn latent representations from low-level stroke and high-level semantic contextual spaces simultaneously via hierarchical contrastive learning on unlabeled text image data. Experiments in un- and semi-supervised learning settings on STR benchmarks demonstrate our proposed framework can yield a more robust representation for both CTC-based and attention-based decoders than other contrastive learning methods. To fully investigate the potential of our method, we also collect a dataset of 100 million unlabeled text images, named UTI-100M, covering 5 scenes and 4 languages. By leveraging hundred-million-level unlabeled data, our PerSec shows significant performance improvement when fine-tuning the learned representation on the labeled data. Furthermore, we observe that the representation learned by PerSec presents great generalization, especially under few labeled data scenes.
Hao Liu 0003, Bin Wang 0070, Zhimin Bao, Mobai Xue, Sheng Kang, Deqiang Jiang, Yinsong Liu, Bo Ren 0002
AAAI3
2022 NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition
abstract
Recently, Vision Transformers (ViT), with the self-attention (SA) as the de facto ingredients, have demon-strated great potential in the computer vision community. For the sake of trade-off between efficiency and performance, a group of works merely perform SA operation within local patches, whereas the global contextual information is abandoned, which would be indispensable for visual recognition tasks. To solve the issue, the subsequent global-local ViTs take a stab at marrying local SA with global one in parallel or alternative way in the model. Nevertheless, the exhaustively combined local and global context may exist redundancy for various visual data, and the receptive field within each layer is fixed. Alternatively, a more graceful way is that global and local context can adaptively contribute per se to accommodate different visual data. To achieve this goal, we in this paper propose a novel ViT architecture, termed NomMer, which can dynamically Nominate the synergistic global-local context in vision transforMer. By investigating the working pattern of NomMer, we further explore what context information is focused. Beneficial from this “dynamic nomination” mechanism, without bells and whistles, the NomMer can not only achieve 84.5% Top-1 classification accuracy on ImageNet with only 73M parameters, but also show promising performance on dense prediction tasks, i.e., object detection and semantic segmentation. The code and models are publicly available at https://github.com/TencentYoutuResearch/VisualRecognition-NomMer.
Hao Liu 0003, Xinghua Jiang, Xin Li 0118, Zhimin Bao, Deqiang Jiang, Bo Ren 0002
CVPR4
2018 Moving object detection via robust background modeling with recurring patterns voting
Chenglong Li 0002, Zhimin Bao, Xiao Wang 0014, Jin Tang 0001
Multim. Tools Appl.2
2017 ReGLe: Spatially Regularized Graph Learning for Visual Tracking
abstract
Weighted patch representation of the target object has been proven to be effective for suppressing the background effects in visual tracking. In this paper, we propose a novel approach, called spatially Regularized Graph Learning (ReGLe), to automatically explore the intrinsic relationship among patches both with global and local cues for robust object representation. In particular, the target object bounding box is partitioned into a set of non-overlapping image patches, which are taken as graph nodes, and each of them is associated with a weight to represent how likely it belongs to the target object. To improve the accuracy of node weight computation, we dynamically learn the edge weights (i.e., the appearance compatibility of two nodes) according to both global and local relationship among patches. First, we pursue the low-rank representation for capturing the global low-dimensional subspace structure of patches. Second, we encode the local information into the low-rank representation by exploiting the fact that neighboring nodes usually have similar appearance. Finally, we utilize the representations to learn their affinities (i.e., graph edge weights). The node and edge weights are jointly optimized by a designed ADMM (Alternating Direction Method of Multipliers) algorithm, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is finally predicted by maximizing the classification score in the structured SVM. Extensive experiments demonstrate the effectiveness of the proposed approach on the tracking benchmark datasets: OTB100 and Temple-Color.
Chenglong Li 0002, Xiaohao Wu, Zhimin Bao, Jin Tang 0001
ACM Multimedia3