VLDB 2026 Research / reviewers in the wild / expert
Jinshui Hu
dblp:308/2417
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0009-0001-3017-973XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QuEPT: Quantized Elastic Precision Transformers with One-Shot Calibration for Multi-Bit SwitchingabstractElastic precision quantization enables multi-bit deployment via a single optimization pass, fitting diverse quantization scenarios. Yet, the high storage and optimization costs associated with the Transformer architecture, research on elastic quantization remains limited, particularly for large language models. This paper proposes QuEPT, an efficient post-training scheme that reconstructs block-wise multi-bit errors with one-shot calibration on a small data slice. It can dynamically adapt to various predefined bit-widths by cascading different low-rank adapters, and supports real-time switching between uniform quantization and mixed precision quantization without repeated optimization. To enhance accuracy and robustness, we introduce Multi-Bit Token Merging (MB-ToMe) to dynamically fuse token features across different bit-widths, improving robustness during bit-width switching. Additionally, we propose Multi-Bit Cascaded Low-Rank adapters (MB-CLoRA) to strengthen correlations between bit-width groups, further improve the overall performance of QuEPT. Extensive experiments demonstrate that QuEPT achieves comparable or better performance to existing state-of-the-art post-training quantization methods. Zhongcheng Li, Jinshui Hu |
AAAI | 5 |
| 2026 | Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentationabstract3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentation methods typically rely on high-dimensional category features, which introduce substantial memory overhead. Moreover, fine-grained segmentation remains challenging due to label space congestion and the lack of stable multi-granularity control mechanisms. To address these limitations, we propose a coarse-to-fine binary encoding scheme for per-Gaussian category representation, which compresses each feature into a single integer via the binary-to-decimal mapping, drastically reducing memory usage. We further design a progressive training strategy that decomposes panoptic segmentation into a series of independent sub-tasks, reducing inter-class conflicts and thereby enhancing fine-grained segmentation capability. Additionally, we fine-tune opacity during segmentation training to address the incompatibility between photometric rendering and semantic segmentation, which often leads to foreground-background confusion. Extensive experiments on multiple benchmarks demonstrate that our method achieves state-of-the-art segmentation performance while significantly reducing memory consumption and accelerating inference. An Yang, Jun Du 0002, Jianqing Gao, Jinshui Hu, Cong Liu 0006 |
AAAI | 6 |
| 2025 | RFL: Simplifying Chemical Structure Recognition with Ring-Free LanguageabstractThe primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current end-to-end methods to learn one-dimensional markup directly. To overcome this limitation, we propose a novel Ring-Free Language (RFL), which utilizes a divide-and-conquer strategy to describe chemical structures in a hierarchical form. RFL allows complex molecular structures to be decomposed into multiple parts, ensuring both uniqueness and conciseness while enhancing readability. This approach significantly reduces the learning difficulty for recognition models. Leveraging RFL, we propose a universal Molecular Skeleton Decoder (MSD), which comprises a skeleton generation module that progressively predicts the molecular skeleton and individual rings, along with a branch classification module for predicting branch information. Experimental results demonstrate that the proposed RFL and MSD can be applied to various mainstream methods, achieving superior performance compared to state-of-the-art approaches in both printed and handwritten scenarios. Qikai Chang, Mingjun Chen, Changpeng Pi, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jinshui Hu |
AAAI | 9 |
| 2025 | Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging ScenariosabstractFace recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale training data, effectively transferring pre-trained face recognition models to these scenarios has become a challenge. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have shown great potential in natural language processing tasks, but their effectiveness in computer vision tasks, especially in face recognition tasks, remains under-explored. In this paper, we propose a Data-Parameter-Efficient Fine-Tuning (DPEFT) approach for the face recognition tasks, encompassing two kinds of fine-tuning strategies. With these strategies, the DPEFT method requires only an additional 2.7% learnable parameters and 20% of the training data during the training phase to achieve competitive results. Moreover, by further integrating the concept of structural re-parameterization, our approach maintains the same model architecture and parameters as the pre-trained model during inference. Extensive experimental results on both holistic and occluded face datasets demonstrate that our approach achieves performance comparable to or better than the fully fine-tuning methods, and significantly lower training costs. Our DPEFT enables the pre-trained face recognition model to adapt efficiently and effectively to a variety of scenarios, indicating its potential in practical applications. Yin Lin, Ziyang Wu, Qidong Huang, Jinshui Hu, Zengfu Wang |
ICASSP | 6 |
| 2025 | Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text RecognitionabstractOnline Handwritten Text Recognition (OLHTR) has gained considerable attention for its diverse range of applications. Current approaches usually treat OLHTR as a sequence recognition task, employing either a single trajectory or image encoder, or multi-stream encoders, combined with a CTC or attention-based recognition decoder. However, these approaches face several drawbacks: 1) single encoders typically focus on either local trajectories or visual regions, lacking the ability to dynamically capture relevant global features in challenging cases; 2) multi-stream encoders, while more comprehensive, suffer from complex structures and increased inference costs. To tackle this, we propose a Collaborative learning-based OLHTR framework, called Col-OLHTR, that learns multimodal features during training while maintaining a single-stream inference process. Col-OLHTR consists of a trajectory encoder, a Point-to-Spatial Alignment (P2SA) module, and an attention-based decoder. The P2SA module is designed to learn image-level spatial features through trajectory-encoded features and 2D rotary position embeddings. During training, an additional image-stream encoder-decoder is collaboratively trained to provide supervision for P2SA features. At inference, the extra streams are discarded, and only the P2SA module is used and merged before the decoder, simplifying the process while preserving high performance. Extensive experimental results on several OLHTR benchmarks demonstrate the state-of-the-art (SOTA) performance, proving the effectiveness and robustness of our design. Jinshui Hu, Jun Du 0002, Qingfeng Liu |
ICASSP | 2 |
| 2025 | Exploring Part-Informed Visual-Language Learning for Person Re-IdentificationabstractRecently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while neglecting supervision on fine-grained part features, thus lacking constraints for local feature semantic consistency. To this end, we propose Part-Informed Visual-language Learning (π-VL) to enhance fine-grained visual features with part-informed language supervisions for ReID tasks. Specifically, π-VL introduces a human parsing-guided prompt tuning strategy and a hierarchical visual-language alignment paradigm to ensure within-part feature semantic consistency. The former combines both identity labels and human parsing maps to constitute pixel-level text prompts, and the latter fuses multi-scale visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our π-VL achieves performance comparable to or better than state-of-the-art methods on four commonly used ReID benchmarks. Notably, it reports 91.0% Rank-1 and 76.9% mAP on the challenging MSMT17 database, without bells and whistles. Yin Lin, Yehansen Chen, Jinshui Hu, Cong Liu 0006, Zengfu Wang |
ICME | 4 |
| 2024 | NAMER: Non-autoregressive Modeling for Handwritten Mathematical Expression Recognition
Jinshui Hu, Mingjun Chen, Cong Liu 0006, Jun Du 0002, Qingfeng Liu |
ECCV (57) | 3 |
| 2024 | ICDAR 2024 Competition on Recognition of Chemical Structures
Mingjun Chen, Hao Wu 0090, Qikai Chang, Hanbo Cheng, Jiefeng Ma, Pengfei Hu 0006, Changpeng Pi, Jinshui Hu, Cong Liu 0006, Jun Du 0002 |
ICDAR (6) | 10 |
| 2024 | 1DFormer: A Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking
Shijie Huan, Shangfei Wang, Jinshui Hu, Cong Liu 0006 |
IJCAI | 4 |
| 2023 | Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object DetectionabstractLiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear. The main challenge arises from that Radar data are extremely sparse and lack height information. Therefore, directly integrating Radar features into LiDAR-centric detection networks is not optimal. In this work, we introduce a bi-directional LiDAR-Radar fusion framework, termed Bi-LRFusion, to tackle the challenges and improve 3D detection for dynamic objects. Technically, Bi-LRFusion involves two steps: first, it enriches Radar's local features by learning important details from the LiDAR branch to alleviate the problems caused by the absence of height information and extreme sparsity; second, it combines LiDAR features with the enhanced Radar features in a unified bird's-eye-view representation. We conduct extensive experiments on nuScenes and ORR datasets, and show that our Bi-LRFusion achieves state-of-the-art performance for detecting dynamic objects. Notably, Radar data in these two datasets have different formats, which demonstrates the generalizability of our method. Codes will be published. Yingjie Wang 0005, Jiajun Deng, Yao Li 0016, Jinshui Hu, Cong Liu 0006, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang |
CVPR | 4 |
| 2023 | Vision-Language Adaptive Mutual Decoder for OOV-STR
Jinshui Hu, Qiandong Yan, Xuyang Zhu, Jiajia Wu 0003, Jun Du 0002, Li-Rong Dai 0001 |
ICIG (2) | 1 |
| 2023 | Handwritten Chemical Structure Image to Structure-Specific Markup Using Random Conditional Guided DecoderabstractSatisfactory recognition performance has been achieved for simple and controllable printed molecular images. However, recognizing handwritten chemical structure images remains unresolved due to the inherent ambiguities in handwritten atoms and bonds, as well as the signifcant challenge of converting projected 2D molecular layouts into markup strings. Target to address these problems, this paper proposes an end-to-end framework for handwritten chemical structure images recognition, with novel structure-specific markup language (SSML) and random conditional guided decoder (RCGD). SSML alleviates ambiguity and complexity in Chemfig syntax by designing an innovative markup language to accurately depict molecular structures. Besides, we propose RCGD to address the issue of multiple path decoding of molecular structures, which is composed of conditional attention guidance, memory classification and path selection mechanisms. In order to fully confirm the effectiveness of the end-to-end method, a new database containing 50,000 handwritten chemical structure images (EDU-CHEMC) has been established. Experimental results demonstrate that compared to traditional SMILES sequences, our SSML can significantly reduces the semantic gap between chemical images and markup strings. It is worth noting that our method can also recognize invalid or non-existent organic molecular structures, making it highly applicable for tasks related to teaching evaluations in the fields of chemistry and biology education. The EDU-CHEMC will be released soon in https://github.com/iFLYTEK-CV/EDU-CHEMC. Jinshui Hu, Hao Wu 0090, Mingjun Chen, Jiajia Wu 0003, Cong Liu 0006, Jun Du 0002, Li-Rong Dai 0001 |
ACM Multimedia | 1 |
| 2023 | OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionabstractBird's-Eye-View (BEV) based 3D visual perception, which formulates a unified space for multi-view representation, has received wide attention in autonomous driving due to its scalability for downstream tasks. However, view transform in transformer-based BEV methods is agnostic of 3D occlusion relationships, resulting in model degradation. To construct a higher-quality BEV space, this paper analyzes the mutual occlusion problems in the view transform process and proposes a new transformer-based method named OccluBEV. OccluBEV alleviates the occlusion issue via point cloud information distillation in both the image and BEV space. Specifically, in the image space, we perform depth estimation for each pixel and utilize it to guide image feature mapping. Further, since predicting depth directly from monocular image is ill-posed, ignoring stereo information such as multi-view and temporal cues, this paper introduces a voxel visibility segmentation task in 3D BEV space. The task explicitly predicts whether each voxel in the 3D BEV grid is occupied or not. In addition, to alleviate the overfitting problem in BEV feature learning under a single task, we design a multi-head learning framework which jointly models multiple strongly-correlated tasks in a unified BEV space. The effectiveness of the proposed method is fully validated on the nuScenes dataset, achieving a competetive NDS/mAP score of 57.5/47.9 on the nuScenes test leaderboard using ResNet101 backbone, which is superior to state-of-the-art camera-based solutions. Ziteng Wen, Jinshui Hu, Xuming He 0001, Fengren Wang, Shun Lou, Haibo Fan |
ACM Multimedia | 5 |
| 2022 | Structural String Decoder for Handwritten Mathematical Expression RecognitionabstractRecently, recognition of handwritten mathematical expression has been greatly improved by employing sequence modeling methods such as encoder-decoder based methods. Existing encoder-decoder models use string decoders or tree decoders to generate markup of mathematical expression recognition. String decoders directly generate LaTeX strings and tree decoders decode expressions into tree structures. The generalization of string decoders is poor on mathematical expressions with complex hierarchical structures, but its language model is better. Tree decoders can deal with the complex hierarchical structures, but its language model is weakened. In order to take advantage of the above two decoders, we propose a novel structural string decoder (SSD) which not only has good generalization but also can make good use of language model. We demonstrate how the proposed SSD outperforms state-of-the-art string decoders and tree decoders through a set of experiments on CROHME database, which is currently the largest benchmark for online handwritten mathematical expression recognition. Jiajia Wu 0003, Jinshui Hu, Mingjun Chen, Li-Rong Dai 0001, Xuejing Niu |
ICPR | 2 |
| 2022 | A multimodal attention fusion network with a dynamic vocabulary for TextVQA
Jiajia Wu 0003, Jun Du 0002, Fengren Wang, Xinzhe Jiang, Jinshui Hu, Jianshu Zhang 0001, Li-Rong Dai 0001 |
Pattern Recognit. | 6 |