VLDB 2026 Research / reviewers in the wild / expert
Yingying Zhu 0005
dblp:40/5552-5
· DBLP profile ↗
20ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-4664-4659ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric ReasoningabstractDeng Linger, Linghao Zhu, Yuliang Liu, Yu Wang, Qunyi Xie, Jingjing Wu, Gang Zhang, Yingying Zhu, Xiang Bai. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Linger Deng, Linghao Zhu, Qunyi Xie, Yingying Zhu 0005, Xiang Bai |
EMNLP | 8 |
| 2025 | Towards Comprehensive Lecture Slides Understanding: Large-Scale Dataset and Effective Method
Enming Zhang, Yingying Zhu 0005, Xiang Bai |
ICCV | 4 |
| 2025 | Hybrid Transformer-Mamba Model for 3D Semantic SegmentationabstractTransformer-based methods have demonstrated remarkable capabilities in 3D semantic segmentation through their powerful attention mechanisms, but the quadratic complexity limits their modeling of long-range dependencies in large-scale point clouds. While recent Mamba-based approaches offer efficient processing with linear complexity, they struggle with feature representation when extracting 3D features. However, effectively combining these complementary strengths remains an open challenge in this field. In this paper, we propose HybridTM, the first hybrid architecture that integrates Transformer and Mamba for 3D semantic segmentation. In addition, we propose the Inner Layer Hybrid Strategy, which combines attention and Mamba at a finer granularity, enabling simultaneous capture of long-range dependencies and fine-grained local features. Extensive experiments demonstrate the effectiveness and generalization of our HybridTM on diverse indoor and outdoor datasets. Furthermore, our HybridTM achieves state-of-the-art performance on ScanNet, ScanNet200, and nuScenes benchmarks. The code will be made available at https://github.com/deepinact/HybridTM. Xinyu Wang 0024, Jinghua Hou, Zhe Liu 0033, Yingying Zhu 0005 |
IROS | 4 |
| 2025 | Enhancing scene text detectors with realistic text image synthesis using diffusion models
Yingying Zhu 0005, Xiang Bai |
Comput. Vis. Image Underst. | 3 |
| 2025 | AVS-Net: Point sampling with adaptive voxel size for 3D scene understanding
Hongcheng Yang, Dingkang Liang, Dingyuan Zhang, Zhe Liu 0033, Zhikang Zou, Xingyu Jiang 0005, Yingying Zhu 0005 |
Neurocomputing | 7 |
| 2025 | Generative compositor for few-shot visual information extractionabstractVisual Information Extraction (VIE), aiming at extracting structured information from visually rich document images , plays a pivotal role in document processing. Considering various layouts, semantic scopes, and languages, VIE encompasses an extensive range of types, potentially numbering in the thousands. However, many of these types suffer from a lack of training data , which poses significant challenges. In this paper, we propose a novel generative model , named Generative Compositor, to address the challenge of few-shot VIE. The Generative Compositor is a hybrid pointer-generator network that emulates the operations of a compositor by retrieving words from the source text and assembling them based on the provided prompts. Furthermore, three pre-training strategies are employed to enhance the model’s perception of spatial context information. Besides, a prompt-aware resampler is specially designed to enable efficient matching by leveraging the entity-semantic prior contained in prompts. The introduction of the prompt-based retrieval mechanism and the pre-training strategies enable the model to acquire more effective spatial and semantic clues with limited training samples . Experiments demonstrate that the proposed method achieves highly competitive results in the full-sample training, while notably outperforms the baseline in the 1-shot, 5-shot, and 10-shot settings. Zhibo Yang 0003, Wei Hua 0005, Sibo Song, Cong Yao, Yingying Zhu 0005, Wenqing Cheng, Xiang Bai |
Pattern Recognit. | 5 |
| 2025 | Layerlink: Bridging remote sensing object detection and large vision models with efficient fine-tuning
Xingkui Zhu, Dingkang Liang, Xingyu Jiang 0005, Yiran Guan, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. | 6 |
| 2024 | ClipComb: Global-Local Composition Network based on CLIP for Composed Image RetrievalabstractIn this paper, we focus on the task of composed image retrieval. To develop a comprehensive understanding of image and text, we propose a novel global-local composition network (ClipComb) based on the vision-language pretraining CLIP model. The two main phases of ClipComb are the fine-tuning training stage and the composition training stage. First, in the fine-tuning training step, we fine-tune the text and image encoders of the CLIP model to transfer CLIP to this task. We also perform contrastive loss alignment on both global and local features. Second, during the composition training phase, we devise a global-local composition module (GLC) based on the fine-tuned composition learning framework. The GLC module make full use of the CLIP’s pretraining knowledge to generate a composite representation aligned with the target representation. Extensive experimental results demonstrate that our method achieves state-of-the-art performance on two benchmark datasets. Yingying Zhu 0005, Dafeng Li |
ICME | 1 |
| 2024 | LATFormer: Locality-Aware Point-View Fusion Transformer for 3D shape recognition
Xinwei He 0001, Silin Cheng 0001, Dingkang Liang, Song Bai 0001, Xi Wang 0044, Yingying Zhu 0005 |
Pattern Recognit. | 6 |
| 2024 | Class-Aware Mask-guided feature refinement for scene text recognition
Minghui Liao, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. | 4 |
| 2024 | Sequential visual and semantic consistency for semi-supervised text recognition
Minghui Liao, Yingying Zhu 0005, Xiang Bai |
Pattern Recognit. Lett. | 4 |
| 2023 | Focal Inverse Distance Transform Maps for Crowd LocalizationabstractIn this paper, we focus on the crowd localization task, a crucial topic of crowd analysis. Most regression-based methods utilize convolution neural networks (CNN) to regress a density map, which can not accurately locate the instance in the extremely dense scene, attributed to two crucial reasons: 1) the density map consists of a series of blurry Gaussian blobs, 2) severe overlaps exist in the dense region of the density map. To tackle this issue, we propose a novel Focal Inverse Distance Transform (FIDT) map for the crowd localization task. Compared with the density maps, the FIDT maps accurately describe the persons' locations without overlapping in dense regions. Based on the FIDT maps, a Local-Maxima-Detection-Strategy (LMDS) is derived to effectively extract the center point for each individual. Furthermore, we introduce an Independent SSIM (I-SSIM) loss to make the model tend to learn the local structural information, better recognizing local maxima. Extensive experiments demonstrate that the proposed method reports state-of-the-art localization performance on six crowd datasets and one vehicle dataset. Additionally, we find that the proposed method shows superior robustness on the negative and extremely dense scenes, which further verifies the effectiveness of the FIDT maps. Dingkang Liang, Wei Xu 0037, Yingying Zhu 0005, Yu Zhou 0016 |
IEEE Trans. Multim. | 3 |
| 2022 | TNDP: Tensor-Based Network Distance Prediction With Confidence IntervalsabstractThe knowledge of network distances, in the form of delay or latency, for example, is beneficial to a number of distributed applications. Notice that it is difficult and expensive to implement global network measurements to obtain network distance, a feasible idea is to predict unknown distances by introducing network coordinates with limited network measurements. The existing solutions always represent the unknown network distances in a rather unique number. However, research and applications indicate that the real network distances are hard to be accurately figured out and changes subtly in an interval over time with the dynamic network environments. Accordingly, this article proposes a tensor-based network distance prediction (TNDP) approach to represent network distance with confidence intervals, by exploiting the random distance tensor and distributed matrix factorization. With a small set of network measurements among the nodes selected randomly, a distance matrix tensor has been established and factorized into the product of two location matrixes with the adaptive SGD-based learning solution. By introducing the important training determinants, including weight matrix, regularization coefficient, and minibatch gradient descent with the exponential decay rates, the unknown distances among nodes can be accurately inferred in the forms of confidence intervals, with quick convergence and less overfitting. Extensive experimental simulations on a wide variety of available data sets demonstrate that TNDP is superior to other approaches in terms of accuracy for network distance prediction. Haojun Huang, Geyong Min, Wang Miao, Yingying Zhu 0005, Yangming Zhao |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Scalable Orchestration of Service Function Chains in NFV-Enabled Networks: A Federated Reinforcement Learning ApproachabstractNetwork function virtualization (NFV) is critical to the scalability and flexibility of various network services in the form of service function chains (SFCs), which refer to a set of Virtual Network Functions (VNFs) chained in a specific order. However, the NFV performance is hard to fulfill the ever-increasing requirements of network services mainly due to the static orchestrations of SFCs. To tackle this issue, a novel Scalable SFC Orchestration (SSCO) scheme is proposed in this paper for NFV-enabled networks via federated reinforcement learning. SSCO has three remarkable characteristics distinguishing from the previous work: (1) A federated-learning-based framework is designed to train a global learning model, with time-variant local model explorations, for scalable SFC orchestration, while avoiding data sharing among stakeholders; (2) SSCO allows for parameter update among local clients and the cloud server just at the first and last epochs of each episode to ensure that distributed clients can make model optimization at a low communication cost; (3) SSCO introduces an efficient deep reinforcement learning (DRL) approach, with the local learning knowledge of available resources and instantiation cost, to map VNFs into networks flexibly. Furthermore, a loss-weight-based mechanism is proposed to generate and exploit reference samples in replay buffers for future training, avoiding the strong relevance of samples. Simulation results obtained from different working scenarios demonstrate that SSCO can significantly reduce placement errors and improve resource utilization ratio to place time-variant VNFs compared with the state-of-the-art mechanisms. Furthermore, the results show that the proposed approach can achieve desirable scalability. Haojun Huang, Yangming Zhao, Geyong Min, Yingying Zhu 0005, Wang Miao, Jia Hu 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2018 | Cascaded Segmentation-Detection Networks for Text-Based Traffic Sign DetectionabstractIn this paper, we propose a novel text-based traffic sign detection framework with two deep learning components. More precisely, we apply a fully convolutional network to segment candidate traffic sign areas providing candidate regions of interest (RoI), followed by a fast neural network to detect texts on the extracted RoI. The proposed method makes full use of the characteristics of traffic signs to improve the efficiency and accuracy of text detection. On one hand, the proposed two-stage detection method reduces the search area of text detection and removes texts outside traffic signs. On the other hand, it solves the problem of multi-scales for the text detection part to a large extent. Extensive experimental results show that the proposed method achieves the state-of-the-art results on the publicly available traffic sign data set: Traffic Guide Panel data set. In addition, we collect a data set of text-based traffic signs including Chinese and English traffic signs. Our method also performs well on this data set, which demonstrates that the proposed method is general in detecting traffic signs of different languages. Yingying Zhu 0005, Minghui Liao, Wenyu Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Scene text detection and recognition: recent advances and future trends
Yingying Zhu 0005, Cong Yao, Xiang Bai |
Frontiers Comput. Sci. | 1 |
| 2016 | Traffic sign detection and recognition using fully convolutional network guided proposals
Yingying Zhu 0005, Chengquan Zhang, Duoyou Zhou, Xinggang Wang, Xiang Bai, Wenyu Liu 0001 |
Neurocomputing | 1 |
| 2013 | Learning context sensitive similarity measure on pair fusion graphabstractIn this paper, we present a new approach for shape/image retrieval by efficiently fusing different shape similarities, called Pair-Graph Diffusion. Different from other algorithms which linearly integrate different similarity measures, our algorithm adopts Tensor Product Graph(TPG) to combine two shape similarities by fusing two single-graphs into a multi-graph for fusion process. In such way, we gain more shape information in a higher order, and the multigraph is able to better reveal the intrinsic relation between shapes especially when the two input similarities are very complementary. We perform the experiments on two popular image datasets: MPEG-7 shape dataset and Nistér and Stewénius (N-S) dataset, and achieve state-of-arts retrieval rates: 98.87% on MPEG-7 dataset and 3.69 on N-S dataset. The results demonstrate that the proposed method can effectively fuse two similarities. In addition, Multi-graph Diffusion is a general similarity learning algorithm, and it can be easily applied other tasks for ranking/retrieval. Cheng Wang 0048, Yingying Zhu 0005, Wenyu Liu 0001 |
ICIP | 3 |
| 2013 | Traffic sign classification using two-layer image representationabstractThis paper makes use of locality-constrained linear coding (LLC) in a two-layer image representation framework for traffic sign recognition. As a multi-category classification problem with unbalanced frequencies and variations, many machine learning approaches have been adopted with some low level features for traffic sign recognition. To the best of our knowledge, this is the first method using coding features for traffic sign recognition. First, we extract features(dense SIFT features, HOG features and LBP features) and encode them with a k-means generated codebook and LLC. Second, each traffic sign image is represented by the features generated by spatial pyramid matching (SPM). Then, all the image representations from each kind of features are concatenated together as the final image representation. Finally, we show that a linear SVM classifier trained with this image representation can achieve the state-of-the-art recognition rate of 99.67% on the well-known German Traffic Sign Recognition Benchmark. Yingying Zhu 0005, Xinggang Wang, Cong Yao, Xiang Bai |
ICIP | 1 |
| 2011 | Image labeling by multiple segmentationabstractIn this paper, we provide a method for image labeling by combining the local features and contextual cues in a multiple segmentation framework. Our main insight is to weight the classification results of each image region in different levels, which are obtained by a series of learned discriminative models based on bag of features. The contextual cues are implicitly embedded as feature selection in learning process. Multiple segmentation framework provides robust representation, allowing a wide variety of cues to contribute to the confidence in each semantic label. Our algorithm has been applied on the lotus hill institute(LHI) 15-class dataset and outperforms other state-of-the-art methods. Quan Zhou 0004, Canxiang Yan, Yingying Zhu 0005, Xiang Bai, Wenyu Liu 0001 |
ICIP | 3 |