EDBT 2026 Demo / reviewers in the wild / expert
Shuqi Mei
dblp:80/6566
· DBLP profile ↗
14ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0002-2849-559XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Topo2Seq: Enhanced Topology Reasoning via Topology Sequence LearningabstractExtracting lane topology from perspective views (PV) is crucial for planning and control in autonomous driving. This approach extracts potential drivable trajectories for self-driving vehicles without relying on high-definition (HD) maps. However, the unordered nature and weak long-range perception of the DETR-like framework can result in misaligned segment endpoints and limited topological prediction capabilities. Inspired by the learning of contextual relationships in language models, the connectivity relations in roads can be characterized as explicit topology sequences. In this paper, we introduce Topo2Seq, a novel approach for enhancing topology reasoning via topology sequences learning. The core concept of Topo2Seq is a randomized order prompt-to-sequence learning between lane segment decoder and topology sequence decoder. The dual-decoder branches simultaneously learn the lane topology sequences extracted from the Directed Acyclic Graph (DAG) and the lane graph containing geometric information. Randomized order prompt-to-sequence learning extracts unordered key points from the lane graph predicted by the lane segment decoder, which are then fed into the prompt design of the topology sequence decoder to reconstruct an ordered and complete lane graph. In this way, the lane segment decoder learns powerful long-range perception and accurate topological reasoning from the topology sequence decoder. Notably, topology sequence decoder is only introduced during training and does not affect the inference efficiency. Experimental evaluations on the OpenLane-V2 dataset demonstrate the state-of-the-art performance of Topo2Seq in topology reasoning. Yiming Yang 0001, Yueru Luo, Bingkun He, Erlong Li, Zhipeng Cao 0002, Chao Zheng 0004, Shuqi Mei, Zhen Li 0026 |
AAAI | 7 |
| 2024 | Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic SegmentationabstractThe core of pointly-supervised panoptic segmentation is estimating accurate dense pseudo labels from sparse point labels to train the panoptic head. Previous works generate pseudo labels mainly based on hand-crafted rules, such as connecting multiple points into polygon masks, or assigning the label information of labeled pixels to unlabeled pixels based on the artificially defined traversing distance. The accuracy of pseudo labels is limited by the quality of the hand-crafted rules (polygon masks are rough at object contour regions, and the traversing distance error will result in wrong pseudo labels). To overcome the limitation of hand-crafted rules, we estimate pseudo labels with a fully data-driven pseudo label branch, which is optimized by point labels end-to-end and predicts more accurate pseudo labels than previous methods. We also train an auxiliary semantic branch with point labels, it assists the training of the pseudo label branch by transferring semantic segmentation knowledge through shared parameters. Experiments on Pascal VOC and MS COCO demonstrate that our approach is effective and shows state-of-the-art performance compared with related works. Codes are available at https://github.com/BraveGroup/FDD. Jing Li 0112, Junsong Fan, Yuran Yang, Shuqi Mei, Jun Xiao 0005, Zhaoxiang Zhang 0001 |
AAAI | 4 |
| 2024 | MemoNav: Working Memory Model for Visual NavigationabstractImage-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without considering the goal-relevant fraction. To address this limitation, we present MemoNav, a novel memory model for image-goal navigation, which utilizes a working memory-inspired pipeline to improve navigation performance. Specifically, we employ three types of navigation memory. The node features on a map are stored in the short-term memory (STM), as these features are dynamically updated. A forgetting module then retains the informative STM fraction to increase efficiency. We also introduce long-term memory (LTM) to learn global scene representations by progressively aggregating STM features. Subsequently, a graph attention module encodes the retained STM and the LTM to generate working memory (WM) which contains the scene features essential for efficient navigation. The synergy among these three memory types boosts navigation performance by enabling the agent to learn and leverage goal-relevant scene features within a topological map. Our evaluation on multi-goal tasks demonstrates that MemoNav significantly outperforms previous methods across all difficulty levels in both Gibson and Matterport3D scenes. Qualitative results further illustrate that MemoNav plans more efficient routes. Xu Yang 0004, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2024 | GraphMLLM: A Graph-Based Multi-level Layout Language-Independent Model for Document Understanding
He-Sen Dai, Xiao-Hui Li 0012, Shuqi Mei, Cheng-Lin Liu 0001 |
ICDAR (1) | 5 |
| 2023 | Flexible 3D Lane Detection by Hierarchical Shape MatchingabstractAs one of the basic while vital technologies for HD map construction, 3D lane detection is still an open problem due to varying visual conditions, complex typologies, and strict demands for precision. In this paper, an end-to-end flexible and hierarchical lane detector is proposed to precisely predict 3D lane lines from point clouds. Specifically, we design a hierarchical network predicting flexible representations of lane shapes at different levels, simultaneously collecting global instance semantics and avoiding local errors. In the global scope, we propose to regress parametric curves w.r.t adaptive axes that help to make more robust predictions towards complex scenes, while in the local vision the structure of lane segment is detected in each of the dynamic anchor cells sampled along the global predicted curves. Moreover, corresponding global and local shape matching losses and anchor cell generation strategies are designed. Experiments on two datasets show that we overwhelm current top methods under high precision standards, and full ablation studies also verify each part of our method. Our codes will be released at https://github.com/Doo-do/FHLD. Zhihao Guan, Ruixin Liu, Zejian Yuan, Ao Liu 0010, Erlong Li, Chao Zheng 0004, Shuqi Mei |
AAAI | 9 |
| 2023 | Social Relation Reasoning Based on Triangular ConstraintsabstractSocial networks are essentially in a graph structure where persons act as nodes and the edges connecting nodes denote social relations. The prediction of social relations, therefore, relies on the context in graphs to model the higher-order constraints among relations, which has not been exploited sufficiently by previous works, however. In this paper, we formulate the paradigm of the higher-order constraints in social relations into triangular relational closed-loop structures, i.e., triangular constraints, and further introduce the triangular reasoning graph attention network (TRGAT). Our TRGAT employs the attention mechanism to aggregate features with triangular constraints in the graph, thereby exploiting the higher-order context to reason social relations iteratively. Besides, to acquire better feature representations of persons, we introduce node contrastive learning into relation reasoning. Experimental results show that our method outperforms existing approaches significantly, with higher accuracy and better consistency in generating social relation graphs. Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001 |
AAAI | 6 |
| 2023 | THMA: Tencent HD Map AI System for Creating HD Map AnnotationsabstractNowadays, autonomous vehicle technology is becoming more and more mature. Critical to progress and safety, high-definition (HD) maps, a type of centimeter-level map collected using a laser sensor, provide accurate descriptions of the surrounding environment. The key challenge of HD map production is efficient, high-quality collection and annotation of large-volume datasets. Due to the demand for high quality, HD map production requires significant manual human effort to create annotations, a very time-consuming and costly process for the map industry. In order to reduce manual annotation burdens, many artificial intelligence (AI) algorithms have been developed to pre-label the HD maps. However, there still exists a large gap between AI algorithms and the traditional manual HD map production pipelines in accuracy and robustness. Furthermore, it is also very resource-costly to build large-scale annotated datasets and advanced machine learning algorithms for AI-based HD map automatic labeling systems. In this paper, we introduce the Tencent HD Map AI (THMA) system, an innovative end-to-end, AI-based, active learning HD map labeling system capable of producing and labeling HD maps with a scale of hundreds of thousands of kilometers. In THMA, we train AI models directly from massive HD map datasets via supervised, self-supervised, and weakly supervised learning to achieve high accuracy and efficiency required by downstream users. THMA has been deployed by the Tencent Map team to provide services to downstream companies and users, serving over 1,000 labeling workers and producing more than 30,000 kilometers of HD map data per day at most. More than 90 percent of the HD map data in Tencent Map is labeled automatically by THMA, accelerating the traditional HD map labeling process by more than ten times. Zhipeng Cao 0002, Erlong Li, Ao Liu 0010, Shengtao Zou, Shuqi Mei, Elena Sizikova, Chao Zheng 0004 |
AAAI | 9 |
| 2023 | Visual Traffic Knowledge Graph Generation from Scene ImagesabstractAlthough previous works on traffic scene understanding have achieved great success, most of them stop at a low-level perception stage, such as road segmentation and lane detection, and few concern high-level understanding. In this paper, we present Visual Traffic Knowledge Graph Generation (VTKGG), a new task for in-depth traffic scene understanding that tries to extract multiple kinds of information and integrate them into a knowledge graph. To achieve this goal, we first introduce a large dataset named CASIA-Tencent Road Scene dataset (RS10K) with comprehensive annotations to support related research. Secondly, we propose a novel traffic scene parsing architecture containing a Hierarchical Graph ATtention network (HGAT) to analyze the heterogeneous elements and their complicated relations in traffic scene images. By hierarchizing the heterogeneous graph and equipping it with cross-level links, our approach exploits the correlation among various elements completely and acquires accurate relations. The experimental results show that our method can effectively generate visual traffic knowledge graphs and achieve state-of-the-art performance. The dataset RS10K is available at http://www.nlpr.ia.ac.cn/pal/RS10K.html. Xiao-Hui Li 0012, Shuqi Mei, Cheng-Lin Liu 0001 |
ICCV | 6 |
| 2023 | Informative Data Mining for One-shot Cross-Domain Semantic SegmentationabstractContemporary domain adaptation offers a practical solution for achieving cross-domain transfer of semantic segmentation between labelled source data and unlabeled target data. These solutions have gained significant popularity; however, they require the model to be retrained when the test environment changes. This can result in unbearable costs in certain applications due to the time-consuming training process and concerns regarding data privacy. One-shot domain adaptation methods attempt to overcome these challenges by transferring the pre-trained source model to the target domain using only one target data. Despite this, the referring style transfer module still faces issues with computation cost and over-fitting problems. To address this problem, we propose a novel framework called Informative Data Mining (IDM) that enables efficient one-shot domain adaptation for semantic segmentation. Specifically, IDM provides an uncertainty-based selection criterion to identify the most informative samples, which facilitates quick adaptation and reduces redundant training. We then perform a model adaptation method using these selected samples, which includes patch-wise mixing and prototype-based information maximization to update the model. This approach effectively enhances adaptation and mitigates the overfitting problem. In general, we provide empirical evidence of the effectiveness and efficiency of IDM. Our approach outperforms existing methods and achieves a new state-of-the-art one-shot performance of 56.7%/55.4% on the GTA5/SYNTHIA to Cityscapes adaptation tasks, respectively. The code will be released at https://github.com/yxiwang/IDM. Yuxi Wang 0001, Jian Liang 0001, Jun Xiao 0005, Shuqi Mei, Yuran Yang, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2023 | SSF: Accelerating Training of Spiking Neural Networks with Stabilized Spiking FlowabstractSurrogate gradient (SG) is one of the most effective approaches for training spiking neural networks (SNNs). While assisting SNNs to achieve classification performance comparable to artificial neural networks, SG suffers from the problem of time-consuming training, preventing it from efficient learning. In this paper, we formally analyze the backward process of classic SG and find that the membrane accumulation through time leads to exponential growth of training time. With this discovery, we propose Stabilized Spiking Flow (SSF), a simple yet effective approach to accelerate training of SG-based SNNs. For each spiking neuron, SSF averages its input and output activations over time to yield stabilized input and output, respectively. Then, instead of back propagating all errors that are related to current neuron and inherently entangled in time domain, the auxiliary gradient is directly propagated from the stabilized output to input through a devised relationship mapping. Additionally, SSF method is suitable to different neuron models. Extensive experiments on both static and neuromorphic datasets demonstrate that SNNs trained with SSF approach can achieve performance comparable to the original counterparts, while reducing the training time significantly. In particular, SSF speeds up the training process of state-of-the-art SNN models up to 10× when time steps equal to 80. Zengjie Song, Yuxi Wang 0001, Jun Xiao 0005, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
ICCV | 6 |
| 2023 | Learning to Detect 3D Lanes by Shape Matching and Embeddingabstract3D lane detection based on LiDAR point clouds is a challenging task that requires precise locations, accurate topologies, and distinguishable instances. In this paper, we propose a dual-level shape attention network (DSANet) with two branches for high-precision 3D lane predictions. Specifically, one branch predicts the refined lane segment shapes and the shape embeddings that encode the approximate lane instance shapes, the other branch detects the coarse-grained structures of the lane instances. In the training stage, two-level shape matching loss functions are introduced to jointly optimize the shape parameters of the twobranch outputs, which are simple yet effective for precision enhancement. Furthermore, a shape-guided segments aggregator is proposed to help local lane segments aggregate into complete lane instances, according to the differences of instance shapes predicted at different levels. Experiments conducted on our BEV-3DLanes dataset demonstrate that our method outperforms previous methods. Ruixin Liu, Zhihao Guan, Zejian Yuan, Ao Liu 0010, Tang Kun, Erlong Li, Chao Zheng 0004, Shuqi Mei |
WACV | 9 |
| 2022 | Trajectory Prediction from EGO View: A Coordinate Transform and Tail-Light Event Driven ApproachabstractTrajectory prediction plays an important role in modern au-tonomous driving system. The multi-modal characteristics of trajectory prediction makes it difficult to accurately predict the driving intention and future trajectory of vehicles, espe-cially in the complex conditions, such as lane changing and road intersection. To improve the prediction performance, ad-ditional features are needed. High-definition (HD) map fea-ture and tail light feature are employed in this paper, which are fused with trajectory feature to assist vehicle trajectory pre-diction. A prediction model based on temporal convolution network (TCN) and graph convolution network (GCN) is constructed with corresponding loss functions. Experiments are carried out on our simulation dataset and the results show the effectiveness of the proposed feature fusion method as well as the prediction model. Wenlong Liao, Huanxi Liu, Junchi Yan, Yingxin Lou, Shuqi Mei |
ICME | 8 |
| 2021 | Learning to Understand Traffic SignsabstractOne of the intelligent transportation system's critical tasks is to understand traffic signs and convey traffic information to humans. However, most related works are focused on the detection and recognition of traffic sign texts or symbols, which is not sufficient for understanding. Besides, there has been no public dataset for traffic sign understanding research. Our work takes the first step towards addressing this problem. First, we propose a "CASIA-Tencent Chinese Traffic Sign Understanding Dataset" (CTSU Dataset), which contains 5000 images of traffic signs with rich semantic descriptions. Second, we introduce a novel multi-task learning architecture that extracts text and symbol information from traffic signs, reasons the relationship between texts and symbols, classifies signs into different categories, and finally, composes the descriptions of the signs. Experiments show that the task of traffic sign understanding is achievable, and our architecture demonstrates state-of-the-art and superior performance. The CTSU Dataset is available at http://www.nlpr.ia.ac.cn/databases/CASIA-Tencent%20CTSU/index.html. Wei Feng 0016, Shuqi Mei, Cheng-Lin Liu 0001 |
ACM Multimedia | 5 |
| 2008 | Directional entropy feature for human detectionabstractIn this paper we propose a novel feature, called directional entropy feature (DEF), to improve the performance of human detection under complicated background in images. DEF describe the regularity of region by computing the entropy value of edge pointspsila spatial distribution in specific direction, so DEF has the discriminating power for regular and random pattern. We combine histogram of oriented gradient (HOG) feature with DEF to construct a human detection classifier to test DEFpsilas performance. Experimental results show that DEF can help HOG to decreases false alarms caused by random complicated and rigid shaped background. Long Meng, Shuqi Mei, Weiguo Wu |
ICPR | 3 |