VLDB 2026 Research / reviewers in the wild / expert
Yuanqi Su
dblp:42/381
· DBLP profile ↗
29ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-7520-7020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible RegionsabstractThis paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the scene completion task into two stages: reconstructing the visible regions (vision) and inferring the invisible regions (hallucination). In the first stage, VisFrontierNet, a visibility-aware projection module, is introduced to accurately trace the visual frontier while preserving fine-grained details. In the second stage, OcclusionMAE, a hallucination network, is employed to generate plausible geometries for the invisible regions using a noise injection mechanism. By decoupling scene completion into these two distinct stages, VisHall3D effectively mitigates feature entanglement and geometric inconsistency, leading to significantly improved reconstruction quality. The effectiveness of VisHall3D is validated through extensive experiments on two challenging benchmarks: SemanticKITTI and SSCBench-KITTI-360. VisHall3D achieves state-of-the-art performance, outperforming previous methods by a significant margin and paves the way for more accurate and reliable scene understanding in autonomous driving and other applications. Haoang Lu, Yuanqi Su, Longjun Gao |
ICCV | 2 |
| 2025 | Edge-Deployable Spatiotemporal Modeling Network for Vehicle Behavior RecognitionabstractVehicle behavior recognition is essential for autonomous vehicles to quickly perceive and respond to their driving environment. In this paper, an edge-deployable spatiotemporal modeling network for vehicle behavior recognition is developed. Firstly, an Efficient SpatioTemporal Modeling (ESTM) Block is designed to extract both long-term evolution features and short-term motion information. Secondly, a Channel-Enhanced Spatial Modeling (CESM) Block is developed to capture the interdependencies among channels in spatial modeling. The proposed network can effectively process video input from onboard cameras while minimizing computational parameters and FLOPs. The combination of the ESTM and CESM blocks can produce rich and effective features for vehicle behavior recognition on edge devices. The experimental results demonstrate the effectiveness of the proposed method in real-world driving scenarios. Gaojie Li, Yaochen Li, Yutong Wang 0008, Sibo Hao, Yuanqi Su |
IV | 6 |
| 2025 | CATIT: Cross-Adaptive Transformer for Road Image TranslationabstractIn the high-resolution image style transfer task of road traffic scenes, how to fully transfer style information while retaining the original content structure information is a challenging problem. In this paper, a novel Transformer-based generative adversarial network for high-resolution un-paired image translation is proposed. Firstly, we design a style transformation module based on cross adaptive Transformer, which dynamically adjusts the content features to achieve statistical alignment between content features and the target style. Meanwhile, an image frequency-domain enhancement module is designed based on cross-attention, which fuses the global information of the low-frequency style with the local details of the high-frequency content information. The detailed texture is then enhanced while the image style consistency is maintained. Furthermore, we design a threshold-guided negative sample screening strategy based on contrastive learning, which can improves the model's transfer effect. The experimental results well demonstrate the effectiveness of the proposed method. Jinhuo Yang, Yaochen Li, Yi Han 0009, Sitong Li, Peijun Chen, Jintao Chang, Yuanqi Su |
IV | 8 |
| 2025 | 3D Shape Adaptation Across Datasets for Weakly Supervised Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is a key yet challenging task that usually involves extensive and expensive manual annotation of 3D boxes. To eliminate the dependence on 3D box labels, weakly supervised M3D (WM3D) has been explored using only 2D annotations, which necessitate the use of extra resources, like LiDAR data, stereo images, and video sequences. However, the strict correspondence and complex calibration between the target image and additional resources limit their applicability. In this work, we propose a simple yet effective framework, 3D Shape Adaptation across datasets for Weakly supervised Monocular 3D Detection (SAWM3D). We observed that directly applying a source-dataset detector to the target dataset results in a significant domain gap, with the primary contribution coming from the 3D location, while orientation and dimensions have a smaller impact. This enables us to view WM3D as 3D shape adaptation optimization on the target dataset. Directly scaling the predicted shape results in a significant reduction of the adaptation gap; fine-tuning on the target dataset using only 2D supervision also yields impressive results. Experiments on the KITTI benchmark demonstrate the effectiveness of our strategies. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 2 |
| 2025 | Rethinking SSIM-Based Optimization in Neural Field TrainingabstractThe Structural Similarity (SSIM) index is a widely used metric for evaluating image quality, with broad applications in areas such as image restoration, 3D reconstruction, and novel view synthesis. A number of previous works have introduced SSIM-based optimization into neural field training to enhance the model's performance. Despite its widespread use, there has been limited research on how to effectively incorporate SSIM loss into the training process. In this work, we explore this gap and provide insights into the role of SSIM loss in neural field training. Our key finding is that SSIM loss is particularly beneficial during the early phase of training, before the model fully learns the luminance information. We show that SSIM loss acts as an effective “guidance” mechanism in the initial training phase, and removing it after the model has learned the luminance does not harm the final performance-in fact, it may improve it. Our experiments demonstrate the effectiveness of our strategy, offering new insights into how SSIM loss can be more efficiently used in neural field training. We believe these findings will not only enhance SSIM's application in neural field training but also inspire further research into more adaptive loss functions for deep learning models. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 2 |
| 2025 | 3D Shape Transfer Learning for Enhanced Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is challenging due to the lack of depth information in the RGB image. Existing works resort to various additional resources to enhance detection performance, including LiDAR data, depth information, CAD models, stereo images, and others, where strict correspondence or synchronization may limit their applicability and scalability. In this work, we propose a simple yet effective framework, 3D Shape Transfer Learning for Enhanced Monocular 3D Object Detection (STLM3D). It views M3D as 3D shape reconstruction and leverages transfer learning across datasets to enhance shape reconstruction capability, thereby enhancing M3D performance. In addition, we design a plug-and-play 3D detection branch for 3D attributes prediction. Experimental results on the KITTI benchmark demonstrate that the proposed method achieves state-of-the-art performance compared to existing approaches. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 2 |
| 2025 | Image Captioning with Multimodal Guidance and Search Space OptimizationabstractImage captioning bridges the gap between visual perception and natural language understanding by transforming image content into descriptive text. While existing methods have made significant progress in visual feature extraction, encoding, and cross-modal semantic alignment, challenges remain in terms of fine-grained feature representation, cross-modal alignment efficiency, and suboptimal search strategies. To address these issues, a multimodal-guided and search space-optimized image captioning model is proposed. In the visual encoding stage, we construct a hierarchical network that integrates regional and grid features through a geometry-constrained multi-layer feature aggregation mechanism, which enhances the model's capability to jointly capture global semantics and local details. In the decoding stage, we introduce a dynamic grouped beam width adjustment strategy to improve semantic path exploration. Additionally, a diversity-driven scoring function is designed to enforce intra-group diversity rewards and inter-group similarity penalties, encouraging the generation of more diverse captions. Finally, we incorporate a two-level pruning algorithm based on syntactic and spatial logic constraints to refine the search space from both hard and soft constraint perspectives, improving both the accuracy and diversity of generated captions. A 3% improvement in CIDEr is achieved by the proposed method over state-of-the-art (SOTA) models, as demonstrated by experiments on the COCO and Flickr30k datasets. Yimou Guo, Yaochen Li, Jingze Liu, Haoyi Lou, Yuanqi Su |
ACM Multimedia | 8 |
| 2025 | UniCuboid: Cuboid-based dense shape supervision for monocular 3D object detection
Yuanqi Su, Haoyue Shi 0002, Haoang Lu, Yuehu Liu, Le Wang 0003 |
Neurocomputing | 2 |
| 2024 | Worst Perception Scenario Search via Recurrent Neural Controller and K-Reciprocal Re-RankingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. A recent theoretical study shows that the generalization on the worst-group of test samples is far more difficult than others. Therefore, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address this, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as the discrete search on the Visual Operation Design Domain (ODD), namely scenario representation, by optimizing LSTM-RNN controller with the worst-performance reward. Moreover, a time-efficient K-reciprocal re-ranking technique is utilized to match the predicted scenario parameters with existing test data. The proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Furthermore, searching performances w.r.t different Visual ODDs are investigated and it is found that visual representations through generative adversarial network contribute to a better performance. Chi Zhang 0020, Xiaoning Ma, Liheng Xu, Haoang Lu, Le Wang 0003, Yuanqi Su, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Point Cloud Segmentation via Edge-fused Local Graph LearningabstractTraditional convolution for capturing local structures and relationships remains a key technical limit in 3D semantic segmentation, which neglects the certain influence of the adjacent points on the central point in the disordered local point clouds. In this paper, we propose a novel joint-edge graph convolution neural network (JEGCN), which can extract the dynamic features of each local area and transfer the edge information between the vertex pairs to the adjacent vertices. In the proposed graph convolution module, the adjacent vertices are selected with high classification confidence which can guide the central vertex, and then reweight these vertices. Considering the lack of texture features in 3D point clouds, we incorporate 2D image features to adjacent feature propagation to effectively extract the local and global features of point clouds. The experimental results based on ScanNet and S3DIS datasets demonstrate the effectiveness of the proposed method. Mengtao Han, Yaochen Li, Liangyu Zuo, Chi Zhang 0020, Yuanqi Su |
ICRA | 6 |
| 2021 | An energy-aware resource deployment algorithm for cloud data centers based on dynamic hybrid machine learning
Bin Liang 0005, Yuanqi Su |
Knowl. Based Syst. | 4 |
| 2020 | A sparse structure for fast circle detection
Yuanqi Su, Bonan Cuan, Yuehu Liu |
Pattern Recognit. | 1 |
| 2018 | Exploiting Vector Fields for Geometric Rectification of Distorted Document Images
Gaofeng Meng, Yuanqi Su, Ying Wu 0001, Shiming Xiang, Chunhong Pan |
ECCV (16) | 2 |
| 2018 | Active target tracking: A simplified view aligning method for binocular camera model
Xinzhao Li, Yuanqi Su, Yuehu Liu, Shaozhuo Zhai, Ying Wu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | A driver fatigue detection method based on multi-sensor signalsabstractFatigue during long-time driving threatens the safety of drivers and transportation. In this paper, we provide an effective method based on multi-sensor signals collected from Kinect2.0 camera and PPG pulse sensor to build a driver fatigue detection system. Unlike most traditional works, we define the transitional process of fatigue and elaborate its effect on training classifiers. The simulation experiments are then designed and 15 groups of data are collected. Our method works in the following steps: 1) feature extraction and fusion, 2) sample labelling and 3) SVM classifier designing. The 10-fold cross-validation accuracy of the classifier is 90.10% and the test accuracy is 83.82%. Experimental results verify that our method to deal with samples in transitional process is universal and more accurate than traditional methods. Moreover, our method based on multi-sensor works better than those dealing with single-sensor. Yuanqi Su, Yuehu Liu, Danchen Zhao |
WACV | 2 |
| 2016 | Iteratively parsing contour fragments for object detection
Yuanqi Su, Yuehu Liu |
Neurocomputing | 2 |
| 2016 | Nonlinear deformation learning for face alignment across expression and pose
Yang Yang 0066, Yuanqi Su, Dongge Cai, Meifeng Xu |
Neurocomputing | 2 |
| 2016 | Fast two-cycle curve evolution with narrow perception of background for object tracking and contour refinement
Yaochen Li, Yuanqi Su, Yuehu Liu |
Signal Process. Image Commun. | 2 |
| 2016 | Three-Dimensional Traffic Scenes Simulation From Road Image SequencesabstractIn this paper, we present a novel framework to allow users to tour simulated traffic scenes from the first-person view. Constructing 3-D scenes from road image sequences is in general difficult, due to the intrinsic complexity of dynamic road scenes, which are composed of a drastically moving background, not to mention numerous other surrounding vehicles. With the definitions of the traffic scene models, we first introduce the construction process of the simple traffic scenes. After the detection of road boundaries by a semantic fast two-cycle (FTC) level set method, we generate the control points on road sides to construct the “floor-wall” background scene that is subsequently propagated to each frame. Furthermore, we approach the cluttered traffic scenes through a three-component processing pipeline as follows: 1) traffic elements segmentation; 2) background images inpainting; and 3) traffic scenes construction. The traffic elements in the cluttered images are segmented by the semantic FTC level set method first. A Gaussian mixture model is then employed to inpaint the occluded background utilizing the optical flows. The cluttered traffic scenes can be constructed after the segmentation and inpainting components. The foreground polygons such as vehicles and traffic signs are then modeled. Users can change their viewpoints according to their own interpretations. We present the evaluations of each technical component, followed by our findings from comprehensive user studies, which well demonstrate the effectiveness of the proposed framework in delivering good touring experience to users. Yaochen Li, Yuehu Liu, Yuanqi Su, Gang Hua 0001, Nanning Zheng 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Contour Guided Hierarchical Model for Shape MatchingabstractFor its simplicity and effectiveness, star model is popular in shape matching. However, it suffers from the loose geometric connections among parts. In the paper, we present a novel algorithm that reconsiders these connections and reduces the global matching to a set of interrelated local matching. For the purpose, we divide the shape template into overlapped parts and model the matching through a part-based layered structure that uses the latent variable to constrain parts' deformation. As for inference, each part is used for localizing candidates by the partial matching. Thanks to the contour fragments, the partial matching can be solved via modified dynamic programming. The overlapped regions among parts of the template are then explored to make the candidates of parts meet at their shared points. The process is fulfilled via a refined procedure based on iterative dynamic programming. Results on ETHZ shape and Inria Horse datasets demonstrate the benefits of the proposed algorithm. Yuanqi Su, Yuehu Liu, Bonan Cuan, Nanning Zheng 0001 |
ICCV | 1 |
| 2015 | Fast Two-Cycle level set tracking with narrow perception of backgroundabstractThe problem of tracking foreground objects in a video sequence with moving background remains challenging. In this paper, we propose the Fast Two-Cycle level set method with Narrow band Background (FTCNB) to automatically extract the foreground objects in such video sequences. The level set curve evolution process consists of two successive cycles: one cycle for data dependent term and a second cycle for smoothness regularization. The curve evolution is implemented by computing the signs of region competition terms on two linked lists of contour pixels rather than solving any Partial Differential Equations (PDEs). Maximum A Posterior (MAP) optimization is applied in the FTCNB method for curve refinement with the assistance of optical flows. The comparison with other level set methods demonstrate the tracking accuracy of our method. The tracking speed of the proposed method also outperforms the traditional level set methods. Yaochen Li, Yuanqi Su, Yuehu Liu |
ICME | 2 |
| 2015 | Edge Propagation KD-Trees: Computing Approximate Nearest Neighbor FieldsabstractPropagation-assisted kd-tree is a state-of-the-art method for computing approximate nearest neighbor (ANN) fields. In this method, each query patch needs descending search in the kd-tree and propagation search in the nearby patches. We observed that the query patches in the edge region need descending search, while other query patches only need propagation search. This can be an opportunity to save plenty of search time. In this letter, we propose edge propagation kd-trees to quickly compute ANNs. Our method can distinguish between edge patches and propagation patches in choosing the proper search. Experiments on public data set VidPairs show that our search method is 2-3 times faster than the propagation-assisted kd-tree search method at nearly the same accuracy. Xiuxiu Bai, Xiaoshe Dong, Yuanqi Su |
IEEE Signal Process. Lett. | 3 |
| 2013 | An iterative parsing approach for contour fragmentsabstractThis paper presents an approach for linking edge points into contour fragments, each of which is an ordered point sequence. In the problem of shape-based object detection, we investigate what characteristics of the contour fragments representation influence the detection performance. It is observed that if interest objects are represented by limited number of contour fragments, with moderate count of noisy points (not belonging to interest objects) brought in, the detection algorithm can reliably obtain interest objects. For this purpose, we propose to iteratively extract a pair of points which preserves the farthest geodesic distance in each connected component of edge map. We conduct experiments on the ETHZ dataset and compare with other typical methods for generating contour fragments. Experimental results demonstrate that the proposed approach outperforms some existing ones in object detection. Yuehu Liu, Yuanqi Su |
ICME | 3 |
| 2013 | A Voting Scheme for Partial Object Extraction under Cluttered EnvironmentabstractShape extraction aims to detect and localize objects via the shape information. The paper presents a novel voting scheme that can extract partially occluded objects under cluttered environments using a single shape. It works by jointly figuring out the boundaries and resolving the geometric configurations. To model the missing part lead by occlusion, we discretize the shape template into a set of its subpart, named portions. Our representation of shape template is through a set of portion together with their interconnections. Instead of forming a fully connected network, our interconnections make the portions consistent with the chain along the boundary of shape template. Based on the representation, we formulate an auto-locked objective function that contains both the unary and pairwise terms and balances the effects of missing parts. Min-sum voting scheme with strategy driven by bottom–up information is then proposed to minimize the objective function. Conducted experiments show that proposed algorithm is promising for shape extraction with occlusion and noisy backgrounds and allows the non-rigid deformations. Yuanqi Su, Yuehu Liu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2011 | A new hybrid method to detect text in natural sceneabstractIn this paper, a new hybrid method for text detection in natural scene is proposed. According the linguistics rules, this algorithm mainly consists of three parts. First, considering both unary property and binary relationship, the conditional random field (CRF) model is introduced for text region detection. Second, connected components (CCs) are extracted by similar stroke width, and filtered coarsely by stroke width analysis. Candidate CCs are then filtered by candidate regions. Finally, text CCs are clustered into words by geometry heuristics. Experiments on the public benchmark ICDAR 2003 dataset show that proposed algorithm can detect text with various font sizes in natural scene. Yuehu Liu, Yuanqi Su |
ICIP | 4 |
| 2011 | Expression transfer for facial sketch animation
Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yuanqi Su, Yoshifumi Nishio |
Signal Process. | 5 |
| 2010 | Optimal Trajectory Space Finding for Nonrigid Structure from Motion
Yuanqi Su, Yuehu Liu, Yang Yang 0066 |
ACIVS (1) | 1 |
| 2010 | A Parameterized Representation for the Cartoon Sample Space
Yuehu Liu, Yuanqi Su, Daitao Jia |
MMM | 2 |
| 2008 | A Method for Deforming-Driven Exaggerated Facial Animation GenerationabstractThis paper proposes a method for automatically generating facial animation with exaggerated features from a facial image. According to facial diversity and identity, an exaggerated face can be determined by the neutral face and the exaggeration effect difference that is represented with the deforming parameters. The proposed method utilizes the central point and the feature vector to represent a facial component and then transforms these feature vectors under the control of deforming parameters to generate the exaggerated facial features. The exaggerated facial animation can be driven to generate by the sequence of the deforming parameters. Experimental results prove that the proposed method has the advantages of simplicity, flexibility and directness, and the generated facial animations are expressive. Ping Wei 0001, Yuehu Liu, Yuanqi Su |
CW | 3 |