VLDB 2026 Research / reviewers in the wild / expert
Lulu Tang
dblp:231/8652
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | You See it, You Got it: Learning 3D Creation on Pose-Free Videos at ScaleabstractRecent 3D generation models typically rely on limited-scale 3D ‘gold-labels’ or 2D diffusion priors for 3D content creation. However, their performance is upper-bounded by constrained 3D priors due to the lack of scalable learning paradigms. In this work, we present See3D, a visual-conditional multi-view diffusion model trained on large-scale Internet videos for open-world 3D creation. The model aims to Get 3D knowledge by solely Seeing the visual contents from the vast and rapidly growing video data — You See it, You Got it. To achieve this, we first scale up the training data using a proposed data curation pipeline that automatically filters out multi-view inconsistencies and insufficient observations from source videos. This results in a high-quality, richly diverse, large-scale dataset of multi-view images, termed WebVi3D, containing 320M frames from 16M video clips. Nevertheless, learning generic 3D priors from videos without explicit 3D geometry or camera pose annotations is nontrivial, and annotating poses for web-scale videos is prohibitively expensive. To eliminate the need for pose conditions, we introduce an innovative visual-condition - a purely 2D-inductive visual signal generated by adding time-dependent noise to the masked video data. Finally, we introduce a novel visual-conditional 3D generation framework by integrating See3D into a warping-based pipeline for high-fidelity 3D generation. Our numerical and visual comparisons on single and sparse reconstruction benchmarks show that See3D, trained on cost- effective and scalable video data, achieves notable zero-shot and open-world generation capabilities, markedly outperforming models trained on costly and constrained 3D datasets. Additionally, our model naturally supports other image-conditioned 3D creation tasks, such as 3D editing, without further fine-tuning. Please refer to our project page at: https://vision.baai.ac.cn/see3d. Baorui Ma, Huachen Gao, Haoge Deng, Tiejun Huang 0001, Lulu Tang |
CVPR | 6 |
| 2025 | PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code ContextualizationabstractMultimodal Large Language Models (MLLMs), which integrate vision and other modalities into Large Language Models (LLMs), significantly enhance AI capabilities but also introduce new security vulnerabilities. By exploiting the vulnerabilities of the visual modality and the long-tail distribution characteristic of code training data, we present PiCo, a novel jailbreaking framework designed to progressively bypass multi-tiered defense mechanisms in advanced MLLMs. PiCo employs a tier-by-tier jailbreak strategy, using token-level typographic attacks to evade input filtering and embedding harmful intent within programming context instructions to bypass runtime monitoring. To comprehensively assess the impact of attacks, a new evaluation metric is further proposed to assess both the toxicity and helpfulness of model outputs post-attack. By embedding harmful intent within code-style visual instructions, PiCo achieves an average Attack Success Rate (ASR) of 84.13% on Gemini-Pro Vision and 52.66% on GPT-4, surpassing previous methods. Experimental results highlight the critical gaps in current defenses, underscoring the need for more robust strategies to secure advanced MLLMs. Content Warning: This paper contains examples that may be offensive. Aofan Liu, Lulu Tang, Yuguo Yin |
ICME | 2 |
| 2025 | Consistent multimodal pre-training for visual tokenization
Lulu Tang, Xin Liu 0044, Shiguang Shan |
Sci. China Inf. Sci. | 2 |
| 2024 | Tokenize Anything via Prompting
Lulu Tang, Shiguang Shan |
ECCV (47) | 2 |
| 2022 | Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingabstractWe present Point-BERT, a new paradigm for learning Transformers to generalize the concept of BERT [8] to 3D point cloud. Inspired by BERT, we devise a Masked Point Modeling (MPM) task to pre-train point cloud Transformers. Specifically, we first divide a point cloud into several local point patches, and a point cloud Tokenizer with a discrete Variational AutoEncoder (dVAE) is designed to generate discrete point tokens containing meaningful local information. Then, we randomly mask out some patches of input point clouds and feed them into the backbone Transformers. The pre-training objective is to recover the original point tokens at the masked locations under the supervision of point tokens obtained by the Tokenizer. Extensive experiments demonstrate that the proposed BERT-style pre-training strategy significantly improves the performance of standard point cloud Transformers. Equipped with our pre-training strategy, we show that a pure Transformer architecture attains 93.8% accuracy on ModelNet40 and 83.1% accuracy on the hardest setting of ScanObjectNN, surpassing carefully designed point cloud models with much fewer hand-made designs. We also demonstrate that the representations learned by Point-BERT transfer well to new tasks and domains, where our models largely advance the state-of-the-art of few-shot point cloud classification task. The code and pre-trained models are available at https://github.com/lulutang0608/Point-BERT. Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 0001, Jie Zhou 0001, Jiwen Lu |
CVPR | 2 |
| 2022 | Spike Transformer: Monocular Depth Estimation for Spiking Camera
Jiyuan Zhang 0005, Lulu Tang, Zhaofei Yu, Jiwen Lu, Tiejun Huang 0001 |
ECCV (7) | 2 |
| 2022 | IMRG: Impedance Matching Oriented Receiver Grouping for MIMO WPT SystemabstractIn recent years, multiple-input multiple-output (MIMO) technology has been imported into magnetic resonance coupled (MRC) enabled wireless power transfer (WPT) systems for concurrent charging of multiple devices. Besides the traditional performance optimization methods (e.g., TX current scheduling, system frequency adjustment, etc.), receiver (RX) grouping will also severely influence the achieved power- delivered-to-load (PDL). In this paper, we investigate the optimal RX grouping issue to maximize the proportional fairness of RX achieved PDL, which is a joint optimization problem involving RX grouping and time-slice allocation among groups. By decoupling the problem, we solve the group generation sub-problem with a impedance-matching based greedy algorithm to generate potential RX group candidates, and we further solve the time slice allocation sub-problem with a genetic algorithm to distribute resources among group candidates. We prototype the proposed system, denoted as IMRG, and conduct extensive experiments to evaluate the performance. The experimental results validate the effectiveness of the proposed algorithm, e.g., IMRG achieves average 59.6% PDL improvement through RX grouping compared to the simultaneous charging scheme. Lulu Tang, Hao Zhou 0001, Weiming Guo, Wangqiu Zhou, Xiaoyan Wang 0003 |
MSN | 1 |
| 2022 | Improving deep learning on point cloud by maximizing mutual information across layers
Di Wang 0032, Lulu Tang, Xu Wang 0037, Luqing Luo, Zhi-Xin Yang 0001 |
Pattern Recognit. | 2 |
| 2022 | Mutual Information Maximization Based Similarity Operation for 3D Point Cloud Completion NetworkabstractReconstructing the incomplete point cloud into the complete and uniform one is a fundamental task in 3D point cloud processing. Some existing studies have explored the use of deep learning networks and representation learning to implement point cloud completion, but the completed shapes still appear unrealistic and non-uniform. To address this issue, we propose the idea of simulating the process of point cloud completion with local-to-global reasoning (LGR). Motivated to achieve the fine LGR, a novel mutual information (MI) maximization-based similarity operation is proposed, which realizes the reconstruction from incomplete point cloud to complete one by maximizing MI between global features and the prior features from the same 3D object. We adopt the Jenson-Shannon MI estimator to maximize the MI between global features and the priors in specific dimensions, which effectively increases the similarity of global and prior features. Compared with other existing works and similarity operations, the superior results indicate the efficacy of the proposed method and its advantages over existing ones in both the synthetic and real-world datasets. Our source code is available athttps://github.com/wendydidi/MISO-PCN. Di Wang 0032, Lulu Tang, Zhi-Xin Yang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Improving Semantic Analysis on Point Clouds via Auxiliary Supervision of Local Geometric PriorsabstractExisting deep learning algorithms for point cloud analysis mainly concern discovering semantic patterns from the global configuration of local geometries in a supervised learning manner. However, very few explore geometric properties revealing local surface manifolds embedded in 3-D Euclidean space to discriminate semantic classes or object parts as additional supervision signals. This article is the first attempt to propose a unique multitask geometric learning network to improve semantic analysis by auxiliary geometric learning with local shape properties, which can be either generated via physical computation from point clouds themselves as self-supervision signals or provided as privileged information. Owing to explicitly encoding local shape manifolds in favor of semantic analysis, the proposed geometric self-supervised and privileged learning algorithms can achieve superior performance to their backbone baselines and other state-of-the-art methods, which are verified in the experiments on the popular benchmarks. Lulu Tang, Ke Chen 0004, Chaozheng Wu, Kui Jia, Zhi-Xin Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | PU-EVA: An Edge-Vector based Approximation Solution for Flexible-scale Point Cloud UpsamplingabstractHigh-quality point clouds have practical significance for point-based rendering, semantic understanding, and surface reconstruction. Upsampling sparse, noisy and non-uniform point clouds for a denser and more regular approximation of target objects is a desirable but challenging task. Most existing methods duplicate point features for upsampling, constraining the upsampling scales at a fixed rate. In this work, the arbitrary point clouds upsampling rates are achieved via edge-vector based affine combinations, and a novel design of Edge-Vector based Approximation for Flexible-scale Point clouds Upsampling (PU-EVA) is proposed. The edge-vector based approximation encodes neighboring connectivity via affine combinations based on edge vectors, and restricts the approximation error within a second-order term of Taylor’s Expansion. Moreover, the EVA upsampling decouples the upsampling scales with network architecture, achieving the arbitrary upsampling rates in one-time training. Qualitative and quantitative evaluations demonstrate that the proposed PU-EVA outperforms the state-of-the-arts in terms of proximity-to-surface, distribution uniformity, and geometric details preservation. Luqing Luo, Lulu Tang, Wanyi Zhou, Shizheng Wang, Zhi-Xin Yang 0001 |
ICCV | 2 |
| 2021 | Automatic representation and detection of fault bearings in in-wheel motors under variable load conditions
Xianbo Wang, Luqing Luo, Lulu Tang, Zhi-Xin Yang 0001 |
Adv. Eng. Informatics | 3 |
| 2021 | Object pose estimation for robot loading in accommodation space using alpha-shape algorithm
Qingda Guo, Lulu Tang, Jianchi Zhang |
Soft Comput. | 2 |