EDBT 2026 Demo / reviewers in the wild / expert
Luming Tang
dblp:203/8352
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-6944-1101ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 30% Transfer learning and domain adaptation · 13% 3D vision · 11% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 76% Rendering · 24% |
Topics — the 25 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.1 | 3 | 2024 | RealFill: Reference-Driven Generation for Authentic Image Completion · ACM Trans. Graph. 2024 Emergent Correspondence from Image Diffusion · NeurIPS 2023 Magic3D: High-Resolution Text-to-3D Content Creation · CVPR 2023 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model · CVPR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning |
0.9 | 1 | 2025 | Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
personalized image generation |
0.8 | 1 | 2024 | RealFill: Reference-Driven Generation for Authentic Image Completion · ACM Trans. Graph. 2024 |
Visual content generation and editing
image completion |
0.8 | 1 | 2024 | RealFill: Reference-Driven Generation for Authentic Image Completion · ACM Trans. Graph. 2024 |
Computer vision › 3D vision › feature matching
correspondence problem |
0.7 | 1 | 2023 | Emergent Correspondence from Image Diffusion · NeurIPS 2023 |
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
0.7 | 1 | 2023 | Emergent Correspondence from Image Diffusion · NeurIPS 2023 |
Visual content generation and editing
3d content generation |
0.7 | 1 | 2023 | Magic3D: High-Resolution Text-to-3D Content Creation · CVPR 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | Magic3D: High-Resolution Text-to-3D Content Creation · CVPR 2023 |
Visual content generation and editing › 3d content generation
text-to-3d generation |
0.7 | 1 | 2023 | Magic3D: High-Resolution Text-to-3D Content Creation · CVPR 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.6 | 1 | 2022 | Visual Prompt Tuning · ECCV (33) 2022 |
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
visual prompt tuning |
0.6 | 1 | 2022 | Visual Prompt Tuning · ECCV (33) 2022 |
Machine learning › Transfer learning and domain adaptation
few-shot classification |
0.5 | 1 | 2021 | Few-Shot Classification With Feature Map Reconstruction Networks · CVPR 2021 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.5 | 1 | 2021 | Few-Shot Classification With Feature Map Reconstruction Networks · CVPR 2021 |
Computer vision › Image recognition and object detection › image classification › fine-grained image classification
few-shot fine-grained classification |
0.4 | 1 | 2020 | Revisiting Pose-Normalization for Fine-Grained Few-Shot Recognition · CVPR 2020 |
Computer vision › Face, body and person analysis
pose normalization |
0.4 | 1 | 2020 | Revisiting Pose-Normalization for Fine-Grained Few-Shot Recognition · CVPR 2020 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.3 | 1 | 2018 | Multi-Entity Dependence Learning With Rich Context via Conditional Variational Auto-Encoder · AAAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.3 | 1 | 2018 | Multi-Entity Dependence Learning With Rich Context via Conditional Variational Auto-Encoder · AAAI 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2018 | Multi-Entity Dependence Learning With Rich Context via Conditional Variational Auto-Encoder · AAAI 2018 |
Computer vision › Face, body and person analysis
person re-identification |
0.3 | 1 | 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identification · ICCV 2017 |
Computer vision › Face, body and person analysis › person re-identification
vehicle re-identification |
0.3 | 1 | 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identification · ICCV 2017 |
Machine learning › Representation and self-supervised learning › visual representation › image representation
image descriptor |
0.2 | 1 | 2023 | Emergent Correspondence from Image Diffusion · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › reconstruction-based representation learning
feature reconstruction |
0.1 | 1 | 2021 | Few-Shot Classification With Feature Map Reconstruction Networks · CVPR 2021 |
Computational science and engineering
computational sustainability |
0.1 | 1 | 2018 | Multi-Entity Dependence Learning With Rich Context via Conditional Variational Auto-Encoder · AAAI 2018 |
Computer vision › Video understanding and tracking
spatio-temporal regularization |
0.1 | 1 | 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identification · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
model personalization · 1.5generative inpainting · 1.5sparse 3d hash grid · 1.3latent diffusion model · 1.3differentiable renderer · 1.3visual prompting · 0.9tracking model · 0.9diffusion feature extraction · 0.7visual prompt · 0.6transformer · 0.6deep neural network · 0.3conditional variational autoencoder · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language ModelabstractMultimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural designs or task-specific fine-tuning to achieve this. We introduce Coarse Correspondences, a simple lightweight method that enhances MLLMs’ spatial-temporal reasoning with 2D images as input, without modifying the architecture or requiring task-specific fine-tuning. Our method uses a lightweight tracking model to identify primary object correspondences between frames in a video or across different image viewpoints, and then conveys this information to MLLMs through visual prompting. We demonstrate that this simple training-free approach brings substantial gains to GPT4-V/O consistently on four benchmarks that require spatial-temporal reasoning, including +20.5% improvement on ScanQA, +9.7% on OpenEQA’s episodic memory subset, +6.0% on the long-form video benchmark EgoSchema, and +11% on the R2R navigation benchmark. Additionally, we show that Coarse Correspondences can also enhance open-source MLLMs’ spatial reasoning (by +6.9% on ScanQA) when applied in both training and inference and that the improvement can generalize to unseen datasets such as SQA3D (+3.1%). Taken together, we show that Coarse Correspondences effectively and efficiently boosts models’ performance on downstream tasks requiring spatial-temporal reasoning. Benlin Liu, Yuhao Dong, Zixian Ma, Yansong Tang, Luming Tang, Yongming Rao, Wei-Chiu Ma, Ranjay Krishna |
CVPR | 6 |
| 2024 | RealFill: Reference-Driven Generation for Authentic Image CompletionabstractRecent advances in generative imagery have brought forth outpainting and inpainting models that can produce high-quality, plausible image content in unknown regions. However, the content these models hallucinate is necessarily inauthentic, since they are unaware of the true scene. In this work, we propose RealFill, a novel generative approach for image completion that fills in missing regions of an image with the content that should have been there. RealFill is a generative inpainting model that is personalized using only a few reference images of a scene. These reference images do not have to be aligned with the target image, and can be taken with drastically varying viewpoints, lighting conditions, camera apertures, or image styles. Once personalized, RealFill is able to complete a target image with visually compelling contents that are faithful to the original scene. We evaluate RealFill on a new image completion benchmark that covers a set of diverse and challenging scenarios, and find that it outperforms existing approaches by a large margin. Project page: https://realfill.github.io. Luming Tang, Nataniel Ruiz, Qinghao Chu, Yuanzhen Li, Aleksander Holynski, David E. Jacobs, Bharath Hariharan, Yael Pritch, Neal Wadhwa, Kfir Aberman, Michael Rubinstein |
ACM Trans. Graph. | 1 |
| 2023 | Magic3D: High-Resolution Text-to-3D Content CreationabstractDreamFusion [31] has recently demonstrated the utility of a pretrained text-to-image diffusion model to optimize Neural Radiance Fields (NeRF) [23], achieving remarkable text-to-3D synthesis results. However, the method has two inherent limitations: (a) extremely slow optimization of NeRF and (b) low-resolution image space supervision on NeRF, leading to low-quality 3D models with a long processing time. In this paper, we address these limitations by utilizing a two-stage optimization framework. First, we obtain a coarse model using a low-resolution diffusion prior and accelerate with a sparse 3D hash grid structure. Using the coarse representation as the initialization, we further optimize a textured 3D mesh model with an efficient differentiable renderer interacting with a high-resolution latent diffusion model. Our method, dubbed Magic3D, can create high quality 3D mesh models in 40 minutes, which is 2× faster than DreamFusion (reportedly taking 1.5 hours on average), while also achieving higher resolution. User studies show 61.7% raters to prefer our approach over DreamFusion. Together with the image-conditioned generation capabilities, we provide users with new ways to control 3D synthesis, opening up new avenues to various creative applications. Chen-Hsuan Lin 0001, Jun Gao 0004, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang 0002, Karsten Kreis, Sanja Fidler, Ming-Yu Liu 0001, Tsung-Yi Lin |
CVPR | 3 |
| 2023 | Emergent Correspondence from Image DiffusionabstractFinding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image features, namely DIffusion FeaTures (DIFT), and use them to establish correspondences between real images. Without any additional fine-tuning or supervision on the task-specific data or annotations, DIFT is able to outperform both weakly-supervised methods and competitive off-the-shelf features in identifying semantic, geometric, and temporal correspondences. Particularly for semantic correspondence, DIFT from Stable Diffusion is able to outperform DINO and OpenCLIP by 19 and 14 accuracy points respectively on the challenging SPair-71k benchmark. It even outperforms the state-of-the-art supervised methods on 9 out of 18 categories while remaining on par for the overall performance. Project page: https://diffusionfeatures.github.io. Luming Tang, Menglin Jia, Cheng Perng Phoo, Bharath Hariharan |
NeurIPS | 1 |
| 2022 | Visual Prompt Tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge J. Belongie, Bharath Hariharan, Ser-Nam Lim |
ECCV (33) | 2 |
| 2021 | Few-Shot Classification With Feature Map Reconstruction NetworksabstractIn this paper we reformulate few-shot classification as a reconstruction problem in latent space. The ability of the network to reconstruct a query feature map from support features of a given class predicts membership of the query in that class. We introduce a novel mechanism for few-shot classification by regressing directly from support features to query features in closed form, without introducing any new modules or large-scale learnable parameters. The resulting Feature Map Reconstruction Networks are both more performant and computationally efficient than previous approaches. We demonstrate consistent and substantial accuracy gains on four fine-grained benchmarks with varying neural architectures. Our model is also competitive on the non-fine-grained mini-ImageNet and tiered-ImageNet benchmarks with minimal bells and whistles.1 Davis Wertheimer, Luming Tang, Bharath Hariharan |
CVPR | 2 |
| 2020 | Revisiting Pose-Normalization for Fine-Grained Few-Shot RecognitionabstractFew-shot, fine-grained classification requires a model to learn subtle, fine-grained distinctions between different classes (e.g., birds) based on a few images alone. This requires a remarkable degree of invariance to pose, articulation and background. A solution is to use pose-normalized representations: first localize semantic parts in each image, and then describe images by characterizing the appearance of each part. While such representations are out of favor for fully supervised classification, we show that they are extremely effective for few-shot fine-grained classification. With a minimal increase in model capacity, pose normalization improves accuracy between 10 and 20 percentage points for shallow and deep architectures, generalizes better to new domains, and is effective for multiple few-shot algorithms and network backbones. Code is available at https://github.com/Tsingularity/PoseNorm_Fewshot. Luming Tang, Davis Wertheimer, Bharath Hariharan |
CVPR | 1 |
| 2018 | Multi-Entity Dependence Learning With Rich Context via Conditional Variational Auto-EncoderabstractMulti-Entity Dependence Learning (MEDL) explores conditional correlations among multiple entities. The availability of rich contextual information requires a nimble learning scheme that tightly integrates with deep neural networks and has the ability to capture correlation structures among exponentially many outcomes. We propose MEDL_CVAE, which encodes a conditional multivariate distribution as a generating process. As a result, the variational lower bound of the joint likelihood can be optimized via a conditional variational auto-encoder and trained end-to-end on GPUs. Our MEDL_CVAE was motivated by two real-world applications in computational sustainability: one studies the spatial correlation among multiple bird species using the eBird data and the other models multi-dimensional landscape composition and human footprint in the Amazon rainforest with satellite images. We show that MEDL_CVAE captures rich dependency structures, scales better than previous methods, and further improves on the joint likelihood taking advantage of very large datasets that are beyond the capacity of previous methods. Luming Tang, Yexiang Xue, Di Chen 0001, Carla P. Gomes |
AAAI | 1 |
| 2017 | Orientation Invariant Feature Embedding and Spatial Temporal Regularization for Vehicle Re-identificationabstractIn this paper, we tackle the vehicle Re-identification (ReID) problem which is of great importance in urban surveillance and can be used for multiple applications. In our vehicle ReID framework, an orientation invariant feature embedding module and a spatial-temporal regularization module are proposed. With orientation invariant feature embedding, local region features of different orientations can be extracted based on 20 key point locations and can be well aligned and combined. With spatial-temporal regularization, the log-normal distribution is adopted to model the spatial-temporal constraints and the retrieval results can be refined. Experiments are conducted on public vehicle ReID datasets and our proposed method achieves state-of-the-art performance. Investigations of the proposed framework is conducted, including the landmark regressor and comparisons with attention mechanism. Both the orientation invariant feature embedding and the spatio-temporal regularization achieve considerable improvements. Zhongdao Wang, Luming Tang, Xihui Liu, Zhuliang Yao, Shuai Yi, Shengjin Wang, Hongsheng Li 0001, Xiaogang Wang 0001 |
ICCV | 2 |