EDBT 2026 Demo / reviewers in the wild / expert
Marco Monteiro
dblp:47/2749
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 30% Generative modeling · 15% Deep learning architectures and training · 13% | |
| Computer graphics and multimedia
1 paper |
Rendering · 87% Image and video processing · 13% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Image recognition and object detection
spatial alignment |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
vision encoder |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.5 | 1 | 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis · CVPR 2021 |
Machine learning › Generative modeling
generative adversarial network |
0.5 | 1 | 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis · CVPR 2021 |
Computer vision › 3D vision
neural rendering |
0.5 | 1 | 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis · CVPR 2021 |
Computer vision › 3D vision › neural rendering
volume rendering |
0.5 | 1 | 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis · CVPR 2021 |
Rendering
hybrid explicit-implicit representation |
0.5 | 1 | 2021 | Acorn: adaptive coordinate networks for neural scene representation · ACM Trans. Graph. 2021 |
Rendering › neural rendering
neural scene representation |
0.5 | 1 | 2021 | Acorn: adaptive coordinate networks for neural scene representation · ACM Trans. Graph. 2021 |
Computer vision › Video understanding and tracking
video classification |
0.3 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › 3D vision › multi-view geometry
multi-view consistency |
0.1 | 1 | 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 0.9alignment method · 0.9periodic activation function · 0.5neural radiance field · 0.5feature grid · 0.5feature decoder · 0.5coordinate network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perception Encoder: The best visual embeddings are not at the output of the networkabstractWe introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models Daniel Bolya, Po-Yao Huang 0001, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei 0005, Tengyu Ma 0005, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Marco Monteiro, Hu Xu 0001, Shiyu Dong, Nikhila Ravi, Shang-Wen Li 0001, Piotr Dollár, Christoph Feichtenhofer |
NeurIPS | 12 |
| 2021 | Pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image SynthesisabstractWe have witnessed rapid progress on 3D-aware image synthesis, leveraging recent advances in generative visual models and neural rendering. Existing approaches how-ever fall short in two ways: first, they may lack an under-lying 3D representation or rely on view-inconsistent rendering, hence synthesizing images that are not multi-view consistent; second, they often depend upon representation network architectures that are not expressive enough, and their results thus lack in image quality. We propose a novel generative model, named Periodic Implicit Generative Adversarial Networks (π-GAN or pi-GAN), for high-quality 3D-aware image synthesis. π-GAN leverages neural representations with periodic activation functions and volumetric rendering to represent scenes as view-consistent radiance fields. The proposed approach obtains state-of-the-art results for 3D-aware image synthesis with multiple real and synthetic datasets. Eric R. Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu 0001, Gordon Wetzstein |
CVPR | 2 |
| 2021 | Acorn: adaptive coordinate networks for neural scene representationabstractNeural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as meshes, point clouds, or volumes they can be flexibly incorporated into differentiable learning-based pipelines. While recent improvements to neural representations now make it possible to represent signals with fine details at moderate resolutions (e.g., for images and 3D shapes), adequately representing large-scale or complex scenes has proven a challenge. Current neural representations fail to accurately represent images at resolutions greater than a megapixel or 3D scenes with more than a few hundred thousand polygons. Here, we introduce a new hybrid implicit-explicit network architecture and training strategy that adaptively allocates resources during training and inference based on the local complexity of a signal of interest. Our approach uses a multiscale block-coordinate decomposition, similar to a quadtree or octree, that is optimized during training. The network architecture operates in two stages: using the bulk of the network parameters, a coordinate encoder generates a feature grid in a single forward pass. Then, hundreds or thousands of samples within each block can be efficiently evaluated using a lightweight feature decoder. With this hybrid implicit-explicit network architecture, we demonstrate the first experiments that fit gigapixel images to nearly 40 dB peak signal-to-noise ratio. Notably this represents an increase in scale of over 1000X compared to the resolution of previously demonstrated image-fitting experiments. Moreover, our approach is able to represent 3D shapes significantly faster and better than previous techniques; it reduces training times from days to hours or minutes and memory requirements by over an order of magnitude. Julien N. P. Martel, David B. Lindell, Connor Z. Lin, Eric R. Chan, Marco Monteiro, Gordon Wetzstein |
ACM Trans. Graph. | 5 |
| 2007 | A Proposal to Delegate GUI Implementation using a Source Code based Model
Marco Monteiro, Paula Oliveira, Ramiro Gonçalves |
SEKE | 1 |