Haithem Turki

dblp:64/10771 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-5634-0918ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Rendering · 91% Geometric modeling and processing · 6% Virtual and augmented reality · 2%
Artificial intelligence
4 papers
3D vision · 65% Generative modeling · 14% Segmentation and scene understanding · 11%
Computer networks
1 paper
Edge and fog computing · 77% Internet of things and sensor networks · 23%

Topics — the 24 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
neural radiance fields
3.452024
HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces · CVPR 2024
SpecNeRF: Gaussian Directional Encoding for Specular Reflections · CVPR 2024
PyNeRF: Pyramidal Neural Radiance Fields · NeurIPS 2023
Rendering
neural rendering
2.032024
HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces · CVPR 2024
SUDS: Scalable Urban Dynamic Scenes · CVPR 2023
Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs · CVPR 2022
Computer vision › 3D vision
3d reconstruction
0.912025
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models · CVPR 2025
Machine learning › Generative modeling › diffusion model › image restoration
artifact correction
0.912025
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models · CVPR 2025
Computer vision › 3D vision
novel view synthesis
0.912025
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models · CVPR 2025
Rendering
directional encoding
0.812024
SpecNeRF: Gaussian Directional Encoding for Specular Reflections · CVPR 2024
Rendering
surface rendering
0.812024
HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces · CVPR 2024
Rendering
view-dependent rendering
0.812024
SpecNeRF: Gaussian Directional Encoding for Specular Reflections · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
3d instance segmentation
0.712023
SUDS: Scalable Urban Dynamic Scenes · CVPR 2023
Computer vision › 3D vision
3d scene reconstruction
0.712023
SUDS: Scalable Urban Dynamic Scenes · CVPR 2023
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.712023
SUDS: Scalable Urban Dynamic Scenes · CVPR 2023
Rendering
antialiasing
0.712023
PyNeRF: Pyramidal Neural Radiance Fields · NeurIPS 2023
Edge and fog computing
distributed learning
0.712023
Low-Bandwidth Self-Improving Transmission of Rare Training Data · MobiCom 2023
Geometric modeling and processing › 3d reconstruction › 3d scene reconstruction
large-scale scene reconstruction
0.612022
Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs · CVPR 2022
Computer vision › Segmentation and scene understanding
semantic segmentation
0.322015
Parameter Estimation and Energy Minimization for Region-Based Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Learning specific-class segmentation from diverse data · ICCV 2011
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.212015
Parameter Estimation and Energy Minimization for Region-Based Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Computer vision › Segmentation and scene understanding › image segmentation
region-based segmentation
0.212015
Parameter Estimation and Energy Minimization for Region-Based Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Computer vision › 3D vision
3d object detection
0.212023
SUDS: Scalable Urban Dynamic Scenes · CVPR 2023
Rendering › neural radiance fields
grid-based neural radiance field
0.212023
PyNeRF: Pyramidal Neural Radiance Fields · NeurIPS 2023
Internet of things and sensor networks › data dissemination
sensor data transmission
0.212023
Low-Bandwidth Self-Improving Transmission of Rare Training Data · MobiCom 2023
Machine learning › Kernel, tree and ensemble methods › support vector machine
latent structural SVM
0.222015
Learning specific-class segmentation from diverse data · ICCV 2011
Parameter Estimation and Energy Minimization for Region-Based Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Parallel and multicore computing › parallel computing › parallel machine learning
data-parallel training
0.212022
Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs · CVPR 2022
Computer vision › Segmentation and scene understanding › semantic segmentation
category-specific segmentation
0.112011
Learning specific-class segmentation from diverse data · ICCV 2011
Machine learning › Learning paradigms
weakly supervised learning
0.112011
Learning specific-class segmentation from diverse data · ICCV 2011

Methods — techniques the papers use, named apart from their topics

optical flow · 1.3hash table factorization · 1.3geometric loss · 1.3feature-metric loss · 1.3single-step diffusion model · 0.9distillation · 0.93d gaussian splatting · 0.9volume rendering · 0.8surface representation · 0.8signed distance function · 0.8prefiltered environment map · 0.8geometry prior · 0.8gaussian directional encoding · 0.8transfer learning · 0.7semi-supervised learning · 0.7self-supervised 2d descriptors · 0.7photometric loss · 0.7few-shot learning · 0.7
YearPublicationVenuePosition
2025 DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
abstract
Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by under-constrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2× improvement in FID score over baselines while maintaining 3D consistency.
Jay Zhangjie Wu, Yuxuan Zhang 0001, Haithem Turki, Xuanchi Ren, Jun Gao 0004, Zheng Shou 0001, Sanja Fidler, Zan Gojcic, Huan Ling
CVPR3
2024 SpecNeRF: Gaussian Directional Encoding for Specular Reflections
abstract
Neural radiance fields have achieved remarkable performance in modeling the appearance of 3D scenes. However, existing approaches still struggle with the view-dependent appearance of glossy surfaces, especially under complex lighting of indoor environments. Unlike existing methods, which typically assume distant lighting like an environment map, we propose a learnable Gaussian directional encoding to better model the view-dependent effects under near-field lighting conditions. Importantly, our new directional encoding captures the spatially-varying nature of near-field lighting and emulates the behavior of prefiltered environment maps. As a result, it enables the efficient evaluation of preconvolved specular color at any 3D location with varying roughness coefficients. We further introduce a data-driven geometry prior that helps alleviate the shape radiance ambiguity in reflection modeling. We show that our Gaussian directional encoding and geometry prior significantly improve the modeling of challenging specular reflections in neural radiance fields, which helps decompose appearance into more physically meaningful components.
Vasu Agrawal, Haithem Turki, Changil Kim 0001, Chen Gao 0003, Pedro V. Sander, Michael Zollhöfer, Christian Richardt
CVPR3
2024 HybridNeRF: Efficient Neural Rendering via Adaptive Volumetric Surfaces
abstract
Neural radiance fields provide state-of-the-art view synthesis quality but tend to be slow to render. One reason is that they make use of volume rendering, thus requiring many samples (and model queries) per ray at render time. Although this representation is flexible and easy to optimize, most real-world objects can be modeled more efficiently with surfaces instead of volumes, requiring far fewer samples per ray. This observation has spurred considerable progress in surface representations, such as signed distance functions, but these may struggle to model semi-opaque and thin structures. We propose a method, HybridNeRF, that leverages the strengths of both representations by rendering most objects as surfaces while modeling the (typically) small fraction of challenging regions volumetrically. We evaluate Hybrid-NeRF against the challenging Eyeful Tower dataset [38] along with other commonly used view synthesis datasets. When comparing to state-of-the-art baselines, including recent rasterization-based approaches, we improve error rates by 15-30% while achieving real-time framerates (at least 36 FPS) for virtual-reality resolutions ($2K\times 2K$). Project page: https://haithemturki.com/hybrid-nerf/.
Haithem Turki, Vasu Agrawal, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder, Deva Ramanan, Michael Zollhöfer, Christian Richardt
CVPR1
2023 SUDS: Scalable Urban Dynamic Scenes
abstract
We extend neural radiance fields (NeRFs) to dynamic large-scale urban scenes. Prior work tends to reconstruct single video clips of short durations (up to 10 seconds). Two reasons are that such methods (a) tend to scale linearly with the number of moving objects and input videos because a separate model is built for each and (b) tend to require supervision via 3D bounding boxes and panoptic labels, obtained manually or via category-specific models. As a step towards truly open-world reconstructions of dynamic cities, we introduce two key innovations: (a) we factorize the scene into three separate hash table data structures to efficiently encode static, dynamic, and far-field radiance fields, and (b) we make use of unlabeled target signals consisting of RGB images, sparse LiDAR, off-the-shelf self-supervised 2D descriptors, and most importantly, 2D optical flow. Operationalizing such inputs via photometric, geometric, and feature-metric reconstruction losses enables SUDS to decompose dynamic scenes into the static background, individual objects, and their motions. When combined with our multi-branch table representation, such reconstructions can be scaled to tens of thousands of objects across 1.2 million frames from 1700 videos spanning geospatial footprints of hundreds of kilometers, (to our knowledge) the largest dynamic NeRF built to date. We present qualitative initial results on a variety of tasks enabled by our representations, including novel-view synthesis of dynamic urban scenes, unsupervised 3D instance segmentation, and unsupervised 3D cuboid detection. To compare to prior work, we also evaluate on KITTI and Virtual KITTI 2, surpassing state-of-the-art methods that rely on ground truth 3D bounding box annotations while being 10x quicker to train.
Haithem Turki, Jason Y. Zhang 0001, Francesco Ferroni, Deva Ramanan
CVPR1
2023 Low-Bandwidth Self-Improving Transmission of Rare Training Data
abstract
A severe bandwidth mismatch between incoming sensor data rate and wireless backhaul bandwidth often exists on unmanned probes when collecting new training data for machine learning (ML). To overcome this mismatch, we describe a self-improving ML-based transmission system called Hawk. Starting from a weak model that is trained on just a few examples, it seamlessly pipelines semi-supervised learning, active learning, and transfer learning, with asynchronous bandwidth-sensitive data transmission to a distant human for labeling. When a significant number of true positives (TPs) have been labeled, Hawk trains an improved model to replace the old model. This iterative workflow, called Live Learning, continues until a sufficient number of TPs have been collected. For very rare events on challenging datasets, and bandwidths as low as 12 kbps, a team of 7 probes using Hawk discovers up to 87% of the TPs that could have been discovered via full preview, transmission and labeling of all mission data. Hawk also uses diversity sampling and few-shot learning.
Shilpa Anna George, Haithem Turki, Ziqiang Feng, Deva Ramanan, Padmanabhan Pillai, Mahadev Satyanarayanan
MobiCom2
2023 PyNeRF: Pyramidal Neural Radiance Fields
abstract
Neural Radiance Fields (NeRFs) can be dramatically accelerated by spatial grid representations. However, they do not explicitly reason about scale and so introduce aliasing artifacts when reconstructing scenes captured at different camera distances. Mip-NeRF and its extensions propose scale-aware renderers that project volumetric frustums rather than point samples. But such approaches rely on positional encodings that are not readily compatible with grid methods. We propose a simple modification to grid-based models by training model heads at different spatial grid resolutions. At render time, we simply use coarser grids to render samples that cover larger volumes. Our method can be easily applied to existing accelerated NeRF methods and significantly improves rendering quality (reducing error rates by 20–90% across synthetic and unbounded real-world scenes) while incurring minimal performance overhead (as each model head is quick to evaluate). Compared to Mip-NeRF, we reduce error rates by 20% while training over 60x faster.
Haithem Turki, Michael Zollhöfer, Christian Richardt, Deva Ramanan
NeurIPS1
2022 Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly- Throughs
abstract
We use neural radiance fields (NeRFs) to build interac-tive 3D environments from large-scale visual captures spanning buildings or even multiple city blocks collected pri-marily from drones. In contrast to single object scenes (on which NeRFs are traditionally evaluated), our scale poses multiple challenges including (1) the need to model thou-sands of images with varying lighting conditions, each of which capture only a small subset of the scene, (2) pro-hibitively large model capacities that make it infeasible to train on a single GPU, and (3) significant challenges for fast rendering that would enable interactive fly-throughs. To address these challenges, we begin by analyzing visi-bility statistics for large-scale scenes, motivating a sparse network structure where parameters are specialized to dif-ferent regions of the scene. We introduce a simple geomet-ric clustering algorithm for data parallelism that partitions training images (or rather pixels) into different NeRF sub-modules that can be trained in parallel. We evaluate our approach on existing datasets (Quad 6k and UrbanScene3D) as well as against our own drone footage, improving training speed by 3x and PSNR by 12%. We also evaluate re-cent NeRF fast renderers on top of Mega-NeRF and intro-duce a novel method that exploits temporal coherence. Our technique achieves a 40x speedup over conventional NeRF rendering while remaining within 0.8 db in PSNR quality, exceeding the fidelity of existing fast renderers.
Haithem Turki, Deva Ramanan, Mahadev Satyanarayanan
CVPR1
2015 Parameter Estimation and Energy Minimization for Region-Based Semantic Segmentation
abstract
We consider the problem of parameter estimation and energy minimization for a region-based semantic segmentation model. The model divides the pixels of an image into non-overlapping connected regions, each of which is to a semantic class. In the context of energy minimization, the main problem we face is the large number of putative pixel-to-region assignments. We address this problem by designing an accurate linear programming based approach for selecting the best set of regions from a large dictionary. The dictionary is constructed by merging and intersecting segments obtained from multiple bottom-up over-segmentations. The linear program is solved efficiently using dual decomposition. In the context of parameter estimation, the main problem we face is the lack of fully supervised data. We address this issue by developing a principled framework for parameter estimation using diverse data. More precisely, we propose a latent structural support vector machine formulation, where the latent variables model any missing information in the human annotation. Of particular interest to us are three types of annotations: (i) images segmented using generic foreground or background classes; (ii) images with bounding boxes specified for objects; and (iii) images labeled to indicate the presence of a class. Using large, publicly available datasets we show that our methods are able to significantly improve the accuracy of the region-based model.
M. Pawan Kumar, Haithem Turki, Dan Preston, Daphne Koller
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Learning specific-class segmentation from diverse data
abstract
We consider the task of learning the parameters of a segmentation model that assigns a specific semantic class to each pixel of a given image. The main problem we face is the lack of fully supervised data. We address this issue by developing a principled framework for learning the parameters of a specific-class segmentation model using diverse data. More precisely, we propose a latent structural support vector machine formulation, where the latent variables model any missing information in the human annotation. Of particular interest to us are three types of annotations: (i) images segmented using generic foreground or background classes; (ii) images with bounding boxes specified for objects; and (iii) images labeled to indicate the presence of a class. Using large, publicly available datasets we show that our approach is able to exploit the information present in different annotations to improve the accuracy of a state-of-the art region-based model.
M. Pawan Kumar, Haithem Turki, Dan Preston, Daphne Koller
ICCV2