Gefen Kohavi

dblp:224/3256 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 59% Vision and language · 12% Knowledge representation and reasoning · 12%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
0.912025
Cubify Anything: Scaling Indoor 3D Object Detection · CVPR 2025
Computer vision › 3D vision
3d scene understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › 3D vision › 3d object detection
indoor 3d object detection
0.912025
Cubify Anything: Scaling Indoor 3D Object Detection · CVPR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › 3D vision › 3d object detection › multimodal 3d object detection
RGB-D 3D object detection
0.912025
Cubify Anything: Scaling Indoor 3D Object Detection · CVPR 2025
Computer vision › 3D vision
spatial understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › Face, body and person analysis › nonverbal behavior analysis
gesture analysis
0.412019
Learning Individual Styles of Conversational Gesture · CVPR 2019
Computer vision › Face, body and person analysis
human pose analysis
0.412019
Learning Individual Styles of Conversational Gesture · CVPR 2019
Computer animation and physical simulation
gesture generation
0.412019
Learning Individual Styles of Conversational Gesture · CVPR 2019
Computer vision › Image recognition and object detection › object detection
detection transformer
0.312025
Cubify Anything: Scaling Indoor 3D Object Detection · CVPR 2025
Computer vision › Image recognition and object detection
object detection
0.312025
Cubify Anything: Scaling Indoor 3D Object Detection · CVPR 2025

Methods — techniques the papers use, named apart from their topics

voxel representation · 0.9transformer · 0.9point cloud · 0.9multimodal large language model · 0.9pose detection · 0.8cross-modal translation · 0.8
YearPublicationVenuePosition
2025 Cubify Anything: Scaling Indoor 3D Object Detection
abstract
We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respect to both data and modeling. First, we establish that existing datasets have significant limitations to scale, accuracy, and diversity of objects. As a result, we introduce the Cubify-Anything 1M (CA-1M) dataset, which exhaustively labels over 400K 3D objects on over 1K highly accurate laser-scanned scenes with near-perfect registration to over 3.5K handheld, egocentric captures. Next, we establish Cubify Transformer (CuTR), a fully Transformer 3D object detection baseline which rather than operating in 3D on point or voxel-based representations, predicts 3D boxes directly from 2D features derived from RGB(-D) inputs. While this approach lacks any 3D inductive biases, we show that paired with CA-1M, CuTR outperforms pointbased methods — accurately recalling over 62%c of objects in 3D, and is significantly more capable at handling noise and uncertainty present in commodity LiDAR-derived depth maps while also providing promising RGB only performance without architecture changes. Furthermore, by pre-training on CA-1M, CuTR can outperform point-based methods on a more diverse variant of SUN RGB-D — supporting the notion that while inductive biases in 3D are useful at the smaller sizes of existing datasets, they fail to scale to the data-rich regime of CA-1M. Overall, this dataset and baseline model provide strong evidence that we are moving towards models which can effectively Cubify Anything.
Justin Lazarow, David Griffiths, Gefen Kohavi, Francisco Crespo, Afshin Dehghan
CVPR3
2025 MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch
ICCV6
2019 Learning Individual Styles of Conversational Gesture
abstract
Human speech is often accompanied by hand and arm gestures. We present a method for cross-modal translation from "in-the-wild" monologue speech of a single speaker to their conversational gesture motion. We train on unlabeled videos for which we only have noisy pseudo ground truth from an automatic pose detection system. Our proposed model significantly outperforms baseline methods in a quantitative comparison. To support research toward obtaining a computational understanding of the relationship between gesture and speech, we release a large video dataset of person-specific gestures.
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, Jitendra Malik
CVPR3