EDBT 2026 Demo / reviewers in the wild / expert
Zhengjia Huang
dblp:209/4879
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 54% Generative modeling · 23% Knowledge representation and reasoning · 23% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 50% Audio and music processing · 50% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
analysis-by-synthesis |
0.3 | 1 | 2017 | Shape and Material from Sound · NIPS 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.3 | 1 | 2017 | Shape and Material from Sound · NIPS 2017 |
Computer vision › 3D vision
physical property estimation |
0.3 | 1 | 2017 | Shape and Material from Sound · NIPS 2017 |
Multimedia analysis and retrieval
audio-visual learning |
0.3 | 1 | 2017 | Generative Modeling of Audible Shapes for Object Perception · ICCV 2017 |
Computer vision › 3D vision
3d shape analysis |
0.1 | 1 | 2017 | Generative Modeling of Audible Shapes for Object Perception · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
physics-based simulation · 0.6physical simulation · 0.6learned mapping · 0.6generative modeling · 0.6analysis-by-synthesis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Generative Modeling of Audible Shapes for Object PerceptionabstractHumans infer rich knowledge of objects from both auditory and visual cues. Building a machine of such competency, however, is very challenging, due to the great difficulty in capturing large-scale, clean data of objects with both their appearance and the sound they make. In this paper, we present a novel, open-source pipeline that generates audiovisual data, purely from 3D object shapes and their physical properties. Through comparison with audio recordings and human behavioral studies, we validate the accuracy of the sounds it generates. Using this generative model, we are able to construct a synthetic audio-visual dataset, namely Sound-20K, for object perception tasks. We demonstrate that auditory and visual information play complementary roles in object perception, and further, that the representation learned on synthetic audio-visual data can transfer to real-world scenarios. Zhoutong Zhang, Jiajun Wu 0001, Qiujia Li, Zhengjia Huang, James Traer, Josh H. McDermott, Josh Tenenbaum, William T. Freeman |
ICCV | 4 |
| 2017 | Shape and Material from SoundabstractHearing an object falling onto the ground, humans can recover rich information including its rough shape, material, and falling height. In this paper, we build machines to approximate such competency. We first mimic human knowledge of the physical world by building an efficient, physics-based simulation engine. Then, we present an analysis-by-synthesis approach to infer properties of the falling object. We further accelerate the process by learning a mapping from a sound wave to object properties, and using the predicted values to initialize the inference. This mapping can be viewed as an approximation of human commonsense learned from past experience. Our model performs well on both synthetic audio clips and real recordings without requiring any annotated data. We conduct behavior studies to compare human responses with ours on estimating object shape, material, and falling height from sound. Our model achieves near-human performance. Zhoutong Zhang, Qiujia Li, Zhengjia Huang, Jiajun Wu 0001, Josh Tenenbaum, William T. Freeman |
NIPS | 3 |