Pratham Nawal

dblp:223/4766 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › visual speech recognition
lip reading
0.312018
Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
silent speech recognition
0.312018
Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018
Natural language and speech › Speech recognition and synthesis
speech reconstruction
0.312018
Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018
Multimedia analysis and retrieval › video content analysis
multi-view video analysis
0.112018
Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed · ACM Multimedia 2018

Methods — techniques the papers use, named apart from their topics

multi-view fusion · 0.7
YearPublicationVenuePosition
2018 Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed
abstract
Speechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in surveillance, Internet telephony, and as an aid to a person with hearing impairments. However, most of the work in speechreading has been limited to text generation from silent videos. Recently, research has started venturing into generating (audio) speech from silent video sequences but there have been no developments thus far in dealing with divergent views and poses of a speaker. Thus although, we have multiple camera feeds for the speech of a user, but we have failed in using these multiple video feeds for dealing with the different poses. To this end, this paper presents the world's first ever multi-view speech reading and reconstruction system. This work encompasses the boundaries of multimedia research by putting forth a model which leverages silent video feeds from multiple cameras recording the same subject to generate intelligent speech for a speaker. Initial results confirm the usefulness of exploiting multiple camera views in building an efficient speech reading and reconstruction system. It further shows the optimal placement of cameras which would lead to the maximum intelligibility of speech. Next, it lays out various innovative applications for the proposed system focusing on its potential prodigious impact in not just security arena but in many other multimedia analytics problems.
Yaman Singla, Mayank Aggarwal, Pratham Nawal, Shin'ichi Satoh 0001, Rajiv Ratn Shah, Roger Zimmermann
ACM Multimedia3