VLDB 2026 Research / reviewers in the wild / expert
Eric Petajan
dblp:47/1827
· DBLP profile ↗
16ranked-venue papers
6as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streaming Video QoE Prediction using Key Quality Indicators from Screen Recordings
Morey Antebi, Katelyn Bandy, Jyotirmoy Banik, Stephanie Berger, Sylvie Covey, Greg Edwards, Eric Petajan, Igor Pruzhansky, Thulasiraman Sampath, Subhabrata Sen |
QoMEX | 7 |
| 2026 | Prototyping QoE-Aware Rate Adaptation in Cellular Networks with Commercial ApplicationsabstractPrior work has shown that QoE-aware resource sharing for real-time interactive video can support up to three times more simultaneous sessions at acceptable quality compared to rate-fair allocation. However, the required capabilities (QoE-targeted encoding, runtime spatial complexity estimation, and rich application-network APIs) are not yet available in commercial deployments. In this paper, we take an evolutionary approach: we design a system that delivers QoE-aware resource allocation using only capabilities that can be assembled in a lab today. We extend the utility-based allocation framework to the radio resource domain by introducing composite spatial complexity, which combines a session's video spatial complexity with its time-variant spectral efficiency into a single resource demand function. To operate with commercial real-time video streaming applications that use rate-based congestion control and lack capability to measure QoE, we use external tooling for QoE measurements. We develop an incremental reallocation algorithm with per-interval limits that encode both the congestion control algorithm's speed constraint and that spatial complexity estimates are reliable only near the current rate. The resulting prototype combines external QoE measurements with congestion-signal-based rate steering and does not require modification to commercial applications. We chart an evolution path from this prototype toward full QoE-aware resource sharing, mapping emerging standards (IETF SCONE, CAMARA, Media over QUIC) to the progressive capabilities they enable. Szilveszter Nádas, Lars Ernström, Dan Druta, Igor Pruzhansky, David Lindero, Jonathan Lynam, Eric Petajan |
QoMEX | 7 |
| 2026 | Mapping the Effects of Resolution Scaling and Compression Level Between VMAF and P.1204.4
Eric Petajan, David Lindero |
QoMEX | 1 |
| 2025 | Automated Mobile Video Objective Testing SystemabstractApplying QoE analysis to optimize usage of cellular spectrum is of high interest to mobile network operators. A key challenge is to be able to perform QoE measurement across very different types of apps, from DASH VoD to interactive applications such as Video Conferencing and Cloud Gaming. This paper presents AMVOTS, a QoE measurement system developed by AT&T, which is flexible enough to support a large range of application types and network conditions. We also discuss using AMVOTS as part of a closed loop to prototype QoE-aware radio resource allocation. Eric Petajan, Jonathan Lynam, Morey Antebi, Hessam Moeini, David Lindero, Lars Ernström, P. Gyanesh Patra, Szilveszter Nádas |
QoMEX | 1 |
| 2020 | What you see is what you get: measure ABR video streaming QoE via on-device screen recordingabstractAnalyzing delivered QoE for Adaptive Bitrate (ABR) streaming over cellular networks is critical for a host of entities including content providers and mobile network providers. However, existing approaches mostly rely on network traffic analysis. In addition to potential accuracy issues, they are challenged by the increasing use of end-to-end network traffic encryption. In this paper, we explore a very different approach to QoE measurement --- utilizing the screen recording capability widely available on commodity devices to record the video displayed on the mobile device screen, and analyzing the recorded video to measure the delivered QoE. We design a novel system VideoEye to conduct such screen-recording-based QoE analysis. We identify the various technical challenges involved, including distortions introduced by the screen recording process that can make such analysis difficult. We develop techniques to accurately measure video QoE from the screen recordings even in the presence of recording distortions. Our evaluations demonstrate that VideoEye accurately detects important QoE indicators including the track played at different points in time, and stall statistics. The maximal error in detected stall duration is 0.5 s. The accuracy of detecting the displayed tracks is higher than 97%. Shichang Xu, Eric Petajan, Subhabrata Sen, Z. Morley Mao |
NOSSDAV | 2 |
| 2017 | BUFFEST: Predicting Buffer Conditions and Real-time Requirements of HTTP(S) Adaptive Streaming ClientsabstractStalls during video playback are perhaps the most important indicator of a client's viewing experience. To provide the best possible service, a proactive network operator may therefore want to know the buffer conditions of streaming clients and use this information to help avoid stalls due to empty buffers. However, estimation of clients' buffer conditions is complicated by most streaming services being rate-adaptive, and many of them also encrypted. Rate adaptation reduces the correlation between network throughput and client buffer conditions. Usage of HTTPS prevents operators from observing information related to video chunk requests, such as indications of rate adaptation or other HTTP-level information. Vengatanathan Krishnamoorthi, Niklas Carlsson, Emir Halepovic, Eric Petajan |
MMSys | 4 |
| 2003 | Introduction to the special issue on image-based modeling, rendering, and animation
Harry Shum, Eric Petajan, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | The communication of virtual human faces using MPEG-4 toolsabstractThe virtual human face is an important type of content which is efficiently represented and coded by the MPEG-4 Face and Body Animation specification. This paper addresses the basic functionality of the standard and describes a facial motion capture system which automatically generates MPEG-4 Face Animation Parameters from video of a real human. Face model design considerations and application scenarios are also covered. Eric Petajan |
ISCAS | 1 |
| 1997 | Acoustic driven viseme identification for face animationabstractUnlike other image templates, visemes have identities in two different media. In audio domain, they are often related to basic linguistic units such as phonemes. In image domain, they are defined by the images of human articulators, such as mouth shapes, chin movements, etc. In this paper, an approach of extracting visemes from both image and acoustic domains is presented. In image domain, the mouth shapes, represented by feature points on inner lip contours, are extracted through face tracking and mouth image analysis. In acoustic domain, viseme segments are obtained automatically by aligning phoneme strings to audio signals through a Viterbi alignment process. Jialin Zhong, Wu Chou, Eric Petajan |
MMSP | 3 |
| 1997 | MPEG-4: Audio/video and synthetic graphics/audio for mixed media
Peter K. Doenges, Tolga K. Çapin, Fabio Lavagetto, Jörn Ostermann, Igor S. Pandzic, Eric Petajan |
Signal Process. Image Commun. | 6 |
| 1996 | Multi-Modal System for Locating Heads and FacesabstractWe designed a modular system using a combination of shape analysis, color segmentation and motion information for locating reliably heads and faces of different sizes and orientations in complex images. The first of the system's three channels does a shape analysis on gray-level images to determine the location of individual facial features as well as the outlines of heads. In the second channel the color space is analyzed with a clustering algorithm to find areas of skin colors. The color space is first calibrated, using the results from the other channels. In the third channel motion information is extracted from frame differences. Head outlines are determined by analyzing the shapes of areas with large motion vectors. All three channels produce lists of shapes, each marking an area of the image where a facial feature or a part of the outline of a head may be present. Combinations of such shapes are evaluated with n-gram searches to produce a list of likely head positions and the locations of facial features. We tested the system for tracking faces of people sitting in front of terminals and video phones and used it to track people entering through a doorway. Hans Peter Graf, Eric Cosatto, David C. Gibbon, Michael Kocheisen, Eric Petajan |
FG | 5 |
| 1996 | Robust face feature analysis for automatic speechreading and character animationabstractThe robust acquisition of facial features needed for visual speech processing is fraught with difficulties which greatly increase the complexity of the machine vision system. This system must extract the inner lip contour from facial images with variations in pose, lighting, and facial hair. This paper describes a face feature acquisition system with robust performance in the presence of extreme lighting variations and moderate variations in pose. Furthermore, system performance is not degraded by facial hair or glasses. To find the position of a face reliably we search the whole image for facial features. These features are then combined and tests are applied, to determine whether any such combination actually belongs to a face. In order to find where the lips are, other features of the face, such as the eyes, must be located as well. Without this information it is difficult to reliably find the mouth in a complex image. Just the mouth by itself is easily missed or other elements in the image can be mistaken for a mouth. If camera position can be constrained to allow the nostrils to be viewed, then nostril tracking is used to both reduce computation and provide additional robustness. Once the nostrils are tracked from frame to frame using a tracking window the mouth area can be isolated and normalized for scale and rotation. A mouth detail analysis procedure is then used to estimate the inner lip contour and teeth and tongue regions. The inner lip contour and head movements are then mapped to synthetic face parameters to generate a graphical talking head synchronized with the original human voice. This information can also be used as the basis for visual speech features in an automatic speechreading system. Similar features were used in our previous automatic speechreading systems. Eric Petajan, Hans Peter Graf |
FG | 1 |
| 1995 | Speech-assisted lip synchronization in audio-visual communicationsabstractWe utilize speech information to improve the quality of audio-visual communications such as video telephony and videoconferencing. We show that the marriage of speech analysis and image processing can solve problems related to lip synchronization. We present a technique called speech-assisted frame-rate conversion, and apply it to coding of talking head video. Demonstration sequences are presented. Extensions and other applications are outlined. Tsuhan Chen, Hans Peter Graf, Barry G. Haskell, Eric Petajan, Yao Wang 0001, Homer H. Chen, Wu Chou |
ICIP | 4 |
| 1995 | The Grand Alliance system for US HDTVabstractThe US HDTV process has fostered substantial research and development activity over the last several years. The Advisory Committee on Advanced Television Service (ACATS) was formed to advise the FCC on the technology and systems suitable for delivery of high definition service over terrestrial broadcast channels. Four digital HDTV systems where tested at the Advanced Television Testing Center. All the systems gave excellent performance, but the results were inconclusive and a plan for a second round of tests was prepared. As each of the four systems where being readied for retest. The proponents of the four individual digital HDTV proposals worked together to define a single HDTV system which incorporated the best technology from the individual systems. The consortium of companies, called the Grand Alliance (GA), announced a combined system and submitted it to ACATS for consideration. After ACATS certification, the GA began construction of a prototype system to submit for laboratory testing at the end of 1994. This paper describes the video compression subsystem and the hardware prototype. The preprocessing, motion estimation, quantization, and rate control subsystems are described. The system uses bidirectional motion compensation, discrete cosine transform, quantization and Huffman coding. The resulting bitstream is input into a transport system which uses fixed length packets. The multiplex transport stream is input into the 8-VSB transmission system. Finally, the specifics of the hardware implementation are described and some simulation results are presented.> Kiran S. Challapali, Xavier Lebègue, Jae S. Lim, Woo H. Paik, Régis Saint-Girons, Eric Petajan, Vinay P. Sathe, Paul Snopko, Joel W. Zdepski |
Proc. IEEE | 6 |
| 1995 | The HDTV Grand Alliance systemabstractThe Grand Alliance (GA) was formed to define and construct a system for the delivery of HDTV using terrestrial broadcast channels. This system is composed of the best components from previously competing systems considered by the FCC. MPEG-2 syntax is used with novel encoding techniques to deliver a set of video scanning formats for a variety of applications. The paper focuses on video compression technology and also describes the important features and concepts embodied in the GA system as compared to the previous digital HDTV systems.> Eric Petajan |
Proc. IEEE | 1 |
| 1988 | An improved automatic lipreading system to enhance speech recognitionabstractCurrent acoustic speech recognition technology performs well with very small vocabularies in noise or with large vocabularies in very low noise. Accurate acoustic speech recognition in noise with vocabularies over 100 words has yet to be achieved. Humans frequently lipread the visible facial speech articulations to enhance speech recognition, especially when the acoustic signal is degraded by noise or hearing impairment. Automatic lipreading has been found to improve significantly acoustic speech recognition and could be advantageous in noisy environments such as offices, aircraft and factories. Eric Petajan, B. Bischoff, David Bodoff, N. M. Brooke |
CHI | 1 |