Oscar N. Garcia

dblp:77/6007 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4Artificial intelligence and machine learning · 3Theory of computation · 3 · 2 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 60% Audio and music processing · 40%
Artificial intelligence
2 papers
Speech recognition and synthesis · 44% Probabilistic and Bayesian machine learning · 44% Face, body and person analysis · 13%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer animation and physical simulation › facial animation
speech-driven facial animation
0.122005
Speech-driven facial animation with realistic dynamics · IEEE Trans. Multim. 2005
Audio/visual mapping with cross-modal hidden Markov models · IEEE Trans. Multim. 2005
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.112005
Audio/visual mapping with cross-modal hidden Markov models · IEEE Trans. Multim. 2005
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.112005
Audio/visual mapping with cross-modal hidden Markov models · IEEE Trans. Multim. 2005
Audio and music processing
audio-visual mapping
0.112005
Audio/visual mapping with cross-modal hidden Markov models · IEEE Trans. Multim. 2005
Computer animation and physical simulation
facial animation
0.112005
Speech-driven facial animation with realistic dynamics · IEEE Trans. Multim. 2005
Audio and music processing
speech processing
0.112005
Speech-driven facial animation with realistic dynamics · IEEE Trans. Multim. 2005
Computer vision › Face, body and person analysis › face modeling
facial performance capture
0.012005
Speech-driven facial animation with realistic dynamics · IEEE Trans. Multim. 2005
Automata and formal languages › formal grammars
context-free grammar
0.011983
Solution of an Open Problem on Probabilistic Grammars · IEEE Trans. Computers 1983
Automata and formal languages
probabilistic grammars
0.011983
Solution of an Open Problem on Probabilistic Grammars · IEEE Trans. Computers 1983
Coding theory › error-correcting codes
arithmetic codes
0.011971
Cyclic and multiresidue codes for arithmetic operations · IEEE Trans. Inf. Theory 1971
Coding theory › error-correcting codes › arithmetic codes
AN codes
0.011971
Cyclic and multiresidue codes for arithmetic operations · IEEE Trans. Inf. Theory 1971
Coding theory › error-correcting codes
single-error-correcting codes
0.011971
Cyclic and multiresidue codes for arithmetic operations · IEEE Trans. Inf. Theory 1971

Methods — techniques the papers use, named apart from their topics

pseudo-muscle model · 0.1nearest-neighbor prediction · 0.1least-mean-squared HMM · 0.1karhunen-loève transform · 0.1cross-modal hidden markov models · 0.1HMM inversion · 0.1relative frequency estimation · 0.0residue arithmetic · 0.0number theory · 0.0
YearPublicationVenuePosition
2006 A comparison of acoustic coding models for speech-driven facial animation
Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Ricardo Gutierrez-Osuna
Speech Commun.3
2005 Audio/visual mapping with cross-modal hidden Markov models
abstract
The audio/visual mapping problem of speech-driven facial animation has intrigued researchers for years. Recent research efforts have demonstrated that hidden Markov model (HMM) techniques, which have been applied successfully to the problem of speech recognition, could achieve a similar level of success in audio/visual mapping problems. A number of HMM-based methods have been proposed and shown to be effective by the respective designers, but it is yet unclear how these techniques compare to each other on a common test bed. In this paper, we quantitatively compare three recently proposed cross-modal HMM methods, namely the remapping HMM (R-HMM), the least-mean-squared HMM (LMS-HMM), and HMM inversion (HMMI). The objective of our comparison is not only to highlight the merits and demerits of different mapping designs, but also to study the optimality of the acoustic representation and HMM structure for the purpose of speech-driven facial animation. This paper presents a brief overview of these models, followed by an analysis of their mapping capabilities on a synthetic dataset. An empirical comparison on an experimental audio-visual dataset consisting of 75 TIMIT sentences is finally presented. Our results show that HMMI provides the best performance, both on synthetic and experimental audio-visual data.
Shengli Fu, Ricardo Gutierrez-Osuna, Anna Esposito, Praveen K. Kakumanu, Oscar N. Garcia
IEEE Trans. Multim.5
2005 Speech-driven facial animation with realistic dynamics
abstract
This work presents an integral system capable of generating animations with realistic dynamics, including the individualized nuances, of three-dimensional (3-D) human faces driven by speech acoustics. The system is capable of capturing short phenomena in the orofacial dynamics of a given speaker by tracking the 3-D location of various MPEG-4 facial points through stereovision. A perceptual transformation of the speech spectral envelope and prosodic cues are combined into an acoustic feature vector to predict 3-D orofacial dynamics by means of a nearest-neighbor algorithm. The Karhunen-Loe/spl acute/ve transformation is used to identify the principal components of orofacial motion, decoupling perceptually natural components from experimental noise. We also present a highly optimized MPEG-4 compliant player capable of generating audio-synchronized animations at 60 frames/s. The player is based on a pseudo-muscle model augmented with a nonpenetrable ellipsoidal structure to approximate the skull and the jaw. This structure adds a sense of volume that provides more realistic dynamics than existing simplified pseudo-muscle-based approaches, yet it is simple enough to work at the desired frame rate. Experimental results on an audiovisual database of compact TIMIT sentences are presented to illustrate the performance of the complete system.
Ricardo Gutierrez-Osuna, Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Adriana Bojórquez, José Luis Castillo, Isaac Rudomín
IEEE Trans. Multim.4
2000 Detecting and Tracking Human Faces in Videos
abstract
A method for detecting and tracking human faces in color videos is presented. The method first uses a chroma chart with information about skin colors of various races to determine regions of skin color in the first frame of a video. A new chroma chart is computed for each region, which more precisely represents the colour contents of that region. Chroma charts for different regions that are similar are combined, while those that are considerably different are kept separate. Model facial patterns are then used to detect faces within the skin regions. Once a face is detected, the particular pattern and color of the face are used to track the face. Regions where facial patterns are not detected are expected to correspond to exposed parts of the body or of the background and are ignored. The proposed method can track faces with a high degree of accuracy once they are identified.
A. Ardeshir Goshtasby, Oscar N. Garcia
ICPR3
1997 A New Methodology for Optimizing Evasive Maneuvers Under Uncertainty in the Extended Two-Dimensional Pursuer/Evader Problem
abstract
Traditional analytic or control-theoretic solutions to the problem of optimizing evasive maneuvers in the extended two-dimensional pursuer/evader problem require the evader to execute specific sequences of maneuvers at precise pursuer/evader distances. These solutions depend upon several pursuer-specific characteristics, and fail to effectively account for uncertainty about the state of the pursuer. This paper describes the implementation of a genetic programming system that evolves optimized solutions to the extended two-dimensional pursuer/evader problem that do not depend upon knowledge of the pursuer's current state. Best-of-run programs execute strategies by which an evader may maneuver to successfully evade a pursuer starting from a wide range of relative initial positions, under conditions where the state of the pursuer is unknown or uncertain.
Frank W. Moore, Oscar N. Garcia
ICTAI2
1995 A Cognitive Framework of Debugging
Byung-do Yoon, Oscar N. Garcia
SEKE2
1989 Foreword - Knowledge and Data Engineering: An Outlook
Oscar N. Garcia
IEEE Trans. Knowl. Data Eng.1
1983 Solution of an Open Problem on Probabilistic Grammars
abstract
It has been proved that when the production probabilities of an unambiguous context-free grammar G are estimated by the relative frequencies of the corresponding productions in a sample S from the language L(G) generated by G, the expected derivation length and the expected word length of the words in L(G) are precisely equal to the mean derivation length and the mean world length of the words in the same S, respectively.
Ranjan Chaudhuri, Son Pham, Oscar N. Garcia
IEEE Trans. Computers3
1978 An approximate and empirical study of the distribution of adder inputs and maximum carry length propagation
abstract
This paper investigates, using sampled data, the commonly used hypothesis that integer operands reaching the adder of a computer are uniformly distributed. Questions raised on the validity of that hypothesis are reinforced and their impact on the calculation of the average of the worst case length of carry propagation is considered. An approximate formula is developed for the worst case carry chain length when the arithmetic operands are restricted in magnitude.
Oscar N. Garcia, Harvey Glass, Stanley C. Haimes
IEEE Symposium on Computer Arithmetic1
1972 Arithmetic codes in a module: A majority decoding approach
abstract
It has been shown how arithmetic codes can be embedded in the structure of a module and how generator and check matrices may be defined. The motivation is to parallel the analogy with communication codes with the purpose of showing majority decoding. The technical difficulties in the derivations are pointed out while some partial results are given.
Oscar N. Garcia, Jorge R. Rodriguez
IEEE Symposium on Computer Arithmetic1
1971 Cyclic and multiresidue codes for arithmetic operations
abstract
In this paper, the cyclic nature ofANcodes is defined after a brief summary of previous work in this area is given. New results are shown in the determination of the range for single-error-correctingANcodes whenAis the product of two odd primesp_1andp_2, given the orders of 2 modulop_1and modulop_2. The second part of the paper treats a more practical class of arithmetic codes known as separate codes. A generalized separate code, called a multiresidue code, is one in which a numberNis represented as \begin{equation} [N, \mid N \mid _ {m1}, \mid N \mid _{m2}, \cdots , \mid N \mid _{mk}] \end{equation} wherem_iare pairwise relatively prime integers. For eachANcode, whereAis composite, a multiresidue code can be derived having error-correction properties analogous to those of theANcode. Under certain natural constraints, multiresidue codes of large distance and large range (i.e., large values ofN) can be implemented. This leads to possible realization of practical single and/or multiple-error-correcting arithmetic units.
T. R. N. Rao, Oscar N. Garcia
IEEE Trans. Inf. Theory2