VLDB 2026 Research / reviewers in the wild / expert
Eric Yeh
dblp:70/1842
· DBLP profile ↗
12ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-8752-4429ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Frontier MoE Training with 3D Integrated OpticsabstractThe unabated growth in AI workload demands is driving the need for concerted advances in compute, memory, and interconnect performance. As traditional semiconductor scaling slows, high-speed interconnects have emerged as the new scaling engine, enabling the creation of larger logical GPUs by linking many GPUs into a single, low-latency, high-bandwidth compute domain. While initial scale-up fabrics leveraged copper interconnects for their power and cost advantages, the maximum reach of passive electrical interconnects (approximately 1 meter) effectively limits the scale-up domain to within a single rack. The advent of 3D-stacked optics and logic offers a transformative, power-efficient scale-up solution for connecting hundreds of GPU packages (thousands of GPUs) across multiple data center racks. This work explores the design tradeoffs of scale-up tech-nologies and demonstrates how frontier LLMs necessitate novel photonic solutions to achieve aggressive power and performance targets. We model the benefits of 3D CPO (Passage) enabled GPUs and switches within the scale-up domain when training Frontier Mixture of Experts (MoE) models exceeding one trillion parameters. Our results show that the substantial increases in bandwidth and radix enabled by 3D CPO allow for an 8X increase in scale-up capability. This affords new opportunities for multi-dimensional parallelism within the scale-up domain and results in a 2.7X reduction in time-to-train, unlocking unprecedented model scaling. Mikhail Bernadskiy, Peter Carson, Thomas Graham, Taylor Groves, Ho John Lee, Eric Yeh |
HOTI | 6 |
| 2023 | An Augmented MetiTarski Dataset for Real Quantifier Elimination Using Machine Learning
John Hester, Briland Hitaj, Grant Olney Passmore, Sam Owre, Natarajan Shankar, Eric Yeh |
CICM | 6 |
| 2023 | CoProver: A Recommender System for Proof Construction
Eric Yeh, Briland Hitaj, Sam Owre, Maena Quemener, Natarajan Shankar |
CICM | 1 |
| 2016 | Diagnosis cloud: Sharing knowledge across cellular networksabstractDiagnosis functionality as a key component for automated Network Management (NM) systems allows rapid, machine-level interpretation of acquired data. In existing work, network diagnosis has focused on building “point solutions” using configuration and performance management, alarm, and topology information from one network. While the use of automated anomaly detection and diagnosis techniques within a single network improves operational efficiency, the knowledge learned by running these techniques across different networks that are managed by the same operator can be further maximized when that knowledge is shared. This paper presents a novel diagnosis cloud framework that enables the extraction and transfer of knowledge from one network to another. It also presents use cases and requirements. We present the implementation details of the diagnosis cloud framework for two specific types of models: topic models and Markov Logic Networks (MLNs). For each, we describe methods for assessing the quality of the local model, ranking models, adapting models to a new network, and performing detection and diagnosis. We performed experiments for the diagnosis cloud framework using real cellular network datasets. Our experiments demonstrate the feasibility of sharing topic models and MLNs. Gabriela F. Ciocarlie, Cherita Corbett, Eric Yeh, Christopher Connolly, Henning Sanneck, Muhammad Naseer ul Islam, Borislava Gajic, Szabolcs Nováczki, Kimmo Hätönen |
CNSM | 3 |
| 2016 | An Annotated Corpus and Method for Analysis of Ad-Hoc Structures Embedded in Text
Eric Yeh, John Niekrasz, Dayne Freitag, Richard Rohwer |
LREC | 1 |
| 2014 | Late fusion and calibration for multimedia event detection using few examplesabstractThe state-of-the-art in example-based multimedia event detection (MED) rests on heterogeneous classifiers whose scores are typically combined in a late-fusion scheme. Recent studies on this topic have failed to reach a clear consensus as to whether machine learning techniques can outperform rule-based fusion schemes with varying amount of training data. In this paper, we present two parametric approaches to late fusion: a normalization scheme for arithmetic mean fusion (logistic averaging) and a fusion scheme based on logistic regression, and compare them to widely used rule-based fusion schemes. We also describe how logistic regression can be used to calibrate the fused detection scores to predict an optimal threshold given a detection prior and costs on errors. We discuss the advantages and shortcomings of each approach when the amount of positives available for training varies from 10 positives (10Ex) to 100 positives (100Ex). Experiments were run using video data from the NIST TRECVID MED 2013 evaluation and results were reported in terms of a ranking metric: the mean average precision (mAP) and R0, a cost-based metric introduced in TRECVID MED 2013. Julien van Hout, Eric Yeh, Dennis C. Koelma, Cees Snoek, Chen Sun 0002, Ramakant Nevatia, Julie Wong, Gregory K. Myers |
ICASSP | 2 |
| 2014 | ISOMER: Informative Segment Observations for Multimedia Event RecountingabstractThis paper describes a system for multimedia event detection and recounting. The goal is to detect a high level event class in unconstrained web videos and generate event oriented summarization for display to users. For this purpose, we detect informative segments and collect observations for them, leading to our ISOMER system. We combine a large collection of both low level and semantic level visual and audio features for event detection. For event recounting, we propose a novel approach to identify event oriented discriminative video segments and their descriptions with a linear SVM event classifier. User friendly concepts including objects, actions, scenes, speech and optical character recognition are used in generating descriptions. We also develop several mapping and filtering strategies to cope with noisy concept detectors. Our system performed competitively in the TRECVID 2013 Multimedia Event Detection task with near 100,000 videos and was the highest performer in TRECVID 2013 Multimedia Event Recounting task. Chen Sun 0002, J. Brian Burns, Ramakant Nevatia, Cees Snoek, Robert C. Bolles, Gregory K. Myers, Wen Wang 0001, Eric Yeh |
ICMR | 8 |
| 2013 | Extracting spoken and acoustic concepts for multimedia event detectionabstractBecause of the popularity of online videos, there has been much interest in recent years in audio processing for the improvement of online video search. In this paper, we explore using acoustic concepts and spoken concepts extracted via audio segmentation/recognition and speech recognition respectively for Multimedia Event Detection (MED). To extract spoken concepts, a segmenter trained on annotated data from user videos segments the audio into three classes: speech, music, and other sounds. The speech segments are passed to an Automatic Speech Recognition (ASR) engine, and words from the 1-best ASR output, as well as posterior-weighted word counts collected from ASR lattices, are used as features to an SVM based classifier. Acoustic concepts are extracted using the 3-gram lattice counts of two Acoustic Concept Recognition (ACR) systems trained on 7 broad classes. MED results are reported on a subset of the NIST 2011 TRECVID data. We find that spoken concepts using lattices yield a 15% relative improvement in Average Pmiss (APM) over 1-best based features. Further, the proposed spoken concepts gave a 30% relative gain in APM over the ACR-based MED system using 7 classes. Lastly, we obtain an 8% relative APM improvement after score-level fusion of both concept types, showing the effective coupling of both approaches. Julien van Hout, Murat Akbacak, Diego Castán, Eric Yeh, Michelle Sanchez |
ICASSP | 4 |
| 2011 | Learning to ask the right questions to help a learner learnabstractIntelligent systems require substantial bodies of problem-solving knowledge. Machine learning techniques hold much appeal for acquiring such knowledge but typically require extensive amounts of user-supplied training data. Alternatively, informed question asking can supplement machine learning by directly eliciting critical knowledge from a user. Question asking can reduce the amount of training data required, and hence the burden on the user; furthermore, focused question asking holds significant promise for faster and more accurate acquisition of knowledge. In previous work, we developed static strategies for question asking that provide background knowledge for a base learner, enabling the learner to make useful generalizations even with few training examples. Here, we extend that work with a learning approach for automatically acquiring question-asking strategies that better accommodate the interdependent nature of questions. We present experiments validating the approach and showing its usefulness for acquiring efficient, context-dependent question-asking strategies. Melinda T. Gervasio, Eric Yeh, Karen L. Myers |
IUI | 2 |
| 2009 | Efficient Online Learning and Prediction of Users' Desktop Actions
Omid Madani, Hung Hai Bui, Eric Yeh |
IJCAI | 3 |
| 2006 | Reducing Human Fatigue in Interactive Evolutionary Computation Through Fuzzy Systems and Machine Learning SystemsabstractWe describe two approaches to reducing human fatigue in interactive evolutionary computation (IEC). A predictor function is used to estimate the human user's score, thus reducing the amount of effort required by the human user during the evolution process. The fuzzy system and four machine learning classifier algorithms are presented. Their performance in a real-world application, the IEC-based design of a micromachine resonating mass, is evaluated. The fuzzy system was composed of four simple rules, but was able to accurately predict the user's score 77% of the time on average. This is equivalent to a 51 % reduction of human effort compared to using IEC without the predictor. The four machine learning approaches tested were k-nearest neighbors, decision tree, AdaBoosted decision tree, and support vector machines. These approaches achieved good accuracy on validation tests, but because of the great diversity in user scoring behavior, were unable to achieve equivalent results on the user test data. Raffi R. Kamalian, Eric Yeh, Alice M. Agogino, Hideyuki Takagi |
FUZZ-IEEE | 2 |
| 2005 | Angler: a collaborative tool for cognitive expansionabstractHumans tend to bias the analysis of situations based on their previous experiences and back-ground. Angler is a tool to help analysts explore, understand, and overcome such biases and to collaborate in expanding their joint cognitive vision. Angler utilizes divergent and convergent techniques, such as brainstorming and clustering or voting, to guide a diverse set of intelligence professionals in completing a complex knowledge task. The tool helps the group through the process of forming consensus, while preserving and quantifying differing ways of thinking. Angler provides a Web-based collaborative environment that allows users distributed by both time and geography to assemble in teams, with the help of a facilitator. Andres C. Rodriguez, Thomas Boyce, John D. Lowrance, Eric Yeh |
K-CAP | 4 |