EDBT 2026 Demo / reviewers in the wild / expert
Eugene Ie
dblp:32/3393
· DBLP profile ↗
18ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Vision and language · 34% Autonomous driving · 18% Generative modeling · 16% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 61% Knowledge graphs · 35% Data mining · 4% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-and-language navigation |
2.1 | 5 | 2020 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020 Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation · ECCV (24) 2020 BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps · ACL 2020 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Improving Rectified Flow with Boundary Conditions · ICCV 2025 |
Machine learning › Generative modeling › diffusion model
rectified flow |
0.9 | 1 | 2025 | Improving Rectified Flow with Boundary Conditions · ICCV 2025 |
Robotics › Autonomous driving
pedestrian behavior prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Robotics › Autonomous driving › trajectory prediction
pedestrian trajectory prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Robotics › Autonomous driving
trajectory prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Computer vision › Vision and language
cross-modal retrieval |
0.4 | 1 | 2020 | Learning to Represent Image and Text with Denotation Graph · EMNLP (1) 2020 |
Machine learning › Reinforcement learning
curriculum reinforcement learning |
0.4 | 1 | 2020 | BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps · ACL 2020 |
Computer vision › Vision and language
image-text retrieval |
0.4 | 1 | 2020 | Learning to Represent Image and Text with Denotation Graph · EMNLP (1) 2020 |
Machine learning › Learning paradigms
multi-task learning |
0.4 | 1 | 2020 | Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation · ECCV (24) 2020 |
Computer vision › Vision and language › multimodal representation
vision-language representation learning |
0.4 | 1 | 2020 | Learning to Represent Image and Text with Denotation Graph · EMNLP (1) 2020 |
Knowledge graphs
knowledge graph construction |
0.4 | 1 | 2020 | Learning to Represent Image and Text with Denotation Graph · EMNLP (1) 2020 |
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning |
0.4 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Machine learning › Representation and self-supervised learning › transferable representation
transferable representation learning |
0.4 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Recommender systems
reinforcement-learning-based recommendation |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Recommender systems › interactive recommendation
slate recommendation |
0.4 | 1 | 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets · IJCAI 2019 |
Natural language and speech › Language models and text generation
instruction following |
0.2 | 2 | 2020 | BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps · ACL 2020 Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Computer vision › Vision and language › visual grounding
language grounding |
0.1 | 1 | 2020 | Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation · ECCV (24) 2020 |
Computer vision › 3D vision
photorealistic simulation |
0.1 | 1 | 2020 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding · EMNLP (1) 2020 |
Bioinformatics and computational biology › protein function prediction
protein classification |
0.1 | 2 | 2007 | Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007 Semi-supervised protein classification using cluster kernels · Bioinform. 2005 |
Robotics › Robot navigation and mapping
embodied navigation |
0.1 | 1 | 2019 | Transferable Representation Learning in Vision-and-Language Navigation · ICCV 2019 |
Machine learning › Learning theory › classification › multiclass classification
error-correcting output codes |
0.1 | 1 | 2007 | Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007 |
Machine learning › Learning theory › classification
multiclass classification |
0.1 | 1 | 2007 | Multi-class Protein Classification Using Adaptive Codes · J. Mach. Learn. Res. 2007 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.1 | 1 | 2005 | Semi-supervised protein classification using cluster kernels · Bioinform. 2005 |
Bioinformatics and computational biology › protein structure prediction › template-based modeling
fold recognition |
0.1 | 1 | 2005 | Multi-class protein fold recognition using adaptive codes · ICML 2005 |
Bioinformatics and computational biology
protein structure prediction |
0.1 | 1 | 2005 | Multi-class protein fold recognition using adaptive codes · ICML 2005 |
Data mining › predictive modeling › classification
multiclass classification |
0.1 | 1 | 2005 | Multi-class protein fold recognition using adaptive codes · ICML 2005 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 1.5boundary conditions · 0.9SDE sampling · 0.9ODE sampling · 0.9contrastive learning · 0.7spatiotemporal grounding · 0.4multimodal pre-training · 0.4memory buffer · 0.4linguistic analysis · 0.4imitation learning · 0.4environment-agnostic learning · 0.4curriculum-based reinforcement learning · 0.4attention layers · 0.4temporal difference learning · 0.4q-learning · 0.4one-vs-all classifiers · 0.1nearest neighbor · 0.1PSI-BLAST · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Rectified Flow with Boundary ConditionsabstractRectified Flow offers a simple and effective approach to high-quality generative modeling by learning a velocity field. However, we identify a limitation in directly modeling the velocity with an unconstrained neural network: the learned velocity often fails to satisfy certain boundary conditions, leading to inaccurate velocity field estimations that deviate from the desired ODE. This issue is particularly critical during stochastic sampling at inference, as the score function's errors are amplified near the boundary. To mitigate this, we propose a Boundary-enforced Rectified Flow Model (Boundary RF Model), in which we enforce boundary conditions with a minimal code modification. Boundary RF Model improves performance over vanilla RF model, demonstrating 8.01% improvement in FID score on ImageNet using ODE sampling and 8.98% improvement using SDE sampling. Xixi Hu 0001, Runlong Liao, Keyang Xu, Bo Liu 0042, Yeqing Li, Eugene Ie, Hongliang Fei, Qiang Liu 0001 |
ICCV | 6 |
| 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human KeypointsabstractAccurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory pre-diction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study. Jiachen Li 0001, Xinwei Shi, Jonathan Stroud, Zhishuai Zhang, Junhua Mao, Jeonhyung Kang, Khaled S. Refaat, Weilong Yang, Eugene Ie |
ICRA | 11 |
| 2021 | On the Evaluation of Vision-and-Language Navigation InstructionsabstractMing Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, Eugene Ie. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Vihan Jain, Su Wang 0001, Alexander Ku, Jason Baldridge, Eugene Ie |
EACL | 7 |
| 2021 | CoMSum and SIBERT: A Dataset and Neural Model for Query-Based Multi-document Summarization
Sayali Kulkarni, Sheide Chammas, Wan Zhu, Fei Sha, Eugene Ie |
ICDAR (2) | 5 |
| 2020 | BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby StepsabstractLearning to follow instructions is of fundamental importance to autonomous agents for vision-and-language navigation (VLN).In this paper, we study how an agent can navigate long paths when learning from a corpus that consists of shorter ones.We show that existing state-of-the-art agents do not generalize well.To this end, we propose BabyWalk, a new VLN agent that is learned to navigate by decomposing long instructions into shorter ones (BabySteps) and completing them sequentially.A special design memory buffer is used by the agent to turn its past experiences into contexts for future steps.The learning process is composed of two phases.In the first phase, the agent uses imitation learning from demonstration to accomplish BabySteps.In the second phase, the agent uses curriculum-based reinforcement learning to maximize rewards on navigation tasks with increasingly longer instructions.We create two new benchmark datasets (of long navigation tasks) and use them in conjunction with existing ones to examine BabyWalk's generalization ability.Empirical results show that BabyWalk achieves state-of-the-art results on several metrics, in particular, is able to follow long instructions better.The codes and the datasets are released on our project page https://github.com/Sha-Lab/babywalk. Wang Zhu 0001, Hexiang Hu, Zhiwei Deng, Vihan Jain, Eugene Ie, Fei Sha |
ACL | 6 |
| 2020 | Environment-Agnostic Multitask Learning for Natural Language Grounded Navigation
Xin Wang 0061, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, Sujith Ravi |
ECCV (24) | 3 |
| 2020 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingabstractWe introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets.It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities.Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators.We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations (Anderson et al., 2018b).We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations.The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments. Alexander Ku, Roma Patel, Eugene Ie, Jason Baldridge |
EMNLP (1) | 4 |
| 2020 | Learning to Represent Image and Text with Denotation GraphabstractLearning to fuse vision and language information and representing them is an important research problem with many applications.Recent progresses have leveraged the ideas of pretraining (from language modeling) and attention layers in Transformers to learn representation from datasets containing images aligned with linguistic expressions that describe the images.In this paper, we propose learning representations from a set of implied, visually grounded expressions between image and text, automatically mined from those datasets.In particular, we use denotation graphs to represent how specific concepts (such as sentences describing images) can be linked to abstract and generic concepts (such as short phrases) that are also visually grounded.This type of generic-to-specific relations can be discovered using linguistic analysis tools.We propose methods to incorporate such relations into learning representation.We show that state-of-the-art multimodal learning models can be further improved by leveraging automatically harvested structural relations.The representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition.Both our codes and the extracted denotation graphs on the Flickr30K and the COCO datasets are publically available on Bowen Zhang 0002, Hexiang Hu, Vihan Jain, Eugene Ie, Fei Sha |
EMNLP (1) | 4 |
| 2020 | Demonstrating Principled Uncertainty Modeling for Recommender Ecosystems with RecSim NGabstractWe develop RecSim NG, a probabilistic platform that supports natural, concise specification and learning of models for multi-agent recommender systems simulation. RecSim NG is a scalable, modular, differentiable simulator implemented in Edward2 and TensorFlow. Martin Mladenov, Vihan Jain, Eugene Ie, Christopher Colby, Nicolas Mayoraz, Hubert Pham, Dustin Tran, Ivan Vendrov, Craig Boutilier |
RecSys | 4 |
| 2019 | Stay on the Path: Instruction Fidelity in Vision-and-Language NavigationabstractAdvances in learning and representations have reinvigorated work that connects language to other modalities.A particularly exciting direction is Vision-and-Language Navigation (VLN), in which agents interpret natural language instructions and visual scenes to move through environments and reach goals.Despite recent progress, current research leaves unclear how much of a role language understanding plays in this task, especially because dominant evaluation metrics have focused on goal completion rather than the sequence of actions corresponding to the instructions.Here, we highlight shortcomings of current metrics for the Room-to-Room dataset (Anderson et al., 2018b) and propose a new metric, Coverage weighted by Length Score (CLS).We also show that the existing paths in the dataset are not ideal for evaluating instruction following because they are direct-to-goal shortest paths.We join existing short paths to form more challenging extended paths to create a new data set, Room-for-Room (R4R).Using R4R and CLS, we show that agents that receive rewards for instruction fidelity outperform agents that focus on goal completion. Vihan Jain, Gabriel Ilharco, Alexander Ku, Ashish Vaswani, Eugene Ie, Jason Baldridge |
ACL (1) | 5 |
| 2019 | Learning Dense Representations for Entity RetrievalabstractDaniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, Diego Garcia-Olano. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019. Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, Diego Garcia-Olano |
CoNLL | 6 |
| 2019 | Transferable Representation Learning in Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task requires competence in several perception problems: successful agents combine spatio-temporal, vision and language understanding to produce appropriate action sequences. Our approach adapts pre-trained vision and language representations to relevant in-domain tasks making them more effective for VLN. Specifically, the representations are adapted to solve both a cross-modal sequence alignment and sequence coherence task. In the sequence alignment task, the model determines whether an instruction corresponds to a sequence of visual frames. In the sequence coherence task, the model determines whether the perceptual sequences are predictive sequentially in the instruction-conditioned latent space. By transferring the domain-adapted representations, we improve competitive agents in R2R as measured by the success rate weighted by path length (SPL) metric. Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Ilharco, Jason Baldridge, Eugene Ie |
ICCV | 7 |
| 2019 | SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation SetsabstractReinforcement learning methods for recommender systems optimize recommendations for long-term user engagement. However, since users are often presented with slates of multiple items---which may have interacting effects on user choice---methods are required to deal with the combinatorics of the RL action space. We develop SlateQ, a decomposition of value-based temporal-difference and Q-learning that renders RL tractable with slates. Under mild assumptions on user choice behavior, we show that the long-term value (LTV) of a slate can be decomposed into a tractable function of its component item-wise LTVs. We demonstrate our methods in simulation, and validate the scalability and effectiveness of decomposed TD-learning on YouTube. Eugene Ie, Vihan Jain, Sanmit Narvekar, Ritesh Agarwal, Rui Wu 0020, Heng-Tze Cheng, Tushar Chandra, Craig Boutilier |
IJCAI | 1 |
| 2011 | Translation-Inspired OCRabstractOptical character recognition is carried out using techniques borrowed from statistical machine translation. In particular, the use of multiple simple feature functions in linear combination, along with minimum-error-rate training, integrated decoding, and N-gram language modeling is found to be remarkably effective, across several scripts and languages. Results are presented using both synthetic and real data in five languages. Dmitriy Genzel, Ashok C. Popat, Nemanja Spasojevic, Michael Jahr, Andrew W. Senior, Eugene Ie, Frank Yung-Fong Tang |
ICDAR | 6 |
| 2007 | SVM-Fold: a tool for discriminative multi-class protein fold and superfamily recognitionabstractBACKGROUND: Predicting a protein's structural class from its amino acid sequence is a fundamental problem in computational biology. Much recent work has focused on developing new representations for protein sequences, called string kernels, for use with support vector machine (SVM) classifiers. However, while some of these approaches exhibit state-of-the-art performance at the binary protein classification problem, i.e. discriminating between a particular protein class and all other classes, few of these studies have addressed the real problem of multi-class superfamily or fold recognition. Moreover, there are only limited software tools and systems for SVM-based protein classification available to the bioinformatics community. RESULTS: We present a new multi-class SVM-based protein fold and superfamily recognition system and web server called SVM-Fold, which can be found at http://svm-fold.c2b2.columbia.edu. Our system uses an efficient implementation of a state-of-the-art string kernel for sequence profiles, called the profile kernel, where the underlying feature representation is a histogram of inexact matching k-mer frequencies. We also employ a novel machine learning approach to solve the difficult multi-class problem of classifying a sequence of amino acids into one of many known protein structural classes. Binary one-vs-the-rest SVM classifiers that are trained to recognize individual structural classes yield prediction scores that are not comparable, so that standard "one-vs-all" classification fails to perform well. Moreover, SVMs for classes at different levels of the protein structural hierarchy may make useful predictions, but one-vs-all does not try to combine these multiple predictions. To deal with these problems, our method learns relative weights between one-vs-the-rest classifiers and encodes information about the protein structural hierarchy for multi-class prediction. In large-scale benchmark results based on the SCOP database, our code weighting approach significantly improves on the standard one-vs-all method for both the superfamily and fold prediction in the remote homology setting and on the fold recognition problem. Moreover, our code weight learning algorithm strongly outperforms nearest-neighbor methods based on PSI-BLAST in terms of prediction accuracy on every structure classification problem we consider. CONCLUSION: By combining state-of-the-art SVM kernel methods with a novel multi-class algorithm, the SVM-Fold system delivers efficient and accurate protein fold and superfamily recognition. Iain Melvin, Eugene Ie, Rui Kuang, Jason Weston, William Stafford Noble, Christina S. Leslie |
BMC Bioinform. | 2 |
| 2007 | Multi-class Protein Classification Using Adaptive Codes
Iain Melvin, Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie |
J. Mach. Learn. Res. | 2 |
| 2005 | Multi-class protein fold recognition using adaptive codesabstractWe develop a novel multi-class classification method based on output codes for the problem of classifying a sequence of amino acids into one of many known protein structural classes, called folds. Our method learns relative weights between one-vs-all classifiers and encodes information about the protein structural hierarchy for multi-class prediction. Our code weighting approach significantly improves on the standard one-vs-all method for the fold recognition problem. In order to compare against widely used methods in protein sequence analysis, we also test nearest neighbor approaches based on the PSI-BLAST algorithm. Our code weight learning algorithm strongly outperforms these PSI-BLAST methods on every structure recognition problem we consider. 1. Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie |
ICML | 1 |
| 2005 | Semi-supervised protein classification using cluster kernelsabstractMOTIVATION: Building an accurate protein classification system depends critically upon choosing a good representation of the input sequences of amino acids. Recent work using string kernels for protein data has achieved state-of-the-art classification performance. However, such representations are based only on labeled data--examples with known 3D structures, organized into structural classes--whereas in practice, unlabeled data are far more plentiful. RESULTS: In this work, we develop simple and scalable cluster kernel techniques for incorporating unlabeled data into the representation of protein sequences. We show that our methods greatly improve the classification performance of string kernels and outperform standard approaches for using unlabeled data, such as adding close homologs of the positive examples to the training data. We achieve equal or superior performance to previously presented cluster kernel methods and at the same time achieving far greater computational efficiency. AVAILABILITY: Source code is available at www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot. The Spider matlab package is available at www.kyb.tuebingen.mpg.de/bs/people/spider. SUPPLEMENTARY INFORMATION: www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot. Jason Weston, Christina S. Leslie, Eugene Ie, Dengyong Zhou, André Elisseeff, William Stafford Noble |
Bioinform. | 3 |