Benjamin Han

dblp:23/5290 · DBLP profile ↗
← Back
13ranked-venue papers
11as first author
5since 2021 · last 2024
0000-0002-2350-7280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Construction of Paired Knowledge Graph - Text Datasets Informed by Cyclic Evaluation
abstract
Datasets that pair Knowledge Graphs (KG) and text together (KG-T) can be used to train forward and reverse neural models that generate text from KG and vice versa. However models trained on datasets where KG and text pairs are not equivalent can suffer from more hallucination and poorer recall. In this paper, we verify this empirically by generating datasets with different levels of noise and find that noisier datasets do indeed lead to more hallucination. We argue that the ability of forward and reverse models trained on a dataset to cyclically regenerate source KG or text is a proxy for the equivalence between the KG and the text in the dataset. Using cyclic evaluation we find that manually created WebNLG is much better than automatically created TeKGen and T-REx. Informed by these observations, we construct a new, improved dataset called LAGRANGE using heuristics meant to improve equivalence between KG and text and show the impact of each of the heuristics on cyclic evaluation. We also construct two synthetic datasets using large language models (LLMs), and observe that these are conducive to models that perform significantly well on cyclic generation of text, but less so on cyclic generation of KGs, probably because of a lack of a consistent underlying ontology.
Ali Mousavi 0003, Xin Zhan, He Bai 0002, Peng Shi 0010, Theodoros Rekatsinas, Benjamin Han, Yunyao Li 0001, Jeffrey Pound, Joshua M. Susskind, Natalie Schluter, Ihab F. Ilyas, Navdeep Jaitly
LREC/COLING6
2022 Real-Time Rideshare Driver Supply Values Using Online Reinforcement Learning
abstract
In this paper, we present Online Supply Values (OSV), a system for estimating the return of available rideshare drivers to match drivers to ride requests at Lyft. Because a future driver state can be accurately predicted from a request destination, it is possible to estimate the expected action value of assigning a ride request to an available driver as a Markov Decision Process using the Bellman Equation. These estimates are updated using temporal difference and are shown to adapt to changing marketplace conditions in real-time. While reinforcement learning has been studied for rideshare dispatch, fully-online approaches without offline priors or other guardrails had never been evaluated in the real world. This work presents the algorithmic changes needed to bridge this gap. OSV is now deployed globally as a core component of Lyft's dispatch matching system. Our A/B user experiments in major US cities measure a +(0.96±0.53)% increase in the request fulfillment rate and a +(0.73±0.22)% increase to profit per passenger session over the previous algorithm.
Benjamin Han, Hyungjun Lee, Sébastien Martin
KDD1
2022 DI-2022: The Third Document Intelligence Workshop
abstract
Business documents are central to the operation of all organizations, and they come in all shapes and sizes: project reports, planning documents, technical specifications, financial statements, meeting minutes, legal agreements, contracts, resumes, purchase orders, invoices, and many more. The ability to read, understand and interpret these documents, referred to here as Document Intelligence (DI), is challenging due to not only many domains of knowledge involved, but also their complex formats and structures, internal and external cross references deployed, and even less-than-ideal quality of scans and OCR oftentimes performed on them. This workshop aims to explore and advance the current state of research and practice in answering these challenges.
Ani Nenkova, Douglas Burdick, Benjamin Han, Dave Lewis 0003, Sandeep Tata, Dan Tecuci
KDD3
2021 Budget Allocation as a Multi-Agent System of Contextual & Continuous Bandits
abstract
Budget allocation for online advertising suffers from multiple complications, including significant delay between the initial ad impression to the call to action as well as cold-start prediction problems for ad campaigns with limited or no historical performance data. To address these issues, we introduce the Contextual Budgeting System (CBS ), a budget allocation framework using a multi-agent system of contextual & continuous Multi-Armed Bandits. Our proposed solution decomposes the problem into a convex optimization problem whose objective is drawn using Thompson Sampling. In order to efficiently deal with context and cold-start, we propose a transfer learning mechanism using supervised learning methods that augment simple parametric models.
Benjamin Han, Carl Arndt
KDD1
2021 DI-2021: The Second Document Intelligence Workshop
abstract
Business documents are central to the operation of all organizations, and they come in all shapes and sizes: project reports, planning documents, technical specifications, financial statements, meeting minutes, legal agreements, contracts, resumes, purchase orders, invoices, and many more. The ability to read, understand and interpret these documents, referred to here as Document Intelligence (DI), is challenging due to not only many domains of knowledge involved, but also their complex formats and structures, internal and external cross references deployed, and even less-than-ideal quality of scans and OCR oftentimes performed on them. This workshop aims to explore and advance the current state of research and practice in answering these challenges.
Benjamin Han, Douglas Burdick, Dave Lewis 0003, Yijuan Lu, Hamid R. Motahari Nezhad, Sandeep Tata
KDD1
2006 Understanding Temporal Expressions in Emails
Benjamin Han, Donna Gates, Lori S. Levin
HLT-NAACL1
2006 From Language to Time: A Temporal Expression Anchorer
abstract
Understanding temporal expressions in natural language is a key step towards incorporating temporal information in many applications. In this paper we describe a system capable of anchoring such expressions in English: system TEA features a constraint-based calendar model and a compact representational language to capture the intensional meaning of temporal expressions. We also report favorable results from experiments conducted on several email datasets.
Benjamin Han, Donna Gates, Lori S. Levin
TIME1
2004 A framework for resolution of time in natural language
abstract
Automatic extraction and reasoning over temporal properties in natural language discourse has not had wide use in practical systems due to its demand for a rich and compositional, yet inference-friendly, representation of time. Motivated by our study of temporal expressions from the Penn Treebank corpora, we address the problem by proposing a two-level constraint-based framework for processing and reasoning over temporal information in natural language. Within this framework, temporal expressions are viewed as partial assignments to the variables of an underlying calendar constraint system, and multiple expressions together describe a temporal constraint-satisfaction problem (TCSP). To support this framework, we designed a typed formal language for encoding natural language expressions. The language can cope with phenomena such as under-specification and granularity change. The constraint problems can be solved using various constraint propagation and search methods, and the solutions can then be used to answer a wide range of time-related queries.
Benjamin Han, Alon Lavie
ACM Trans. Asian Lang. Inf. Process.1
1999 A Model-Based Diagnosis System for Identifying Faulty Components in Digital Circuits
Benjamin Han, Shie-Jue Lee, Hsin-Tai Yang
Appl. Intell.1
1999 A Genetic Algorithm Approach to Measurement Prescription in Fault Diagnosis
Benjamin Han, Shie-Jue Lee
Inf. Sci.1
1999 Comments on the Theory of Measurement in Diagnosis from First Principles
Benjamin Han, Shie-Jue Lee, Hsin-Tai Yang
Inf. Sci.1
1999 Deriving minimal conflict sets by CS-trees with mark set in diagnosis from first principles
abstract
To discriminate among all possible diagnoses using Hou's theory of measurement in diagnosis from first principles, one has to derive all minimal conflict sets from a known conflict set. However, the result derived from Hou's method depends on the order of node generation in CS-trees. We develop a derivation method with mark set to overcome this drawback of Hou's method. We also show that our method is more efficient in the sense that no redundant tests have to be done. An enhancement to our method with the aid of extra information is presented. Finally, a discussion on top-down and bottom-up derivations is given.
Benjamin Han, Shie-Jue Lee
IEEE Trans. Syst. Man Cybern. Part B1
1995 A Hybrid Diagnosis System for Digital Circuits
Benjamin Han, Hsin-Tai Yang, Shie-Jue Lee
IEA/AIE1