Emre Barut

dblp:141/7475 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0003-3064-6227ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 46% Question answering and dialogue systems · 18% Deep learning architectures and training · 15%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.612022
Building Goal-Oriented Dialogue Systems with Situated Visual Context · AAAI 2022
Machine learning › Deep learning architectures and training › feedforward neural network
convex neural network
0.512021
Uncertainty Quantification in CNN Through the Bootstrap of Convex Neural Networks · AAAI 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
Statistically Consistent Saliency Estimation · ICCV 2021
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
model-agnostic explanation
0.512021
Statistically Consistent Saliency Estimation · ICCV 2021
Computer vision › Segmentation and scene understanding
saliency detection
0.512021
Statistically Consistent Saliency Estimation · ICCV 2021
Machine learning › Trustworthy machine learning
uncertainty estimation
0.512021
Uncertainty Quantification in CNN Through the Bootstrap of Convex Neural Networks · AAAI 2021

Methods — techniques the papers use, named apart from their topics

multimodal conversational framework · 0.6joint action and argument prediction · 0.6transfer learning · 0.5perturbation scheme · 0.5linear programming · 0.5convexified neural networks · 0.5bootstrap · 0.5
YearPublicationVenuePosition
2024 Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
abstract
We introduce a novel framework, LM-Guided CoT, that leverages a lightweight (i.e., <1B) language model (LM) for guiding a black-box large (i.e., >10B) LM in reasoning tasks. Specifically, the lightweight LM first generates a rationale for each input instance. The Frozen large LM is then prompted to predict a task output based on the rationale generated by the lightweight LM. Our approach is resource-efficient in the sense that it only requires training the lightweight LM. We optimize the model through 1) knowledge distillation and 2) reinforcement learning from rationale-oriented and task-oriented reward signals. We assess our method with multi-hop extractive question answering (QA) benchmarks, HotpotQA, and 2WikiMultiHopQA. Experimental results show that our approach outperforms all baselines regarding answer prediction accuracy. We also find that reinforcement learning helps the model to produce higher-quality rationales with improved QA performance.
Fan Yang 0155, Emre Barut, Kai-Wei Chang 0001
LREC/COLING5
2022 Building Goal-Oriented Dialogue Systems with Situated Visual Context
abstract
Goal-oriented dialogue agents can comfortably utilize the conversational context and understand its users' goals. However, in visually driven user experiences, these conversational agents are also required to make sense of the screen context in order to provide a proper interactive experience. In this paper, we propose a novel multimodal conversational framework where the dialogue agent's next action and their arguments are derived jointly conditioned both on the conversational and the visual context. We demonstrate the proposed approach via a prototypical furniture shopping experience for a multimodal virtual assistant.
Sanchit Agarwal, Jan Jezabek, Arijit Biswas, Emre Barut, Bill Gao, Tagyoung Chung
AAAI4
2022 Build Connections Between Two Groups of Images Using Deep Learning Method
abstract
Finding relationships between two related groups of images has been a challenging question in many fields. E.g., it is informative in neuron science studies to build the connection between experiment animals' neuron activity images and their behavior images. Very few previous works have achieved this task since most generative models focus on reconstructing the output images similar to input images. We proposed a novel framework in this paper to accomplish this goal, which could map images from one group to images from another group. We apply the singular value decomposition (SVD) method to remove the original images' background noise. Next, we combine two deep learning approaches, variational autoencoder (VAE) and convolutional neural networks (CNN), to directly connect two groups of images. We test our framework on images from a neuron science experiment. Results show that the proposed framework could generate mice paw movement images given the mice neuron images, which are very close to the ground truth images. In terms of capturing the paw gestures in paw movement images, experiment results demonstrate that our framework outperforms the state-of-art paw location detection method.
Emre Barut, Juntao Su
COMPSAC2
2022 Semantic VL-BERT: Visual Grounding via Attribute Learning
abstract
In recent years, Smart Home Assistants have expanded into tens of thousands of devices and transformed from a voice only assistant to a much more versatile smart assistant, that uses a connected display to provide a multi-modal customer experience. In order to further improve on the multi-modality experience, comprehension systems need models that can work with multisensory inputs. We focus on the problem of visual grounding, which allows customers to interact with and manipulate items displayed on a screen via voice. We propose a novel learning approach that improves upon a lightweight single stream transformer architecture by adjusting it to better align the visual input features with the referring expressions. Our approach learns to cluster parts of the image along spatial and channel dimensions based on descriptive attributes in the query, and takes advantage of the information in separate clusters more efficiently, as demonstrated by a 1.32% absolute accuracy improvement on a public dataset over the baseline. Given that modern-day Smart Home Assistants have very stringent memory and latency requirements, we restrict our focus to a family of lightweight single stream transformer architectures - our focus is not to beat the ever improving state-of-the-art in visual grounding but to improve upon a lightweight transformer architecture which leads to a model that is easy to train and deploy while having improved semantic awareness.
Prashan Wanigasekara, Kechen Qin, Emre Barut, Weitong Ruan, Chengwei Su
IJCNN3
2021 Uncertainty Quantification in CNN Through the Bootstrap of Convex Neural Networks
abstract
Despite the popularity of Convolutional Neural Networks (CNN), the problem of uncertainty quantification (UQ) of CNN has been largely overlooked. Lack of efficient UQ tools severely limits the application of CNN in certain areas, such as medicine, where prediction uncertainty is critically important. Among the few existing UQ approaches that have been proposed for deep learning, none of them has theoretical consistency that can guarantee the uncertainty quality. To address this issue, we propose a novel bootstrap based framework for the estimation of prediction uncertainty. The inference procedure we use relies on convexified neural networks to establish the theoretical consistency of bootstrap. Our approach has a significantly less computational load than its competitors, as it relies on warm-starts at each bootstrap that avoids refitting the model from scratch. We further explore a novel transfer learning method so our framework can work on arbitrary neural networks. We experimentally demonstrate our approach has a much better performance compared to other baseline CNNs and state-of-the-art methods on various image datasets.
Emre Barut, Fang Jin
AAAI2
2021 Introducing Deep Reinforcement Learning to Nlu Ranking Tasks
abstract
Natural language understanding (NLU) models in production systems rely heavily on human annotated data, which involves an expensive, time-consuming, and error-prone process. Moreover, the model release process requires offline supervised learning and human in-the-loop. Together, these factors prolong the model update cycle and result in sub-standard model performance in situations where usage behavior is non-stationary. In this paper, we address these issues with a deep reinforcement learning approach that ranks suggestions from multiple experts in an online fashion. Our proposed method removes the reliance on annotated data, and can effectively adapt to recent changes in the data distribution. The efficiency of the new approach is demonstrated through simulation experiments using logged data from voice-based virtual assistants. Our results show that our algorithm, without any reliance on annotation, outperforms offline supervised learning methods.
Emre Barut, Chengwei Su
ICASSP2
2021 Statistically Consistent Saliency Estimation
abstract
The growing use of deep learning for a wide range of data problems has highlighted the need to understand and diagnose these models appropriately, making deep learning interpretation techniques an essential tool for data analysts. The numerous model interpretation methods proposed in recent years are generally based on heuristics, with little or no theoretical guarantees. Here, we present a statistical framework for saliency estimation for black-box computer vision models. Our proposed model-agnostic estimation procedure, which is statistically consistent and capable of passing sanity checks, has polynomial-time computational efficiency since it only requires solving a linear program. An upper bound is established on the number of model evaluations needed to recover regions of importance with high probability through our theoretical analysis. Furthermore, a new perturbation scheme is presented for the estimation of local gradients that is more efficient than commonly used random perturbation schemes. The validity and excellence of our new method are demonstrated experimentally via sensitivity analyses.
Shunyan Luo, Emre Barut, Fang Jin
ICCV2
2014 Optimal learning for sequential sampling with non-parametric beliefs
Emre Barut, Warren B. Powell
J. Glob. Optim.1