VLDB 2026 Research / reviewers in the wild / expert
Dipanjan Das 0001
dblp:90/3182-1
· DBLP profile ↗
42ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0003-2325-0646ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 7 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dolomites: Domain-Specific Long-Form Methodical TasksabstractAbstract Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a teacher writing a lesson plan for students, these tasks are pervasive, requiring to methodically generate structured long-form output for a given input. We develop a typology of methodical tasks structured in the form of a task objective, procedure, input, and output, and introduce DoLoMiTes, a novel benchmark with specifications for 519 such tasks elicited from hundreds of experts from across 25 fields. Our benchmark further contains specific instantiations of methodical tasks with concrete input and output examples (1,857 in total) which we obtain by collecting expert revisions of up to 10 model-generated examples of each task. We use these examples to evaluate contemporary language models, highlighting that automating methodical tasks is a challenging long-form generation problem, as it requires performing complex inferences, while drawing upon the given context as well as domain knowledge. Our dataset is available at https://dolomites-benchmark.github.io/. Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das 0001, Mirella Lapata, Christopher Alberti |
Trans. Assoc. Comput. Linguistics | 8 |
| 2023 | Query Refinement Prompts for Closed-Book Long-Form QAabstractReinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das, Shashi Narayan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Reinald Kim Amplayo, Kellie Webster, Michael Collins 0001, Dipanjan Das 0001, Shashi Narayan |
ACL (1) | 4 |
| 2023 | SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization EvaluationabstractElizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur Parikh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das 0001, Ankur P. Parikh |
EMNLP | 9 |
| 2023 | Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 0002, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das 0001, Jason Wei |
ICLR | 11 |
| 2023 | Measuring Attribution in Natural Language Generation ModelsabstractAbstract Large neural models have brought a new challenge to natural language generation (NLG): It has become imperative to ensure the safety and reliability of the output of models that generate freely. To this end, we present an evaluation framework, Attributable to Identified Sources (AIS), stipulating that NLG output pertaining to the external world is to be verified against an independent, provided source. We define AIS and a two-stage annotation pipeline for allowing annotators to evaluate model output according to annotation guidelines. We successfully validate this approach on generation datasets spanning three tasks (two conversational QA datasets, a summarization dataset, and a table-to-text dataset). We provide full annotation guidelines in the appendices and publicly release the annotated data at https://github.com/google-research-datasets/AIS. Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins 0001, Dipanjan Das 0001, Slav Petrov, Gaurav Tomar, Iulia Turc, David Reitter |
Comput. Linguistics | 6 |
| 2023 | QAmeleon: Multilingual QA with Only 5 ExamplesabstractAbstract The availability of large, high-quality datasets has been a major driver of recent progress in question answering (QA). Such annotated datasets, however, are difficult and costly to collect, and rarely exist in languages other than English, rendering QA technology inaccessible to underrepresented languages. An alternative to building large monolingual training datasets is to leverage pre-trained language models (PLMs) under a few-shot learning setting. Our approach, QAmeleon, uses a PLM to automatically generate multilingual data upon which QA models are fine-tuned, thus avoiding costly annotation. Prompt tuning the PLM with only five examples per language delivers accuracy superior to translation-based baselines; it bridges nearly 60% of the gap between an English-only baseline and a fully-supervised upper bound fine-tuned on almost 50,000 hand-labeled examples; and consistently leads to improvements compared to directly fine-tuning a QA model on labeled examples in low resource settings. Experiments on the TyDiqa-GoldP and MLQA benchmarks show that few-shot prompt tuning for data synthesis scales across languages and is a viable alternative to large-scale annotation.1 Priyanka Agrawal, Christopher Alberti, Fantine Huot, Joshua Maynez, Ji Ma 0004, Sebastian Ruder, Kuzman Ganchev, Dipanjan Das 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 8 |
| 2023 | Conditional Generation with a Question-Answering BlueprintabstractAbstract The ability to convey relevant and faithful information is critical for many tasks in conditional generation and yet remains elusive for neural seq-to-seq models whose outputs often reveal hallucinations and fail to correctly cover important details. In this work, we advocate planning as a useful intermediate representation for rendering conditional generation less opaque and more grounded. We propose a new conceptualization of text plans as a sequence of question-answer (QA) pairs and enhance existing datasets (e.g., for summarization) with a QA blueprint operating as a proxy for content selection (i.e., what to say) and planning (i.e., in what order). We obtain blueprints automatically by exploiting state-of-the-art question generation technology and convert input-output pairs into input-blueprint-output tuples. We develop Transformer-based models, each varying in how they incorporate the blueprint in the generated output (e.g., as a global plan or iteratively). Evaluation across metrics and datasets demonstrates that blueprint models are more factual than alternatives which do not resort to planning and allow tighter control of the generation output. Shashi Narayan, Joshua Maynez, Reinald Kim Amplayo, Kuzman Ganchev, Annie Louis, Fantine Huot, Anders Sandholm 0001, Dipanjan Das 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 8 |
| 2022 | A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional GenerationabstractShashi Narayan, Gonçalo Simões, Yao Zhao, Joshua Maynez, Dipanjan Das, Michael Collins, Mirella Lapata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shashi Narayan, Gonçalo Simões, Joshua Maynez, Dipanjan Das 0001, Michael Collins 0001, Mirella Lapata |
ACL (1) | 5 |
| 2022 | The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick |
ICLR | 11 |
| 2021 | Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable FeaturesabstractHannah Rashkin, David Reitter, Gaurav Singh Tomar, Dipanjan Das. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hannah Rashkin, David Reitter, Gaurav Tomar, Dipanjan Das 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Decontextualization: Making Sentences Stand-AloneabstractAbstract Models for question answering, dialogue agents, and summarization often interpret the meaning of a sentence in a rich context and use that meaning in a new context. Taking excerpts of text can be problematic, as key pieces may not be explicit in a local window. We isolate and define the problem of sentence decontextualization: taking a sentence together with its context and rewriting it to be interpretable out of context, while preserving its meaning. We describe an annotation procedure, collect data on the Wikipedia corpus, and use the data to train models to automatically decontextualize sentences. We present preliminary studies that show the value of sentence decontextualization in a user-facing task, and as preprocessing for systems that perform document understanding. We argue that decontextualization is an important subtask in many downstream applications, and that the definitions and resources provided can benefit tasks that operate on sentences that occur in a richer context. Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das 0001, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2020 | Syntactic Data Augmentation Increases Robustness to Inference HeuristicsabstractPretrained neural models such as BERT, when fine-tuned to perform natural language inference (NLI), often show high accuracy on standard datasets, but display a surprising lack of sensitivity to word order on controlled challenge sets.We hypothesize that this issue is not primarily caused by the pretrained model's limitations, but rather by the paucity of crowdsourced NLI examples that might convey the importance of syntactic structure at the finetuning stage.We explore several methods to augment standard training sets with syntactically informative examples, generated by applying syntactic transformations to sentences from the MNLI corpus.The best-performing augmentation method, subject/object inversion, improved BERT's accuracy on controlled examples that diagnose sensitivity to word order from 0.28 to 0.73, without affecting performance on the MNLI test set.This improvement generalized beyond the particular construction used for data augmentation, suggesting that augmentation causes BERT to recruit abstract syntactic representations. Junghyun Min, Tom McCoy 0001, Dipanjan Das 0001, Emily Pitler, Tal Linzen |
ACL | 3 |
| 2020 | BLEURT: Learning Robust Metrics for Text GenerationabstractText generation has made significant advances in the last few years.Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate poorly with human judgments.We propose BLEURT, a learned evaluation metric based on BERT that can model human judgments with a few thousand possibly biased training examples.A key aspect of our approach is a novel pre-training scheme that uses millions of synthetic examples to help the model generalize.BLEURT provides state-ofthe-art results on the last three years of the WMT Metrics shared task and the WebNLG Competition dataset.In contrast to a vanilla BERT-based approach, it yields superior results even when the training data is scarce and out-of-distribution. Thibault Sellam, Dipanjan Das 0001, Ankur P. Parikh |
ACL | 2 |
| 2020 | ToTTo: A Controlled Table-To-Text Generation DatasetabstractAnkur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Ankur P. Parikh, Xuezhi Wang 0002, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das 0001 |
EMNLP (1) | 7 |
| 2019 | Handling Divergent Reference Texts when Evaluating Table-to-Text GenerationabstractAutomatically constructed datasets for generating text from semi-structured data (tables), such as WikiBio (Lebret et al., 2016), often contain reference texts that diverge from the information in the corresponding semistructured data.We show that metrics which rely solely on the reference texts, such as BLEU and ROUGE, show poor correlation with human judgments when those references diverge.We propose a new metric, PAR-ENT, which aligns n-grams from the reference and generated texts to the semi-structured data before computing their precision and recall.Through a large scale human evaluation study of table-to-text models for WikiBio, we show that PARENT correlates with human judgments better than existing text generation metrics.We also adapt and evaluate the information extraction based evaluation proposed in Wiseman et al. (2017), and show that PAR-ENT has comparable correlation to it, while being easier to use.We show that PARENT is also applicable when the reference texts are elicited from humans using the data from the WebNLG challenge.1 * Work done during an internship at Google. Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das 0001, William W. Cohen |
ACL (1) | 5 |
| 2019 | BERT Rediscovers the Classical NLP PipelineabstractPre-trained text encoders have rapidly advanced the state of the art on many NLP tasks.We focus on one such model, BERT, and aim to quantify where linguistic information is captured within the network.We find that the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way, and that the regions responsible for each step appear in the expected sequence: POS tagging, parsing, NER, semantic roles, then coreference.Qualitative analysis reveals that the model can and often does adjust this pipeline dynamically, revising lowerlevel decisions on the basis of disambiguating information from higher-level representations. Ian Tenney, Dipanjan Das 0001, Ellie Pavlick |
ACL (1) | 2 |
| 2019 | What do you learn from context? Probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia 0002, Berlin Chen, Adam Poliak, Tom McCoy 0001, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das 0001, Ellie Pavlick |
ICLR (Poster) | 10 |
| 2018 | Learning To Split and Rephrase From Wikipedia Edit HistoryabstractSplit and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning.We extract a rich new dataset for this task by mining Wikipedia's edit history: WikiSplit contains one million naturally occurring sentence rewrites, providing sixty times more distinct split examples and a ninety times larger vocabulary than the WebSplit corpus introduced by Narayan et al. (2017) as a benchmark for this task.Incorporating WikiSplit as training data produces a model with qualitatively better predictions that score 32 BLEU points above the prior best result on the WebSplit benchmark. Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge, Dipanjan Das 0001 |
EMNLP | 5 |
| 2018 | Identifying Well-formed Natural Language QuestionsabstractUnderstanding search queries is a hard problem as it involves dealing with "word salad" text ubiquitously issued by users.However, if a query resembles a well-formed question, a natural language processing pipeline is able to perform more accurate interpretation, thus reducing downstream compounding errors.Hence, identifying whether or not a query is well formed can enhance query understanding.Here, we introduce a new task of identifying a well-formed natural language question.We construct and release a dataset of 25,100 publicly available questions classified into well-formed and non-wellformed categories and report an accuracy of 70.7% on the test set.We also show that our classifier can be used to improve the performance of neural sequence-to-sequence models for generating questions for reading comprehension. Manaal Faruqui, Dipanjan Das 0001 |
EMNLP | 2 |
| 2018 | WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and DiscourseabstractWe release a corpus of 43 million atomic edits across 8 languages.These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single contiguous phrase from, an existing sentence.We use the collected data to show that the language generated during editing differs from the language that we observe in standard corpora, and that models trained on edits encode different aspects of semantics and discourse than models trained on raw, unstructured text.We release the full corpus as a resource to aid ongoing research in semantics, discourse, and representation learning. Manaal Faruqui, Ellie Pavlick, Ian Tenney, Dipanjan Das 0001 |
EMNLP | 4 |
| 2016 | A Decomposable Attention Model for Natural Language InferenceabstractWe propose a simple neural architecture for natural language inference.Our approach uses attention to decompose the problem into subproblems that can be solved separately, thus making it trivially parallelizable.On the Stanford Natural Language Inference (SNLI) dataset, we obtain state-of-the-art results with almost an order of magnitude fewer parameters than previous work and without relying on any word-order information.Adding intra-sentence attention that takes a minimum amount of order into account yields further improvements. Ankur P. Parikh, Oscar Täckström, Dipanjan Das 0001, Jakob Uszkoreit |
EMNLP | 3 |
| 2016 | Transforming Dependency Structures to Logical Forms for Semantic ParsingabstractThe strongly typed syntax of grammar formalisms such as CCG, TAG, LFG and HPSG offers a synchronous framework for deriving syntactic structures and semantic logical forms. In contrast—partly due to the lack of a strong type system—dependency structures are easy to annotate and have become a widely used form of syntactic analysis for many languages. However, the lack of a type system makes a formal mechanism for deriving logical forms from dependency structures challenging. We address this by introducing a robust system based on the lambda calculus for deriving neo-Davidsonian logical forms from dependency trees. These logical forms are then used for semantic parsing of natural language to Freebase. Experiments on the Free917 and Web-Questions datasets show that our representation is superior to the original dependency trees and that it outperforms a CCG-based representation on this task. Compared to prior work, we obtain the strongest result to date on Free917 and competitive results on WebQuestions. Siva Reddy, Oscar Täckström, Michael Collins 0001, Tom Kwiatkowski, Dipanjan Das 0001, Mark Steedman, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 5 |
| 2015 | Semantic Role Labeling with Neural Network FactorsabstractWe present a new method for semantic role labeling in which arguments and semantic roles are jointly embedded in a shared vector space for a given predicate.These embeddings belong to a neural network, whose output represents the potential functions of a graphical model designed for the SRL task.We consider both local and structured learning methods and obtain strong results on standard PropBank and FrameNet corpora with a straightforward product-of-experts model.We further show how the model can learn jointly from PropBank and FrameNet annotations to obtain additional improvements on the smaller FrameNet dataset. Nicholas FitzGerald, Oscar Täckström, Kuzman Ganchev, Dipanjan Das 0001 |
EMNLP | 4 |
| 2015 | Efficient Inference and Structured Learning for Semantic Role LabelingabstractWe present a dynamic programming algorithm for efficient constrained inference in semantic role labeling. The algorithm tractably captures a majority of the structural constraints examined by prior work in this area, which has resorted to either approximate methods or off-the-shelf integer linear programming solvers. In addition, it allows training a globally-normalized log-linear model with respect to constrained conditional likelihood. We show that the dynamic program is several times faster than an off-the-shelf integer linear programming solver, while reaching the same solution. Furthermore, we show that our structured model results in significant improvements over its local counterpart, achieving state-of-the-art results on both PropBank- and FrameNet-annotated corpora. Oscar Täckström, Kuzman Ganchev, Dipanjan Das 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2014 | Semantic Frame Identification with Distributed Word RepresentationsabstractA computer-implemented technique can include receiving, at a server, labeled training data including a plurality of groups of words, each group of words having a predicate word, each word having generic word embeddings. The technique can include extracting, at the server, the plurality of groups of words in a syntactic context of their predicate words. The technique can include concatenating, at the server, the generic word embeddings to create a high dimensional vector space representing features for each word. The technique can include obtaining, at the server, a model having a learned mapping from the high dimensional vector space to a low dimensional vector space and learned embeddings for each possible semantic frame in the low dimensional vector space. The technique can also include outputting, by the server, the model for storage, the model being configured to identify a specific semantic frame for an input. Karl Moritz Hermann, Dipanjan Das 0001, Jason Weston, Kuzman Ganchev |
ACL (1) | 2 |
| 2014 | Learning Compact Lexicons for CCG Semantic ParsingabstractWe present methods to control the lexicon size when learning a Combinatory Categorial Grammar semantic parser.Existing methods incrementally expand the lexicon by greedily adding entries, considering a single training datapoint at a time.We propose using corpus-level statistics for lexicon learning decisions.We introduce voting to globally consider adding entries to the lexicon, and pruning to remove entries no longer required to explain the training data.Our methods result in state-of-the-art performance on the task of executing sequences of natural language instructions, achieving up to 25% error reduction, with lexicons that are up to 70% smaller and are qualitatively less noisy.* This research was carried out at Google. Yoav Artzi, Dipanjan Das 0001, Slav Petrov |
EMNLP | 2 |
| 2014 | Frame-Semantic ParsingabstractFrame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised learning to improve frame disambiguation for targets unseen at training time. The second stage finds the target's locally expressed semantic arguments. At inference time, a fast exact dual decomposition algorithm collectively predicts all the arguments of a frame at once in order to respect declaratively stated linguistic constraints, resulting in qualitatively better structures than naïve local predictors. Both components are feature-based and discriminatively trained on a small set of annotated frame-semantic parses. On the SemEval 2007 benchmark data set, the approach, along with a heuristic identifier of frame-evoking targets, outperforms the prior state of the art by significant margins. Additionally, we present experiments on the much larger FrameNet 1.5 data set. We have released our frame-semantic parser as open-source software. Dipanjan Das 0001, Desai Chen, André F. T. Martins, Nathan Schneider 0001, Noah A. Smith |
Comput. Linguistics | 1 |
| 2013 | Cross-Lingual Discriminative Learning of Sequence Models with Posterior RegularizationabstractWe present a framework for cross-lingual transfer of sequence information from a resource-rich source language to a resourceimpoverished target language that incorporates soft constraints via posterior regularization.To this end, we use automatically word aligned bitext between the source and target language pair, and learn a discriminative conditional random field model on the target side.Our posterior regularization constraints are derived from simple intuitions about the task at hand and from cross-lingual alignment information.We show improvements over strong baselines for two tasks: part-of-speech tagging and namedentity segmentation. Kuzman Ganchev, Dipanjan Das 0001 |
EMNLP | 2 |
| 2013 | Token and Type Constraints for Cross-Lingual Part-of-Speech TaggingabstractWe consider the construction of part-of-speech taggers for resource-poor languages. Recently, manually constructed tag dictionaries from Wiktionary and dictionaries projected via bitext have been used as type constraints to overcome the scarcity of annotated data in this setting. In this paper, we show that additional token constraints can be projected from a resource-rich source language to a resource-poor target language via word-aligned bitext. We present several models to this end; in particular a partially observed conditional random field model, where coupled token and type constraints provide a partial signal for training. Averaged across eight previously studied Indo-European languages, our model achieves a 25% relative error reduction over the prior state of the art. We further present successful results on seven additional languages from different families, empirically demonstrating the applicability of coupled token and type constraints across a diverse set of languages. Oscar Täckström, Dipanjan Das 0001, Slav Petrov, Ryan T. McDonald, Joakim Nivre |
Trans. Assoc. Comput. Linguistics | 2 |
| 2012 | A Universal Part-of-Speech Tagset
Slav Petrov, Dipanjan Das 0001, Ryan T. McDonald |
LREC | 2 |
| 2012 | Graph-Based Lexicon Expansion with Sparsity-Inducing Penalties
Dipanjan Das 0001, Noah A. Smith |
HLT-NAACL | 1 |
| 2011 | Unsupervised Part-of-Speech Tagging with Bilingual Graph-Based Projections
Dipanjan Das 0001, Slav Petrov |
ACL | 1 |
| 2011 | Semi-Supervised Frame-Semantic Parsing for Unknown Predicates
Dipanjan Das 0001, Noah A. Smith |
ACL | 1 |
| 2011 | Unsupervised Structure Prediction with Non-Parallel Multilingual Guidance
Shay B. Cohen, Dipanjan Das 0001, Noah A. Smith |
EMNLP | 2 |
| 2010 | Distributed Asynchronous Online Learning for Natural Language Processing
Kevin Gimpel, Dipanjan Das 0001, Noah A. Smith |
CoNLL | 2 |
| 2010 | Probabilistic Frame-Semantic Parsing
Dipanjan Das 0001, Nathan Schneider 0001, Desai Chen, Noah A. Smith |
HLT-NAACL | 1 |
| 2010 | Movie Reviews and Revenues: An Experiment in Text Regression
Mahesh Joshi, Dipanjan Das 0001, Kevin Gimpel, Noah A. Smith |
HLT-NAACL | 2 |
| 2009 | Paraphrase Identification as Probabilistic Quasi-Synchronous Recognition
Dipanjan Das 0001, Noah A. Smith |
ACL/IJCNLP | 1 |
| 2008 | Stacking Dependency Parsers
André F. T. Martins, Dipanjan Das 0001, Noah A. Smith, Eric P. Xing |
EMNLP | 2 |
| 2008 | Automatic Extraction of Briefing Templates
Dipanjan Das 0001, Alexander I. Rudnicky |
IJCNLP | 1 |
| 2007 | Combating information overload in non-visual web access using contextabstractWeb sites are designed for graphical mode of interaction. Sighted users can visually segment Web pages and quickly identify relevant information. In contrast, visually-disabled individuals have to use screen readers to browse the Web. Screen readers process pages sequentially and read through everything, making Web browsing time-consuming and strenuous. The use of shortcut keys and searching offers some improvements, but the problem still remains. In this paper, we address this problem using the notion of context. When a user follows a link, we capture the context of the link, and use it to identify relevant information on the next page. The content of this page is rearranged, so that the relevant information is read out first. We conducted a series experiments to compare the performance of our prototype system with the state-of-the-art JAWS screen reader. Our results show that the use of context can potentially save browsing time as well as improve browsing experience of visually disabled individuals. Jalal Mahmud, Yevgen Borodin, Dipanjan Das 0001, I. V. Ramakrishnan |
IUI | 3 |
| 2006 | Improving non-visual web access using contextabstractTo browse the Web, blind people have to use screen readers, which process pages sequentially, making browsing timeconsuming. We present a prototype system, CSurf, which provides all features of a regular screen reader, but when a user follows a link, CSurf captures the context of the link and uses it to identify relevant information on the next page. CSurf rearranges the content of the next page, so, that the relevant information is read out first. A series experiments have been conducted to evaluate the performance of CSurf. Jalal Mahmud, Yevgen Borodin, Dipanjan Das 0001, I. V. Ramakrishnan |
ASSETS | 3 |