Pankaj Dhoolia

dblp:54/6684 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 12 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
6 papers
Program analysis · 36% Requirements engineering and software design · 33% Debugging and program repair · 18%
Artificial intelligence
2 papers
Question answering and dialogue systems · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 57% Parallel and multicore computing · 43%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.512021
Doc2Bot: Document grounded Bot Framework · AAAI 2021
Program analysis
dynamic analysis
0.322013
Distributed program tracing · ESEC/SIGSOFT FSE 2013
Fault localization for data-centric programs · SIGSOFT FSE 2011
Debugging and program repair
fault localization
0.222011
Fault localization for data-centric programs · SIGSOFT FSE 2011
Automated support for repairing input-model faults · ASE 2010
Programming languages and type systems
development environment
0.212015
Smart Programming Playgrounds · ICSE (2) 2015
Cloud and datacenter computing › resource allocation
cloud resource allocation
0.212015
Smart Programming Playgrounds · ICSE (2) 2015
Program analysis › dynamic analysis
program tracing
0.212013
Distributed program tracing · ESEC/SIGSOFT FSE 2013
Information retrieval › document processing › document analysis
document understanding
0.112021
Doc2Bot: Document grounded Bot Framework · AAAI 2021
Requirements engineering and software design
model-driven engineering
0.112011
Using MATCON to generate CASE tools that guide deployment of pre-packaged applications · ICSE 2011
Debugging and program repair
automated program repair
0.112010
Automated support for repairing input-model faults · ASE 2010
Program analysis › static analysis
taint analysis
0.112010
Automated support for repairing input-model faults · ASE 2010
Program analysis › static analysis
dependency analysis
0.112015
Smart Programming Playgrounds · ICSE (2) 2015
Program analysis
static analysis
0.112015
Smart Programming Playgrounds · ICSE (2) 2015

Methods — techniques the papers use, named apart from their topics

document digestion · 1.5dialog generation · 1.5user emulation · 0.5agent emulation · 0.5resource binding injection · 0.4context resolution · 0.4program instrumentation · 0.3edge-based profiling · 0.3program slicing · 0.1metamodeling · 0.1execution trace analysis · 0.1dynamic tainting · 0.1
YearPublicationVenuePosition
2021 Bootstrapping Dialog Models from Human to Human Conversation Logs
abstract
State-of-the-art commercial dialog platforms provide powerful tools to build a conversational agent. These platforms provide complete control to the dialog designer to model user-agent interactions. However, a dialog designer needs to rely on domain experts to manually build the dialog model -- by creating dialog flow nodes and modeling user intents. This process is laborious, time consuming and expensive and does not allow the designer to exploit human to human conversation logs effectively. In this work, we present a research prototype that can ingest human-to-human conversation logs between an end-user and an agent, and suggest user-intents and agent-responses, given a conversation context. We utilize human to human conversation logs to build two emulators: user and agent. An agent emulator models an agent response given the conversation context so far, and a user emulator outputs possible user responses. Our system is able to recommend conversational intents as well as conversation flow using emulators based on real-world data, thus making the process of designing a bot more efficient. To the best our knowledge this is the first system that enables data-driven dialog model creation by emulating users and agents.
Pankaj Dhoolia, Danish Contractor, Sachindra Joshi
AAAI1
2021 Doc2Bot: Document grounded Bot Framework
abstract
Conversational agents, or chatbots, are widely used to provide customer care and other informational support. Currently, the development of chatbots using standard frameworks requires a lot of manual crafting by subject matter experts (SMEs). On the other hand, while learning-based approaches to dialog have made significant advancements, they require training with a large volume of dialog data, which chatbot developers typically do not have access to. To tackle these challenges, we introduce DOC2BOT, a system that supports the automated construction of chatbots by digesting various forms of documents such as business manuals, HowTos, and customer support pages that organizations own. In addition to that, DOC2BOT provides a user-friendly experience to SMEs, and to minimize their effort by supporting intuitive interactions and streamlining their workflow.
Kshitij Fadnis, Pankaj Dhoolia, Qingzi Vera Liao, Steven Ross, Nathaniel Mills, Sachindra Joshi, Luis A. Lastras
AAAI2
2015 Smart Programming Playgrounds
abstract
Modern IDEs contain sophisticated components for inferring missing types, correcting bad syntax and completing partial expressions in code, but they are limited to the context that is explicitly defined in a project's configuration. These tools are ill-suited for quick prototyping of incomplete code snippets, such as those found on the Web in Q&A forums or walk-through tutorials, since such code snippets often assume the availability of external dependencies and may even contain implicit references to an execution environment that provides data or compute services. We propose an architecture for smart programming playgrounds that can facilitate rapid prototyping of incomplete code snippets through a semi-automatic context resolution that involves identifying static dependencies, provisioning external resources on the cloud and injecting resource bindings to handles in the original code fragment. Such a system could be potentially useful in a range of different scenarios, from sharing code snippets on the Web to experimenting with new ideas during traditional software development.
Rohan Padhye, Pankaj Dhoolia, Senthil Mani, Vibha Sinha
ICSE (2)2
2015 The Synergy between Voting and Acceptance of Answers on StackOverflow - Or the Lack Thereof
abstract
StackOverflow's primary goal is to serve as a platform for users to solicit answers regarding programming questions, though its archives are often used by other users who face similar issues and thus it serves a secondary purpose of documenting common problems. The two driving mechanisms for filtering out low quality posts and highlighting the best answers are community votes and the mark of acceptance by the original question asker. But does the asker's choice always match the popular vote? If so, is the asker's choice influenced by the community vote or is the community vote biased towards the accepted answer? And if the asker and community disagree, then can we determine any particular characteristics of posts that influence the choice of the asker and community differently, such as its size, readability, presence of code snippets and external links as well as similarity to the original question? In this paper, we explore the answers to these questions by studying a data-set of all posts on StackOverflow from its launch in September 2008 to September 2014.
Neelamadhav Gantayat, Pankaj Dhoolia, Rohan Padhye, Senthil Mani, Vibha Sinha
MSR2
2015 Detecting and Mitigating Secret-Key Leaks in Source Code Repositories
abstract
Several news articles in the past year highlighted incidents in which malicious users stole API keys embedded in files hosted on public source code repositories such as GitHub and Bit Bucket in order to drive their own work-loads for free. While some service providers such as Amazon have started taking steps to actively discover such developer carelessness by scouting public repositories and suspending leaked API keys, there is little support for tackling the problem from the code sharing platforms themselves. In this paper, we discuss practical solutions to detecting, preventing and fixing API key leaks. We first outline a handful of methods for detecting API keys embedded within source code, and evaluate their effectiveness using a sample set of projects from GitHub. Second, we enumerate the mechanisms which could be used by developers to prevent or fix key leaks in code repositories manually. Finally, we outline a possible solution that combines these techniques to provide tool support for protecting against key leaks in version control systems.
Vibha Sinha, Diptikalyan Saha, Pankaj Dhoolia, Rohan Padhye, Senthil Mani
MSR3
2013 Distributed program tracing
abstract
Dynamic program analysis techniques depend on accurate program traces. Program instrumentation is commonly used to collect these traces, which causes overhead to the program execution. Various techniques have addressed this problem by minimizing the number of probes/witnesses used to collect traces. In this paper, we present a novel distributed trace collection framework wherein, a program is executed multiple times with the same input for different sets of witnesses. The partial traces such obtained are then merged to create the whole program trace. Such divide-and-conquer strategy enables parallel collection of partial traces, thereby reducing the total time of collection. The problem is particularly challenging as arbitrary distribution of witnesses cannot guarantee correct formation of traces. We provide and prove a necessary and sufficient condition for distributing the witnesses which ensures correct formation of trace. Moreover, we describe witness distribution strategies that are suitable for parallel collection. We use the framework to collect traces of field SAP-ABAP programs using breakpoints as witnesses as instrumentation cannot be performed due to practical constraints. To optimize such collection, we extend Ball-Larus' optimal edge-based profiling algorithm to an optimal node-based algorithm. We demonstrate the effectiveness of the framework for collecting traces of SAP-ABAP programs.
Diptikalyan Saha, Pankaj Dhoolia, Gaurab Paul
ESEC/SIGSOFT FSE2
2011 Using MATCON to generate CASE tools that guide deployment of pre-packaged applications
abstract
The complex process of adapting pre-packaged applications, such as Oracle or SAP, to an organization's needs is full of challenges. Although detailed, structured, and well-documented methods govern this process, the consulting team implementing the method must spend a huge amount of manual effort to make sure the guidelines of the method are followed as intended by the method author. MATCON breaks down the method content, documents, templates, and work products into reusable objects, and enables them to be cataloged and indexed so these objects can be easily found and reused on subsequent projects. By using models and meta-modeling the reusable methods, we automatically produce a CASE tool to apply these methods, thereby guiding consultants through this complex process. The resulting tool helps consultants create the method deliverables for the initial phases of large customization projects. Our MATCON output, referred to as Consultant Assistant, has shown significant savings in training costs, a 20 - 30% improvement in productivity, and positive results in large Oracle and SAP implementations.
Elad Fein, Natalia Razinkov, Shlomit Shachor, Pietro Mazzoleni, SweeFen Goh, Richard Goodwin, Manisha Bhandar, Shyh-Kwei Chen, Juhnyoung Lee, Vibha Sinha, Senthil Mani, Debdoot Mukherjee, Biplav Srivastava, Pankaj Dhoolia
ICSE14
2011 Fault localization for data-centric programs
abstract
In this paper we present an automated technique for localizing faults in data-centric programs. Data-centric programs primarily interact with databases to get collections of content, process each entry in the collection(s), and output another collection or write it back to the database. One or more entries in the output may be faulty. In our approach, we gather the execution trace of a faulty program. We use a novel, precise slicing algorithm to break the trace into multiple slices, such that each slice maps to an entry in the output collection. We then compute the semantic difference between the slices that correspond to correct entries and those that correspond to incorrect ones. The "diff" helps to identify potentially faulty statements.
Diptikalyan Saha, Mangala Gowri Nanda, Pankaj Dhoolia, V. Krishna Nandivada, Vibha Sinha, Satish Chandra 0001
SIGSOFT FSE3
2010 From Informal Process Diagrams to Formal Process Models
Debdoot Mukherjee, Pankaj Dhoolia, Saurabh Sinha 0001, Aubrey J. Rembert, Mangala Gowri Nanda
BPM2
2010 Debugging Model-Transformation Failures Using Dynamic Tainting
Pankaj Dhoolia, Senthil Mani, Vibha Sinha, Saurabh Sinha 0001
ECOOP1
2010 Automated support for repairing input-model faults
abstract
Model transforms are a class of applications that convert a model to another model or text. The inputs to such transforms are often large and complex; therefore, faults in the models that cause a transformation to generate incorrect output can be difficult to identify and fix. In previous work, we presented an approach that uses dynamic tainting to help locate input-model faults. In this paper, we present techniques to assist with repairing input-model faults. Our approach collects runtime information for the failing transformation, and computes repair actions that are targeted toward fixing the immediate cause of the failure. In many cases, these repair actions result in the generation of the correct output. In other cases, the initial fix can be incomplete, with the input model requiring further repairs. To address this, we present a pattern-analysis technique that identifies correct output fragments that are similar to the incorrect fragment and, based on the taint information associated with such fragments, computes additional repair actions. We present the results of empirical studies, conducted using real model transforms, which illustrate the applicability and effectiveness of our approach for repairing different types of faults.
Senthil Mani, Vibha Sinha, Pankaj Dhoolia, Saurabh Sinha 0001
ASE3
2009 Efficient Testing of Service-Oriented Applications Using Semantic Service Stubs
abstract
Service-oriented applications can be expensive to test because services are hosted remotely, are potentially shared among many users, and may have costs associated with their invocation. In this paper, we present an approach for reducing the costs of testing such applications. The key observation underlying our approach is that certain aspects of an application can be tested using locally deployed semantic service stubs, instead of actual remote services.A semantic service stub incorporates some of the service functionality, such as verifying preconditions and generating output messages based on post conditions. We illustrate how semantic stubs can enable the client test suite to be partitioned into subsets, some of which need not be executed using remote services. We also present a case study that demonstrates the feasibility of the approach, and potential cost savings for testing. The main benefits of our approach are that it can (1) reduce the number of test cases that need to be run to invoke remote services, (2) ensure that certain aspects of application functionality are well-tested before service integration occurs.
Senthil Mani, Vibha Sinha, Saurabh Sinha 0001, Pankaj Dhoolia, Debdoot Mukherjee, Soham Chakraborty 0001
ICWS4
2008 Siena: From PowerPoint to Web App in 5 Minutes
David Cohn, Pankaj Dhoolia, Terry Heath, Florian Pinel, John Vergo
ICSOC2
2008 Toward the Development of Contextually Aware Business Applications via Model-Driven Transformations
abstract
While the traditional model driven development techniques are useful for building solutions in a reusable manner, they do not say much about how the existing assets in a client environment can be leveraged effectively and efficiently. In this work, we enhance model driven transformation techniques to generate implementation artifacts on a given platform from platform independent models while leveraging the existing assets in a client environment. We apply semantic Web service matching technology to achieve automatic binding of generated artifacts with available client assets. By generating implementation artifacts that are bound where appropriate with clientspsila existing functionality, our approach helps cut down the development time during project implementations and thereby resulting in reduced project durations and costs. We demonstrate the feasibility of the two platforms: IBM WebSphere and SAP NetWeaver. Lessons learned are presented.
Rama Akkiraju, Tilak Mitra, Pankaj Dhoolia, Wei Zhao 0003, Shiwa Fu, Manisha Bhandar, Nilay Ghosh, Dipankar Saha
ICWS3