VLDB 2026 Research / reviewers in the wild / expert
Ajay Dholakia
dblp:27/5574
· DBLP profile ↗
22ranked-venue papers
6as first author
7since 2021 · last 2026
0009-0007-8973-6063ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Computer networks · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model ImprovementabstractData drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines fast pattern recognition (System 1) with deliberate incremental updates (System 2). Our approach implements structured memory systems enabling System 1 to leverage successful solutions previously discovered by System 2, achieving efficient pattern-based responses without costly re-computation. KC-Agent incorporates atomic change principles and rollback capabilities to ensure reliable, verifiable updates in production environments. We evaluate our method on five datasets including real-world NASA turbofan data with authentic temporal degradation and synthetic datasets with controlled drift scenarios. KC-Agent achieves state-of-the-art performance (76.8% accuracy) while maintaining optimal efficiency (13.2s execution time), outperforming established cognitive architectures: CodeAct (+2.4%), Tree of Thoughts (+3.6%), ReAct (+8.0%), and Reflexion (+8.9%). Consensus evaluation by a panel of state-of-the-art LLMs confirms superior strategic efficacy (8.33/10 Smartness score), significantly outperforming baseline agents. The knowledge consolidation mechanism delivers 91% speedup over the slow variant while maintaining higher accuracy. Our approach demonstrates both theoretical foundations and practical viability for cognitive-inspired automated ML improvement systems capable of handling complex real-world data drift scenarios. Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain |
COMPSAC | 3 |
| 2025 | Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Rodrigo M. Carrillo-Larco, Ajay Dholakia, David Ellison |
AAMAS | 5 |
| 2024 | TADIL: Task-Agnostic Domain-Incremental Learning Through Task-ID Inference Using Transformer Nearest-Centroid Embeddings
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison |
ICPR (29) | 4 |
| 2022 | Scanflow-K8s: Agent-based Framework for Autonomic Management and Supervision of ML Workflows in Kubernetes ClustersabstractMachine Learning (ML) projects are currently heavily based on workflows composed of some reproducible steps and executed as containerized pipelines to build or deploy ML models efficiently because of the flexibility, portability, and fast delivery they provide to the ML life-cycle. However, deployed models need to be watched and constantly managed, supervised, and debugged to guarantee their availability, validity, and robustness in unexpected situations. Therefore, containerized ML workflows would benefit from leveraging flexible and diverse autonomic capabilities. This work presents an architecture for autonomic ML workflows with abilities for multi-layered control, based on an agent-based approach that enables autonomic management and supervision of ML workflows at the application layer and the infrastructure layer (by collaborating with the orchestrator). We redesign the Scanflow ML framework to support such multi-agent approach by using triggers, primitives, and strategies. We also implement a practical platform, so-called Scanflow-K8s, that enables autonomic ML workflows on Kubernetes clusters based on the Scanflow agents. MNIST image classification and MLPerf ImageNet classification benchmarks are used as case studies to show the capabilities of Scanflow-K8s under different scenarios. The experimental results demonstrate the feasibility and effectiveness of our proposed agent approach and the Scanflow-K8s platform for the autonomic management of ML workflows in Kubernetes clusters at multiple layers. Peini Liu, Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak |
CCGRID | 4 |
| 2022 | Human-in-the-loop online multi-agent approach to increase trustworthiness in ML models through trust scores and data augmentationabstractIncreasing a ML model accuracy is not enough, we must also increase its trustworthiness. This is an important step for building resilient AI systems for safety-critical applications such as automotive, finance, and healthcare. For that purpose, we propose a multi-agent system that combines both machine and human agents. In this system, a checker agent calculates a trust score of each instance (which penalizes overconfidence in predictions) using an agreement-based method and ranks it; then an improver agent filters the anomalous instances based on a human rule-based procedure (which is considered safe), gets the human labels, applies geometric data augmentation, and retrains with the augmented data using transfer learning. We evaluate the system on corrupted versions of the MNIST and FashionMNIST datasets. We get an improvement in accuracy and trust score with just few additional labels compared to a baseline approach. Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak |
COMPSAC | 4 |
| 2022 | Scanflow: A multi-graph framework for Machine Learning workflow management, supervision, and debugginabstractMachine Learning (ML) is more than just training models, the whole workflow must be considered. Once deployed, a ML model needs to be watched and constantly supervised and debugged to guarantee its validity and robustness in unexpected situations. Debugging in ML aims to identify (and address) the model weaknesses in not trivial contexts. Several techniques have been proposed to identify different types of model weaknesses, such as bias in classification, model decay, adversarial attacks, etc., yet there is not a generic framework that allows them to work in a collaborative, modular, portable, iterative way and, more importantly, flexible enough to allow both human- and machine-driven techniques. In this paper, we propose a novel containerized directed graph framework to support and accelerate end-to-end ML workflow management, supervision, and debugging. The framework allows defining and deploying ML workflows in containers, tracking their metadata, checking their behavior in production, and improving the models by using both learned and human-provided knowledge. We demonstrate these capabilities by integrating in the framework two hybrid systems to detect data drift distribution which identify the samples that are far from the latent space of the original distribution, ask for human intervention, and whether retrain the model or wrap it with a filter to remove the noise of corrupted data at inference time. We test these systems on MNIST-C, CIFAR-10-C, and FashionMNIST-C datasets, obtaining promising accuracy results with the help of human involvement. Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Jeffrey Falkanger, Miro Hodak |
Expert Syst. Appl. | 4 |
| 2021 | Recent Efficiency Gains in Deep Learning: Performance, Power, and SustainabilityabstractDeep learning (DL) continues to develop at a rapid pace with improvements coming from both hardware and software sides. In this work we evaluate the strides made over the last 2 years, during which a new generation of GPU accelerators has been introduced and significant algorithmic progress has been made. We find a dramatic improvement in runtime and power usage for a standard AI training workload. Specifically, a 4x improvement in runtime and energy consumption is demonstrated. The improvements are about equally split between hardware and algorithms. Additionally, we examine further ways to improve AI training power consumption on data center servers and identity 3 system level tunings that make most difference. These yield up to 20% more energy savings without any changes to the user code. Implications for the field and ways to make DL more energy-efficient going forward are also discussed. (Abstract) Miro Hodak, Ajay Dholakia |
IEEE BigData | 2 |
| 2020 | Big data deployment in containerized infrastructures through the interconnection of network namespacesabstractSummary Big Data applications tackle the challenge of fast handling of large streams of data. Their performance is not only dependent on the data frameworks implementation and the underlying hardware but also on the deployment scheme and its potential for fast scaling. Consequently, several efforts have focused on the ease of deployment of Big Data applications, notably through the use of containerization. This technology was indeed raised to bring multitenancy and multiprocessing out of clusters, providing high deployment flexibility through lightweight container images. Recent studies have focused mostly on Docker containers. Notwithstanding, this article is actually interested in recent Singularity containers as they provide more security and support high‐performance computing (HPC) environments and, in this way, they can make Big Data applications benefit from the specialized hardware of HPC. Singularity 2.x, however, does not isolate network resources as required by most Big Data components. Singularity 3.x allows allocating each container with isolated network resources, but their interconnection requires a nontrivial amount of configuration effort. In this context, this article makes a functional contribution in the form of a deployment scheme based on the interconnection of network namespaces, through underlay and overlay networking approaches, to make Big Data applications easily deployable inside Singularity containers. We provide detailed account of our deployment scheme when using both interconnection approaches in the form of a “how‐to‐do‐it” report, and we evaluate it by comparing three Big Data applications based on Hadoop when performing on a bare‐metal infrastructure and on scenarios involving Singularity and Docker instances. Carla Sauvanaud, Ajay Dholakia, Jordi Guitart, Chulho Kim, Peter Mayes |
Softw. Pract. Exp. | 2 |
| 2019 | Towards Power Efficiency in Deep Learning on Data Center HardwareabstractDeep learning (DL) is a computationally intensive workload that is expected to grow rapidly in data centers in the near future. Its high energy demand necessitates finding ways to improve computational efficiency. In this work, we directly measure power used by the whole system as well as that used by GPU, CPU, and RAM during DL training to determine their contributions to the overall energy consumption. We find that while GPUs use most of the power - about 70 % - the consumption of other components is also significant and their optimizations can bring important power savings. Evaluating a multitude of options, we identify the parameters that bring in the most power savings. Overall, an energy savings of over 20% of can be obtained by adjusting system settings alone without changing the workload, at the cost of a minor increase in runtime. Alternatively, if runtime needs to stay constant, an 18% energy savings is identified. In distributed multi-server DL, we find that scale-out overhead has only a small energy cost, making distributed training more energy-efficient than expected. Implications for the field and ways to make DL more energy-efficient going forward are also discussed. Miro Hodak, Masha Gorkovenko, Ajay Dholakia |
IEEE BigData | 3 |
| 2018 | Performance Implications of Big Data in Scalable Deep Learning: On the Importance of Bandwidth and CachingabstractDeep learning techniques have revolutionized many areas including computer vision and speech recognition. While such networks require tremendous amounts of data, the requirement for and connection to Big Data storage systems is often undervalued and not well understood. In this paper, we explore the relationship between Big Data storage, networking, and Deep Learning workloads to understand key factors for designing Big Data/Deep Learning integrated solutions. We find that storage and networking bandwidths are the main parameters determining Deep Learning training performance. Local data caching can provide a performance boost and eliminate repeated network transfers, but it is mainly limited to smaller datasets that fit into memory. On the other hand, local disk caching is an intriguing option that is overlooked in current state-of-the-art systems. Finally, we distill our work into guidelines for designing Big Data/Deep Learning solutions. Miro Hodak, David Ellison, Peter Seidel, Ajay Dholakia |
IEEE BigData | 4 |
| 2017 | Designing a high performance cluster for large-scale SQL-on-hadoop analyticsabstractExecuting and optimizing SQL analytics on Data Lakes and Enterprise Data Warehouses (EDW) are areas of significant and growing interest. Achieving high performance for SQL analytics on large-scale data repositories remains a key challenge for data practitioners. The SQL-on-Hadoop Analytics solution described in this paper is very well suited for implementing the infrastructure to support these modern analytics initiatives while meeting requirements such as higher performance, lower cost, more efficient data center footprint, lower power consumption, appropriate storage needs and increased reliability. By using a TPC-DS derived workload applied to 100 TB of data, the work demonstrates for the first time the feasibility of designing such an extremely high-performance cluster. Furthermore, it enables investigation of large-scale SQL-on-Hadoop systems as the Spark SQL framework matures and enables similar investigations into machine learning and related Spark capabilities. Ajay Dholakia, Prasad Venkatachar, Kshitij A. Doshi, Ravikanth Durgavajhala, Stewart Tate, Berni Schiefer, Matthew Sheard, Ramnath Sai Sagar |
IEEE BigData | 1 |
| 2008 | A new intra-disk redundancy scheme for high-reliability RAID storage systems in the presence of unrecoverable errorsabstractToday's data storage systems are increasingly adopting low-cost disk drives that have higher capacity but lower reliability, leading to more frequent rebuilds and to a higher risk of unrecoverable media errors. We propose an efficient intradisk redundancy scheme to enhance the reliability of RAID systems. This scheme introduces an additional level of redundancy inside each disk, on top of the RAID redundancy across multiple disks. The RAID parity provides protection against disk failures, whereas the proposed scheme aims to protect against media-related unrecoverable errors. In particular, we consider an intradisk redundancy architecture that is based on an interleaved parity-check coding scheme, which incurs only negligible I/O performance degradation. A comparison between this coding scheme and schemes based on traditional Reed--Solomon codes and single-parity-check codes is conducted by analytical means. A new model is developed to capture the effect of correlated unrecoverable sector errors. The probability of an unrecoverable failure associated with these schemes is derived for the new correlated model, as well as for the simpler independent error model. We also derive closed-form expressions for the mean time to data loss of RAID-5 and RAID-6 systems in the presence of unrecoverable errors and disk failures. We then combine these results to characterize the reliability of RAID systems that incorporate the intradisk redundancy scheme. Our results show that in the practical case of correlated errors, the interleaved parity-check scheme provides the same reliability as the optimum, albeit more complex, Reed--Solomon coding scheme. Finally, the I/O and throughput performances are evaluated by means of analysis and event-driven simulation. Ajay Dholakia, Evangelos Eleftheriou, Xiao-Yu Hu, Ilias Iliadis, Jai Menon 0001, K. K. Rao |
ACM Trans. Storage | 1 |
| 2005 | Reduced-Complexity Decoding of LDPC CodesabstractVarious log-likelihood-ratio-based belief-propagation (LLR-BP) decoding algorithms and their reduced-complexity derivatives for low-density parity-check (LDPC) codes are presented. Numerically accurate representations of the check-node update computation used in LLR-BP decoding are described. Furthermore, approximate representations of the decoding computations are shown to achieve a reduction in complexity by simplifying the check-node update, or symbol-node update, or both. In particular, two main approaches for simplified check-node updates are presented that are based on the so-called min-sum approximation coupled with either a normalization term or an additive offset term. Density evolution is used to analyze the performance of these decoding algorithms, to determine the optimum values of the key parameters, and to evaluate finite quantization effects. Simulation results show that these reduced-complexity decoding algorithms for LDPC codes achieve a performance very close to that of the BP algorithm. The unified treatment of decoding techniques for LDPC codes presented here provides flexibility in selecting the appropriate scheme from performance, latency, computational-complexity, and memory-requirement perspectives. Jinghu Chen, Ajay Dholakia, Evangelos Eleftheriou, Marc P. C. Fossorier, Xiao-Yu Hu |
IEEE Trans. Commun. | 2 |
| 2004 | Rate-compatible low-density parity-check codes for digital subscriber linesabstractRate-compatible low-density parity-check (LDPC) codes obtained from the class of array LDPC codes are presented. The design methodology described herein retains practical advantages of array LDPC codes such as excellent performance and efficient encodability across all the codes in a rate-compatible family. Different codes in the rate-compatible family can be specified by a small number of parameters and constructed algebraically with a small amount of preprocessing. The rate-compatible codes can be decoded using a generic decoder architecture, leading to efficient implementations. These properties make the codes attractive for use in DSL systems that need to support a large number of code parameters to cope with channel variability. Ajay Dholakia, Sedat Ölçer |
ICC | 1 |
| 2004 | Rate-compatible array LDPC codesabstractRate-compatible low-density parity-check (LDPC) codes obtained from the class of array LDPC codes are presented. Our approach builds on the deterministic array LDPC code construction and allows efficient encodability across all codes in a rate-compatible family. Furthermore, the use of permutation matrices as building blocks of the parity-check matrix permits efficient decoder implementation. The suitability of rate-compatible array (RCA) LDPC codes for various applications is examined. Ajay Dholakia, Sedat Ölçer |
ISIT | 1 |
| 2003 | A Nanotechnology-based Approach to Data Storage
Evangelos Eleftheriou, Peter Bächtold, Giovanni Cherubini, Ajay Dholakia, Christoph Hagleitner, Teddy Loeliger, Angeliki Pantazi, Haralampos Pozidis, T. R. Albrecht, Gerd Karl Binnig, Michel Despont, Ute Drechsler, Urs Dürig, Bernd Gotsmann, Daniel Jubin, Walter Häberle, Mark A. Lantz, Hugo E. Rothuizen, Richard Stutz, Peter Vettiger, Dorothea Wiesmann |
VLDB | 4 |
| 2001 | Efficient implementations of the sum-product algorithm for decoding LDPC codesabstractEfficient implementations of the sum-product algorithm (SPA) are presented for decoding low-density parity-check (LDPC) codes using log-likelihood ratios (LLR) as messages between symbol and parity-check nodes. Various reduced-complexity derivatives of the LLR-SPA are proposed. Both serial and parallel implementations are investigated, leading to trellis and tree topologies, respectively. Furthermore, by exploiting the inherent robustness of LLRs, it is shown, via simulations, that coarse quantization tables are sufficient to implement complex core operations with negligible or no loss in performance. The unified treatment of decoding techniques for LDPC codes presented here provides flexibility in selecting the appropriate design point in high-speed applications from a performance, latency and computational complexity perspective. Xiao-Yu Hu, Evangelos Eleftheriou, Dieter-Michael Arnold, Ajay Dholakia |
GLOBECOM | 4 |
| 2001 | Application of high-rate tail-biting codes to generalized partial response channelsabstractThe performance of high-rate tail-biting convolutional codes serially concatenated with generalized partial response channels is studied. The effect of precoders on the overall performance is investigated. Extrinsic information transfer charts are used to guide the selection of appropriate tail-biting codes and precoders. Simulation results for a magnetic recording system modeled as a serial concatenation of tail-biting codes with a generalized partial response channel are presented. In particular, rate-8/9 and -16/17 short- and long-block-length tail-biting codes are studied. In the former case, hard-decision decoded interleaved Reed-Solomon (RS) codes are used as the outer-most code, whereas in the latter case the sector-size tail-biting codes replace the RS codes traditionally used in storage systems. The results indicate that high-rate tail-biting codes deliver significant performance gains when used in conjunction with a rate-1 precoder and iterative detection/decoding. The results also show that long tail-biting codes can outperform hard-decision decoding of RS codes by 2 dB at a sector error rate of approx. 10/sup -4/. Michael Tüchler, Christian Weiss, Evangelos Eleftheriou, Ajay Dholakia, Joachim Hagenauer |
GLOBECOM | 4 |
| 1998 | Performance evaluation of burst-error-correcting codes on a Gilbert-Elliott channelabstractThe performance of single burst-error-correcting (BEC) codes used over bursty channels is evaluated. The channel is represented by the Gilbert-Elliott (1960, 1963) model, which has been used by numerous authors to evaluate the performance of random-error-correcting (REC) codes over bursty channels. Recursive expressions are derived, which are used in evaluating the probability of a codeword error. These expressions and an approximate closed-form expression are applied to the performance of a single (23,12) BEC code. Gaurav Sharma 0001, Amer A. Hassan, Ajay Dholakia |
IEEE Trans. Commun. | 3 |
| 1998 | On Locally Invertible Rate-1/n Convolutional EncodersabstractA locally invertible convolutional encoder has a local inverse defined as a full rank w/spl times/w matrix that specifies a one-to-one mapping between equal-length blocks of information and encoded bits. In this correspondence, it is shown that a rate-1/n convolutional encoder is nondegenerate and noncatastrophic if and only if it is locally invertible. Local invertibility is used to obtain upper and lower bounds on the number of consecutive zero-weight branches in a convolutional codeword. Further, existence of a local inverse can be used as an alternate test for noncatastrophicity instead of the usual approach involving computation of the greatest common divisor of n polynomials. Donald L. Bitzer, Ajay Dholakia, Havish Koorapaty, Mladen A. Vouk |
IEEE Trans. Inf. Theory | 2 |
| 1995 | Table based decoding of rate one-half convolutional codesabstractTable based error correction and decoding of rate one-half convolutional codes is described. A new class of fast-decodeable locally invertible convolutional codes based on a one-to-one mapping between information and encoded blocks of equal lengths is defined. The syndrome is used as an address to access a correction table which stores pre-computed correction information. The correction table generation process is described and a specific table based correction algorithm is given. Performance of this scheme is analyzed and simulation results are presented.> Ajay Dholakia, Mladen A. Vouk, Donald L. Bitzer |
IEEE Trans. Commun. | 1 |
| 1994 | A variable-redundancy hybrid ARQ scheme using invertible convolutional codesabstractNonstationary channels (e.g., digital mobile communication channels) require adaptive error control schemes for reliable communication. On these channels, a particular transmission may encounter no errors. Hence, it is desirable to split an encoded sequence into subsequences that are sent in successive transmissions such that each subsequence contains all the information necessary to recover the original message in case of no errors. When errors are present, the subsequences are combined to perform error correction. We show that invertible convolutional codes have this property, and can be used to provide incremental redundancy in a variable-redundancy hybrid ARQ (VR-HARQ) scheme. Invertible convolutional codes provide an alternative to polynomial division otherwise required to extract the original message in convolutional VR-HARQ schemes.> Ajay Dholakia, Mladen A. Vouk, Donald L. Bitzer |
VTC | 1 |