Vijay Sundaresan

dblp:51/2959 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0006-9342-4356ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 3 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VFlatten: Selective Value-Object Flattening using Hybrid Static and Dynamic Analysis
abstract
Object flattening is a non-trivial optimization that inlines the fields of an object inside its containers. Owing to its direct applicability for immutable objects, Java would soon allow programmers to mark compatible classes as "value types", and Java Virtual Machines (JVMs) to transparently flatten their instances (value objects). Expectations include reduced memory footprint, faster field access, and overall improvement in performance. This paper describes the surprises and challenges we faced while experimenting with value types and object flattening on a real-world JVM, and presents the design of an efficient strategy that selectively flattens profitable value objects, using a novel combination of static and dynamic analyses.Our value-object flattening strategy is based on insights that span the source program, the just-in-time (JIT) compiler employed by the JVM, as well as the underlying hardware. The first insight identifies source-level patterns that favour and oppose value-object flattening. The second insight finds an interesting dependence of object flattening on object scalarization, and estimates the capability of the JIT in avoiding overheads using escape analysis. Finally, the third insight correlates container objects with cache-line size, based on the load semantics of object fields. In order to develop an efficient strategy to flatten potentially profitable objects, we capture these insights in a tool called VFLATTEN that uses a novel combination of static and dynamic analyses and flattens value objects selectively in a production Java runtime.
Arjun H. Kumar, Bhavya Hirani, Tobi Ajila, Vijay Sundaresan, Daryl Maier, Manas Thakur
CGO5
2026 In-Production Characterization of an Open Source Serverless Platform and New Scaling Strategies
abstract
Serverless computing has become more popular and evolved to support more complex tasks than the original Function as a Service (FaaS) model. The design of serverless systems has advanced to accommodate application demands and offer flexibility. Careful characterization of modern serverless systems and understanding of current gaps are warranted. Publicly available datasets on workloads in select production serverless systems do not fully represent all offerings or capture traces at the required time resolution to identify changes in application-level request-response patterns.
Nima Nasiri, Nalin Munshi, Simon Moser, Marius Pirvu, Vijay Sundaresan, Daryl Maier, Thatta Premnath, Norman Böwing, Sathish Gopalakrishnan, Mohammad Shahrad
EuroSys5
2025 CoSSJIT: Combining Static Analysis and Speculation in JIT Compilers
abstract
Just-in-time (JIT) compilers typically sacrifice the precision of program analysis for efficiency, but are capable of performing sophisticated speculative optimizations based on run-time profiles to generate code that is specialized to a given execution. On the contrary, ahead-of-time static compilers can often afford precise flow-sensitive interprocedural analysis, but produce conservative results in scenarios where higher precision could be derived from run-time specialization. In this paper, we propose the first-of-its-kind approach to enrich static analysis with the possibility of speculative optimization during JIT compilation, as well as its usage to perform aggressive stack allocation on a production Java Virtual Machine (JVM). Our approach of combining static analysis with JIT speculation – named CoSSJIT – involves three key contributions. First, we identify the scenarios where a static analysis would make conservative assumptions but a JIT could deliver precision based on run-time speculation. Second, we present the notion of ‘‘speculative conditions’’ and plug them into a static interprocedural dataflow analyzer (whose aim is to identify heap objects that can be allocated on stack), to generate partial results that can be specialized at run-time. Finally, we extend a production JIT compiler to read and enrich static-analysis results with the resolved values of speculative conditions, leading to a practical approach that efficiently combines the best of both worlds. Cherries on the cake: Using CoSSJIT , we obtain 5.7× improvement in stack allocation (translating to performance), while building on a system that ensures functional correctness during JIT compilation.
Aditya Anand 0002, Vijay Sundaresan, Daryl Maier, Manas Thakur
Proc. ACM Program. Lang.2
2024 Identification of Java lock contention anti-patterns based on run-time performance data
abstract
Locks play a crucial role in multi-threaded applications, offering an effective solution for synchronizing shared resources. Yet, mishandling locks and threads can result in contention, leading to performance deterioration and compromising the scalability of software applications. In this study, several machine learning models were evaluated on how well they could detect the Java lock contention anti-pattern that caused the lock contention fault based on run time performance data. We trained the machine learning models with performance data generated from the execution of eight Java lock contention anti-patterns and tested the prediction of the models against 30% of the training data as well as performance data from six applications in the Dacappo benchmark that exhibit lock contention. Our results show that we can accurately identify the lock contention anti-pattern based on runtime performance data with an accuracy close to 90%.
Aritra Ahmed, Ramiro Liscano, Akramul Azim, Yee-Kang Chang, Vijay Sundaresan
AST5
2024 Optimistic Stack Allocation and Dynamic Heapification for Managed Runtimes
abstract
The runtimes of managed object-oriented languages such as Java allocate objects on the heap, and rely on automatic garbage collection (GC) techniques for freeing up unused objects. Most such runtimes also consist of just-in-time (JIT) compilers that optimize memory access and GC times by employing escape analysis: an object that does not escape (outlive) its allocating method can be allocated on (and freed up with) the stack frame of the corresponding method. However, in order to minimize the time spent in JIT compilation, the scope of such useful analyses is quite limited, thereby restricting their precision significantly. On the contrary, even though it is feasible to perform precise program analyses statically, it is not possible to use their results in a managed runtime without a closed-world assumption. In this paper, we propose a static+dynamic scheme that allows one to harness the results of a precise static escape analysis for allocating objects on stack, while taking care of both soundness and efficiency concerns in the runtime. Our scheme comprises of three key ideas. First, using the results of a statically performed escape analysis, it performs optimistic stack allocation during JIT compilation. Second, it handles the challenges associated with features that may invalidate the optimism, using a novel idea of dynamic heapification. Third, it uses another novel notion of stack ordering, again supported by a static analysis, to reduce the overheads associated with the checks that determine the need for heapification. The static and the runtime components of our approach are implemented in the Soot optimization framework and in the tiered infrastructure of the Eclipse OpenJ9 VM, respectively. To evaluate the benefits, we compare our scheme with the existing escape analysis and find that it succeeds in allocating a much larger number of objects on the stack. Furthermore, the enhanced stack allocation leads to a significant reduction in the number of GC cycles and brings decent performance improvements, especially suited for constrained-memory environments.
Aditya Anand 0002, Solai Adithya, Swapnil Rustagi, Priyam Seth, Vijay Sundaresan, Daryl Maier, V. Krishna Nandivada, Manas Thakur
Proc. ACM Program. Lang.5
2024 The ART of Sharing Points-to Analysis: Reusing Points-to Analysis Results Safely and Efficiently
abstract
Data-flow analyses like points-to analysis can vastly improve the precision of other analyses, and enable powerful code optimizations. However, whole-program points-to analysis of large Java programs tends to be expensive – both in terms of time and memory. Consequently, many compilers (both static and JIT) and program-analysis tools tend to employ faster – but more conservative – points-to analyses to improve usability. As an alternative to such trading of precision for performance, various techniques have been proposed to perform precise yet expensive fixed-point points-to analyses ahead of time in a static analyzer, store the results, and then transmit them to independent compilation/program-analysis stages that may need them. However, an underlying concern of safety affects all such techniques – can a compiler (or program analysis tool) trust the points-to analysis results generated by another compiler/tool? In this work, we address this issue of trust in the context of Java, while accounting for the issue of performance. We propose ART : Analysis-Results Representation Template – a novel scheme to efficiently and concisely encode results of flow-sensitive, context-insensitive points-to analysis computed by a static analyzer for use in any independent system that may benefit from such a precise points-to analysis. ART also allows for fast regeneration of the encoded sound analysis results in such systems. Our scheme has two components: (i) a producer that can statically perform expensive points-to analysis and encode the same concisely, (ii) a consumer that, on receiving such encoded results (called art work), can regenerate the points-to analysis results encoded by the art work if it is deemed “safe”. The regeneration scheme completely avoids fixed-point computations and thus can help consumers like static analyzers and JIT compilers to obtain precise points-to information without paying a prohibitively high cost. We demonstrate the usage of ART by implementing a producer (in Soot) and two consumers (in Soot and the Eclipse OpenJ9 JIT compiler). We have evaluated our implementation over various benchmarks from the DaCapo and SPECjvm2008 suites. Our results demonstrate that using ART, a consumer can obtain precise flow-sensitive, context-insensitive points-to analysis results in less than (average) 1% of the time taken by a static analyzer to perform the same analysis, with the storage overhead of ART representing a small fraction of the program size (average around 4%).
Shashin Halalingaiah, Vijay Sundaresan, Daryl Maier, V. Krishna Nandivada
Proc. ACM Program. Lang.2
2023 A Lock Contention Classifier Based on Java Lock Contention Anti-Patterns
abstract
Locks are essential in multi-threaded applications as they provide a solution to synchronization of shared resources. However, improper management of locks and threads can lead to contention and surface as run-time performance degradation in the application. Nowadays, performance engineers use legacy tools and their experience to determine causes of lock contention but it takes significant expertise to use these tools. In this paper, a data clustering approach is presented to help identify lock contention faults. The classifier is trained leveraging run-time performance data acquired from a catalog of lock contention Java anti-patterns and code smells. The K-means unsupervised classifier algorithm was used to create the classification model and the results show that lock contentions can be classified into three clusters that can be identified into those caused by a) threads spending too much time inside the critical section, b) threads blocked because of high frequency access requests, and c) threads with a low contention. This classifier is intended to be used to help tailor recommendations to the developer based on the lock contention anti-patterns and type of lock contention.
Ramiro Liscano, Aritra Ahmed, Joseph Robertson, Akramul Azim, Vijay Sundaresan, Yee-Kang Chang
ICMLA5
2022 Mining Annotation Usage Rules: A Case Study with MicroProfile
abstract
While Application Programming Interfaces (APIs) allow easier reuse of existing functionality, developers might make mistakes in using these APIs (a.k.a. API misuses). If an API usage specification exists, then automatically detecting such misuses becomes feasible. Since manually encoding specifications is a tedious process, there has been a lot of research regarding pattern-based specification mining. However, while annotations are widely used in Java enterprise microservices frameworks, most of these pattern-based rule discovery techniques have not considered annotation-based API usage rules. In this industrial case study of MicroProfile, an open-source Java microservices framework developed by IBM and others, we investigate whether the idea of pattern-based discovery of rules can be applied to annotation-based API usages. We find that our pattern-based approach mines 23 candidate rules, among which 4 are fully valid specifications and 8 are partially valid specifications. Overall, our technique mines 12 valid rules, 10 of which are not even documented in the official MicroProfile documentation. To evaluate the usefulness of the mined rules, we scan MicroProfile client projects for violations. We find 100 violations of 5 rules in 16 projects. Our results suggest that the mined rules can be useful in detecting and preventing annotation-based API misuses.
Batyr Nuryyev, Ajay Kumar Jha, Sarah Nadi, Yee-Kang Chang, Emily Jiang, Vijay Sundaresan
ICSME6
2022 Lock Contention Performance Classification for Java Intrinsic Locks
Nahid Hasan Khan, Joseph Robertson, Ramiro Liscano, Akramul Azim, Vijay Sundaresan, Yee-Kang Chang
RV5
2022 JITServer: Disaggregated Caching JIT Compiler for the JVM in the Cloud
Alexey Khrabrov 0002, Marius Pirvu, Vijay Sundaresan, Eyal de Lara
USENIX ATC3
2013 Experiences in designing a robust and scalable interpreter profiling framework
abstract
Profile directed feedback (PDF) is a well known technique used to drive many compiler optimizations like basic block ordering and guarded devirtualization. These optimizations are particularly crucial in order to achieve good throughput performance in JEE applications that have a large code footprint. To effectively apply optimizations that rely on profiling information, a just-in-time (JIT) compiler must have access to profiling information that is accurate. One common source of profiling information in a Java virtual machine (JVM) is the interpreter. Typically methods are interpreted as a program ramps up, during which profiling information can be collected. However, obtaining useful and accurate information for large enterprise-class applications can be a challenge because of the memory and performance overhead associated with collecting and processing the large volumes of profiling data that is generated. This paper describes the challenges in maintaining the balance between throughput performance and profiling overhead in a production JIT compiler that is used by the IBM JDK. The scope of the performance overhead in terms of throughput, memory footprint and startup speed for large JEE class applications is introduced and various engineering solutions that were tried are detailed and compared in terms of experimental results. We found that the throughput improvement due to interpreter profiling (IP) can be as high as 58%, whereas the overhead measured in terms of application startup time could cost up to 57%. Our solutions to reducing profiling overhead managed to reduce the startup cost to only a few percent while maintaining the full throughput benefit. By discussing these approaches, this paper offers a balanced and practical overview on how to make PDF work well for enterprise-class applications in a production JIT compiler.
Ian Gartley, Marius Pirvu, Vijay Sundaresan, Nikola Grcevski
CGO3
2008 Removing redundancy via exception check motion
abstract
Partial redundancy elimination aims to reduce the number of times an expression is computed more than once. The traditional Lazy Code Motion (LCM) algorithm formulated by Knoop, Ruthing and Steffen, through its reliance on unordered bit vectors, is severely limited in its ability to remove redundancy when precise exception semantics are required because bit vectors cannot express the order of exception checks. This paper describes our new PRE algorithm Exception Check Motion that uses the LCM algorithm to treat and optimize exception checks in a similar way to any other expression. Unlike earlier techniques that can remove only the compare instruction of a partially redundant exception check, our solution can eliminate both the compare and trap instructions without any run time code patching or expensive recovery operations. Since it is the trap instructions that restrict subsequent code motions, our technique gives downstream optimizations more flexibility to improve the performance of the resulting code once the partially redundant checks are eliminated. Our analysis has been implemented in the IBM® Testarossa (TR) just-in-time (JIT) compiler in the IBM Developer Kit for Java Release 5.0 as part of the J9 Virtual Machine. We measure performance improvements up to 7.6% and averaging 2.5% across 22 SPEC and DaCapo benchmarks on 4-way IBM pSeries (PowerPC) hardware.
Vijay Sundaresan, Mark G. Stoodley, Pramod Ramarao
CGO1
2007 Power invariant secure IC design methodology using reduced complementary dynamic and differential logic
abstract
Security of cryptographic devices (secure ICs) like smart cards has come under threat from powerful side channel attacks like Differential Power Analysis (DPA). DPA uses power consumption information leaked from the secure IC in conjunction with statistical correlation techniques to retrieve the secret key stored in the secure IC. The most effective countermeasure to resist DPA attacks is to make the power consumption of the secure IC invariant, hence uncorrelated to the input data (secret key). In hardware implementations, this can be achieved by designing the secure IC using Dynamic and Differential Logic (DDL) style. In this paper, we present a novel methodology to design DPA-resistant power invariant secure ICs using Reduced Complementary Dynamic and Differential Logic (RCDDL). The proposed methodology involves strategies to design: 1) RCDDL gates, and 2) secure circuits using RCDDL gates. Experiments show significant improvements in security strength, average power consumption and area, when compared with a similar secure DDL and non-secure static-CMOS logic design styles.
Vijay Sundaresan, Srividhya Rammohan, Ranga Vemuri
VLSI-SoC1
2006 Experiences with Multi-threading and Dynamic Class Loading in a Java Just-In-Time Compiler
abstract
In this paper, we describe the techniques that have been implemented in the IBM TestaRossa (TR) just-in-time (JIT) compiler to safely perform aggressive code patching and collect accurate profiles in the context of a Java application employing multiple threads and dynamic class loading and unloading. Previous work in these areas either did not account for the synchronization cost of safety or dynamic class loading/unloading effects in a heavily multithreaded program or did not consider how different patching techniques may be required for different platforms where instruction cache coherence guarantees vary. We evaluate the space and time overhead to make our profiling framework correct, showing that privatizing the profiling variables to achieve correctness impacts execution time only minimally but it can grow the stack frames for profiled methods by less than 15% on average for the SPECjvm98 and SPECjbb2000 benchmarks. Since methods are profiled for only a brief time and the stack frames themselves are not large, we do not consider this growth to be prohibitive. The techniques reported in this paper are implemented in the 1.5.0 release of the IBM Developer Kit for Java targeting 12 different processor-operating system platforms.
Vijay Sundaresan, Daryl Maier, Pramod Ramarao, Mark G. Stoodley
CGO1
2005 Automatically Reducing Repetitive Synchronization with a Just-in-Time Compiler for Java
abstract
We describe an automatic technique to remove repetitive synchronization in Java/spl trade/ programs by removing selected MONITORENTER/EXIT operations. Once these operations are removed, parts of a method that were not originally locked become protected by a lock. If it is unsafe to synchronize the code between the original locked regions, however, the code is not transformed. Scalability is also protected by not allowing a lock to be held for a significantly longer time than it would be held in the original program code. Our base algorithm improved the throughput of the industry-standard SPECjbb2000 benchmark by 2% to 5% for three different platforms. We also describe an extension to our algorithm to better handle virtual calls, which are prevalent in Java code, and this extension provides up to a further 5% improvement. Our computationally efficient algorithm was implemented and evaluated in a production Just-In-Time (JIT) compiler.
Mark G. Stoodley, Vijay Sundaresan
CGO2
2000 Optimizing Java Bytecode Using the Soot Framework: Is It Feasible?
Raja Vallée-Rai, Etienne M. Gagnon, Laurie J. Hendren, Patrick Lam 0001, Patrice Pominville, Vijay Sundaresan
CC6
2000 Practical virtual method call resolution for Java
abstract
This paper addresses the problem of resolving virtual method and interface calls in Java bytecode. The main focus is on a new practical technique that can be used to analyze large applications. Our fundamental design goal was to develop a technique that can be solved with only one iteration, and thus scales linearly with the size of the program, while at the same time providing more accurate results than two popular existing linear techniques, class hierarchy analysis and rapid type analysis.We present two variations of our new technique, variable-type analysis and a coarser-grain version called declared-type analysis. Both of these analyses are inexpensive, easy to implement, and our experimental results show that they scale linearly in the size of the program.We have implemented our new analyses using the Soot frame-work, and we report on empirical results for seven benchmarks. We have used our techniques to build accurate call graphs for complete applications (including libraries) and we show that compared to a conservative call graph built using class hierarchy analysis, our new variable-type analysis can remove a significant number of nodes (methods) and call edges. Further, our results show that we can improve upon the compression obtained using rapid type analysis.We also provide dynamic measurements of monomorphic call sites, focusing on the benchmark code excluding libraries. We demonstrate that when considering only the benchmark code, both rapid type analysis and our new declared-type analysis do not add much precision over class hierarchy analysis. However, our finer-grained variable-type analysis does resolve significantly more call sites, particularly for programs with more complex uses of objects.
Vijay Sundaresan, Laurie J. Hendren, Chrislain Razafimahefa, Raja Vallée-Rai, Patrick Lam 0001, Etienne M. Gagnon, Charles Godin
OOPSLA1