Kshitiz Malik

dblp:42/5439 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
6since 2021 · last 2024
0009-0003-9300-6793ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 59% Transfer learning and domain adaptation · 22% Optimization for machine learning · 19%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Processor architecture and microarchitecture · 78% Energy-efficient computing · 10% Parallel and multicore computing · 9%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
1.222023
Where to Begin? On the Impact of Pre-Training and Initialization in Federated Learning · ICLR 2023
Federated Learning with Partial Model Personalization · ICML 2022
Human-AI interaction
intelligent assistant
0.712023
Towards Next-Generation Intelligent Assistants Leveraging LLM Techniques · KDD 2023
Human-AI interaction › intelligent assistant
virtual assistants
0.712023
Towards Next-Generation Intelligent Assistants Leveraging LLM Techniques · KDD 2023
Machine learning › Optimization for machine learning
convergence analysis
0.612022
Federated Learning with Partial Model Personalization · ICML 2022
Machine learning › Efficient and distributed learning › federated learning
model personalization
0.612022
Federated Learning with Partial Model Personalization · ICML 2022
Processor architecture and microarchitecture
branch prediction
0.222008
Fetch-Criticality Reduction through Control Independence · ISCA 2008
PaCo: Probability-based path confidence prediction · HPCA 2008
Processor architecture and microarchitecture › instruction-level parallelism
control independence
0.222008
Fetch-Criticality Reduction through Control Independence · ISCA 2008
Branch-mispredict level parallelism (BLP) for control independence · HPCA 2008
Processor architecture and microarchitecture › branch prediction
branch misprediction
0.112008
Branch-mispredict level parallelism (BLP) for control independence · HPCA 2008
Processor architecture and microarchitecture › branch prediction
branch misprediction recovery
0.112008
Fetch-Criticality Reduction through Control Independence · ISCA 2008
Processor architecture and microarchitecture
instruction-level parallelism
0.112008
Branch-mispredict level parallelism (BLP) for control independence · HPCA 2008
Energy-efficient computing › power management › fine-grain power management
pipeline gating
0.112008
PaCo: Probability-based path confidence prediction · HPCA 2008
Concurrent programming
speculative execution
0.112007
Exploiting Postdominance for Speculative Parallelization · HPCA 2007
Parallel and multicore computing
speculative parallelization
0.112007
Exploiting Postdominance for Speculative Parallelization · HPCA 2007
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.012008
PaCo: Probability-based path confidence prediction · HPCA 2008
Processor architecture and microarchitecture
speculative execution
0.012008
Branch-mispredict level parallelism (BLP) for control independence · HPCA 2008

Methods — techniques the papers use, named apart from their topics

weight initialization · 0.7pre-training · 0.7large language model · 0.7non-convex optimization · 0.6convergence analysis · 0.6simulation · 0.2immediate postdominance · 0.1dynamic reconvergence prediction · 0.1probability estimation · 0.1prediction · 0.1criticality modeling · 0.1
YearPublicationVenuePosition
2024 Effective Long-Context Scaling of Foundation Models
abstract
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001
NAACL-HLT15
2023 Joint Federated Learning and Personalization for on-Device ASR
abstract
In this paper, we propose a joint federated learning (FL) and personalization method for on-device ASR adaptation. Starting with a Conformer-based RNN-T as the ASR model backbone that is pretrained on public data and shared across devices, we propose to adapt to user data on-device by ① collectively finetuning the backbone on all user data by FL, and ② for each user, augmenting the backbone with a personalized adapter that is trained on on-device data and stored locally on their devices. As ground-truth transcriptions are not available, we use pseudo-label training, which can be completely performed on-device. Our joint method combines the best of both FL and personalization and achieves maximum effect for both heavy and light users: For users with 50+ adaptation utterances, our recipe achieves $27.4 \%$ relative WER reduction, largely due to personalization; for users with no adaptation utterance, our recipe achieves $8.9 \%$ relative WER reduction purely due to FL.
Junteng Jia, Ke Li 0023, Mani Malek 0001, Kshitiz Malik, Jay Mahadeokar, Ozlem Kalinli, Frank Seide
ASRU4
2023 Where to Begin? On the Impact of Pre-Training and Initialization in Federated Learning
John Nguyen, Kshitiz Malik, Maziar Sanjabi, Michael G. Rabbat
ICLR3
2023 Towards Next-Generation Intelligent Assistants Leveraging LLM Techniques
abstract
Virtual Intelligent Assistants take user requests in the voice form, perform actions such as setting an alarm, turning on a light, and answering a question, and provide answers or confirmations in the voice form or through other channels such as a screen. Assistants have become prevalent in the past decade, and users have been taking services from assistants like Amazon Alexa, Apple Siri, Google Assistant, and Microsoft Cortana.
Xin Dong 0001, Seungwhan Moon, Yifan Ethan Xu, Kshitiz Malik, Zhou Yu 0005
KDD4
2022 Federated Learning with Buffered Asynchronous Aggregation
abstract
Scalability and privacy are two critical concerns for cross-device federated learning (FL) systems. In this work, we identify that synchronous FL – cannot scale efficiently beyond a few hundred clients training in parallel. It leads to diminishing returns in model performance and training speed, analogous to large-batch training. On the other hand, asynchronous aggregation of client updates in FL (i.e., asynchronous FL) alleviates the scalability issue. However, aggregating individual client updates is incompatible with Secure Aggregation, which could result in an undesirable level of privacy for the system. To address these concerns, we propose a novel buffered asynchronous aggregation method, FedBuff, that is agnostic to the choice of optimizer, and combines the best properties of synchronous and asynchronous FL. We empirically demonstrate that FedBuff is $3.3\times$ more efficient than synchronous FL and up to $2.5\times$ more efficient than asynchronous FL, while being compatible with privacy-preserving technologies such as Secure Aggregation and differential privacy. We provide theoretical convergence guarantees in a smooth non-convex setting. Finally, we show that under differentially private training, FedBuff can outperform FedAvgM at low privacy settings and achieve the same utility for higher privacy settings.
John Nguyen, Kshitiz Malik, Hongyuan Zhan, Ashkan Yousefpour, Michael G. Rabbat, Mani Malek 0001, Dzmitry Huba
AISTATS2
2022 Federated Learning with Partial Model Personalization
abstract
We consider two federated learning algorithms for training partially personalized models, where the shared and personal parameters are updated either simultaneously or alternately on the devices. Both algorithms have been proposed in the literature, but their convergence properties are not fully understood, especially for the alternating variant. We provide convergence analyses of both algorithms in the general nonconvex setting with partial participation and delineate the regime where one dominates the other. Our experiments on real-world image, text, and speech datasets demonstrate that (a) partial personalization can obtain most of the benefits of full model personalization with a small fraction of personal parameters, and, (b) the alternating update algorithm outperforms the simultaneous update algorithm by a small but consistent margin.
Krishna Pillutla, Kshitiz Malik, Abdel-rahman Mohamed, Michael G. Rabbat, Maziar Sanjabi, Lin Xiao 0003
ICML2
2008 PaCo: Probability-based path confidence prediction
abstract
A path confidence estimate indicates the likelihood that the processor is currently fetching correct path instructions. Accurate path confidence prediction is critical for applications like pipeline gating and confidence-based SMT fetch prioritization. Previous work in this domain uses a threshold-and-count predictor, where the number of unresolved, low-confidence branches serves as an estimate of path confidence. This approach is inaccurate since it implicitly assumes that all low-confidence branches have the same mispredict rate, and that high-confidence branches never mispredict. We propose an alternative path confidence predictor designed from first principles, called PaCo, that directly estimates the probability that the processor is on the goodpath, and considers contributions from all branches, both high and low confidence. Even though it uses only modest hardware, PaCo can estimate the processor’s goodpath likelihood with very high accuracy, with an RMS error of 3.8%. We show that PaCo significantly outperforms threshold-and-count predictors in pipeline gating and SMT fetch prioritization. In pipeline gating, while the best conventional predictor can reduce badpath instructions executed by 7% with a small loss in performance, PaCo can reduce bad-path instructions by 32% without any performance loss. In SMT fetch prioritization, using PaCo instead of conventional path confidence predictors improves performance by up to 23%, and 5.5% on average.
Kshitiz Malik, Mayank Agarwal, Vikram Dhar, Matthew I. Frank
HPCA1
2008 Branch-mispredict level parallelism (BLP) for control independence
abstract
A microprocessorpsilas performance is fundamentally limited by the rate at which it can resolve branch mispredictions. Control independence (CI) architectures look for useful control and data independent instructions to fetch and execute in the shadow of a branch misprediction. This paper demonstrates that CI architectures can be guided to exploit substantial branch-mispredict level parallelism (BLP) in existing control intensive applications. A program has branch-mispredict level parallelism when its dynamic execution trace contains hard-to-predict branches that are both control and data independent, and thus could, potentially, be resolved in parallel. Although applications have a high degree of inherent BLP, we find that the amount of BLP exploited by naive CI architectures tends to be quite small. We show that spawn selection and data dependence handling policies in a CI architecture should make choices that explicitly aim to maximize branch-mispredict level parallelism. We demonstrate that with BLP-focussed policies, CI architectures can expose high amounts of branch-mispredict level parallelism and achieve 50% to 90% improvements in IPC on several of the SPEC 2000 Integer benchmarks.
Kshitiz Malik, Mayank Agarwal, Sam S. Stone, Kevin M. Woley, Matthew I. Frank
HPCA1
2008 Fetch-Criticality Reduction through Control Independence
abstract
Architectures that exploit control independence (CI) promise to remove in-order fetch bottlenecks, like branch mispredicts, instruction-cache misses and fetch unit stalls, from the critical path of single-threaded execution. By exposing more fetch options, however, CI architectures also expose more performance tradeoffs. These tradeoffs make it hard to design policies that deliver good performance. This paper presents a criticality-based model for reasoning about CI architectures, and uses that model to describe the tradeoffs between gains from control independence versus increased costs of honoring data dependences. The model is then used to derive the design of a criticality-aware task selection policy that strikes the right balance between fetch-criticality and execute-criticality. Finally, the paper validates the model by attacking branch-misprediction induced fetch-criticality through the above derived spawn policy. This leads to as high as 100% improvements in performance, and in the region of 40% or more improvements for four of the benchmarks where this is the main problem. Criticality analysis shows that this improvement arises due to reduced fetch-criticality.
Mayank Agarwal, Nitin Navale, Kshitiz Malik, Matthew I. Frank
ISCA3
2007 Exploiting Postdominance for Speculative Parallelization
abstract
Task-selection policies are critical to the performance of any architecture that uses speculation to extract parallel tasks from a sequential thread. This paper demonstrates that the immediate postdominators of conditional branches provide a larger set of parallel tasks than existing task-selection heuristics, which are limited to programming language constructs (such as loops or procedure calls). Our evaluation shows that postdominance-based task selection achieves, on average, more than double the speedup of the best individual heuristic, and 33% more speedup than the best combination of heuristics. The specific contributions of this paper include, first, a description of task selection based on immediate post-dominance for a system that speculatively creates tasks. Second, our experimental evaluation demonstrates that existing task-selection heuristics based on loops, procedure calls, and if-else statements are all subsumed by compiler-generated immediate postdominators. Finally, by demonstrating that dynamic reconvergence prediction closely approximates immediate postdominator analysis, we show that the notion of immediate postdominators may also be useful in constructing dynamic task selection mechanisms
Mayank Agarwal, Kshitiz Malik, Kevin M. Woley, Sam S. Stone, Matthew I. Frank
HPCA2