
I am a fourth-year Ph.D. student at Georgia Tech, advised by Professor Moinuddin Qureshi in the FAST Lab. My research focuses on improving the performance of CPU and GPU workloads by leveraging the intricacies of modern memory systems and proposing architectural enhancements to overcome memory bottlenecks. Currently, I work on low-latency LLM inference, where I focus on the communication and interconnect bottlenecks that arise in distributed serving.
News
- Nov, 2025
- Our paper MIRZA: Efficiently Mitigating Rowhammer with Randomization and ALERT accepted at HPCA 2026
- Nov, 2025
- I’ll be at Meta in the Bay Area next summer, working with the AI and Systems Co-Design team!
- Aug, 2025
- Finished Summer Internship at AMD Research!
- March, 2025
- Our paper DREAM: Enabling Low-Overhead Rowhammer Mitigation via Directed Refresh Management accepted at ISCA 2025
- Feb, 2025
- I will be in Austin at AMD RAD working on GPU memory systems during the upcoming summers!
- Aug, 2024
- Finished Summer Internship at Astera Labs!
- Aug, 2023
- Our paper Hot Pixels: Frequency, Power, and Temperature Attacks on GPUs and ARM SoCs appeared at Usenix Security 2023
- Apple assigned us CVE-2023-38599 for Hot Pixels.
- Aug, 2022
- Started Ph.D at Georgia Tech
- Jan, 2020
- Started as a Low Latency Infrastructure Developer in Plutus Research, Bangalore, India
- Dec, 2019
- Graduated with B.Tech in CS from IIT Kanpur
- Jul, 2019
- Finished 2 months Summer Internship at Nutanix, Bangalore, India
Publications
BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference
arXiv:2609.13592, September 2026Preprint
From Fleet to Lab: Revisiting the Security and Complexity of Industrial Rowhammer Mitigation
arXiv:2608.26072, August 2026Preprint
SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference
arXiv:2607.08973, July 2026Preprint
TileLens: Efficiently Using Large-Granularity Memory Systems with Transparent Two-Dimensional Memory Layout
arXiv:2607.04031, July 2026Preprint
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance
arXiv:2606.09441, June 2026Preprint
DREAM: Enabling Low-Overhead Rowhammer Mitigation via Directed Refresh Management
ISCA 2025
Utility-Driven Speculative Decoding for Mixture-of-Experts
arXiv:2506.20675, June 2025Preprint
RogueRFM: Attacking Refresh Management for Covert-Channel and Denial-of-Service
arXiv:2501.06646, January 2025Preprint