← All writing
PhD applications

My PhD Statement of Purpose

The statement I submitted to UT Austin, paired with general lessons that apply when writing a research-focused SOP for any university.

My core research interests are in Generative Models and Efficient ML. The computational and memory demands of state-of-the-art foundation models create significant barriers to their practical, sustainable, and democratized deployment. As a pre-doctoral researcher at Google DeepMind (GDM), I have been actively working to mitigate the high computational and memory demands of frontier generative models, specifically Gemini and Veo. This experience has given me an invaluable perspective on critical challenges inherent in scaling large-scale foundation models. By coupling these internal contributions at GDM with published research (ICML [1]) on efficient inference of generative models, I have solidified my long-term goal to drive foundational research that significantly advances the efficiency-quality frontier.

Recent research [5], [6] has observed that emerging properties arise in multimodal models as we scale and integrate diverse learning objectives, including multi-modality and multi-tasking. Specifically, these models demonstrate the ability to generalize to unseen tasks, effectively acting as unified, generalist models. While unified models show promise, due to their prohibitive costs, specialized models remain the practical standard. My goal is to bridge this gap between current systems and true unified, multimodal models by focusing on the following key areas: i) Efficient Architectures and Algorithms: Fundamentally mitigating the extreme computational and memory demands of foundation models to make generalist capabilities accessible. ii) Multimodal Reasoning: Going beyond natural language for robust reasoning, investigating how distinct modalities can synergistically enhance the reasoning process of unified models. iii) Long-Horizon Generation: Developing ability to efficiently generate long content while maintaining consistency. iv) Embodied Intelligence: Extending these unified capabilities to embodied agents, enabling real-time perception and planning. My research at GDM, under the guidance of Dr. Prateek Jain, Dr. Sujoy Paul, and Dr. Aditya Kusupati, directly tackles this challenge by developing a spectrum of efficiency techniques, ranging from adaptive computation (inference-time efficiency) to model compression (training-time efficiency) algorithms.

Adaptive Compute for Visual Generation

A critical bottleneck of visual generative models (such as diffusion and masked generative models) is their iterative sampling process, which incurs high computational costs during inference. While recent research in this area focuses on reducing the number of sampling steps, it still treats all steps as equally complex. Our orthogonal approach is built on a key insight: not all sampling steps require the same computational cost. In this work, we used nested transformers [7] to allocate different amount of compute to different steps of the generation process. In a single training run, we extract a series of high quality, aligned nested models of different sizes. We perform online distillation, progressively from a nested model to its immediate smaller counterpart. We use smaller nested models for easier sampling steps and larger ones for difficult ones. Our method (ICML [1]) led to ~3x speedup for MaskGIT [8] on academic benchmarks without compromising visual quality or increasing memory footprint. Translating these lessons from my research, I have been implementing similar efficiency techniques for Google’s SOTA video generation model, Veo. These large-scale experiments have provided me with end-to-end hands-on experience in optimizing a flagship model.

Motivated by the positive findings from the aforementioned project, I wanted to push the boundaries of parameter-efficiency for ultra resource constrained scenarios such as on-device visual generation. We exploited recursive transformers and replaced deep transformer stacks with a single, recursively applied fixed point layer to maximize parameter efficiency. We developed a stochastic joint optimization strategy to stabilize this recurrent architecture, yielding a family of free elastic student models from a single training run. We devise an adaptive inference schedule that allocates compute (fixed point iterations) based on sampling step difficulty, further optimizing the speed-quality trade-off.

Compressing Large-Scale Pre-trained Models

Having gained immense experience optimizing frontier foundational models, I sought to use this deep expertise for leading a project end-to-end. I wanted to define the strategic direction for our next efficiency initiative and lead its execution. Given that Google already has large pre-trained models, these models can be effectively utilized to derive compressed models at minimal extra training cost. These derived models offer a two-fold advantage. First, they achieve significantly better quality compared to models with equivalent parameters (iso-params) and training compute (iso-FLOPs). Second, they resolve the extreme memory footprint inherent to flagship models, enabling broader application and democratized deployment. I subsequently proposed and initiated a project focused on the compression of large pre-trained Mixture of Experts (MoE) models. Existing compression literature largely relies on naive merging approaches such as weighted combination or SVD, effective for dense models but scale poorly when applied to the sparse, distributed nature of MoE. My research proposes a stage-wise approach to compression: structured pruning for strong initialization and layer-wise distillation for feature alignment, in compute-constrained settings. These techniques are key exploratory pathways for optimizing Google’s flagship Gemini model.

Automated Design Evaluation and Refinement

My early career research was instrumental in shaping my core research philosophy, spanning diverse domains and fostering key research insights. My internship at Adobe Research focused on efficient architectures in the complex, semantic visual domain of graphic designs, where I developed an automated framework for scoring and refinement of graphic designs. Initially, I made the common early-career mistake of chasing trends, attempting to leverage computationally expensive Vision-Language Models (VLMs). However, I quickly learned that complexity does not equate to effectiveness, particularly because these VLMs were primarily optimized for semantic alignment and lacked the deep compositional reasoning required to evaluate nuanced graphic designs. This led me to pivot my research methodology. Instead of scaling up, I focused on deep problem-specific engineering. Under the guidance of Dr. Joseph KJ and Dr. Balaji Vasan Srinivasan, I designed a highly efficient CNN-based architecture (0.5M parameters) by meticulously crafting the dataset and loss function. This tiny specialized model significantly outperformed multimodal LLMs like GPT-4o and LLaVA-NeXT (7B) in ranking graphic designs. My undergraduate background in genetic algorithms was instrumental in advancing the refinement process, leading to the first unified framework (WACV [2]) to score and refine graphic designs, with potential integration into Adobe Express.

Multimodal Talking Face Generation

My first research project started with an undergraduate interest in talking face generation. This interest was sparked while watching a dubbed movie, highlighting the critical need for accurate lip synchronization and natural visual expression. Driven by this challenge, I contacted MIDAS-IIITD Lab, leveraging Dr. Rajiv Ratn Shah’s expertise in multimodal learning to initiate a project. We identified a critical gap that the existing literature focused solely on lip-synchronization while ignoring visual emotions. We started building a talking face generation network that can effectively fuse the information from three different modalities (vision, speech, and emotion). My initial approach, a direct concatenation of embeddings from different modalities, was simple but ineffective. I subsequently realized that we could concatenate the embeddings sequentially in a skip-connections style, which removed the black dot artifacts from the results. Our approach, the first to integrate emotions in face generation, resulted in Best Demo Paper [4] at ACM MM Asia ’22 and subsequent presentations at workshops in ACM MM ’23 [3] and ICCV ’23.

I effectively prioritized and executed research alongside demanding coursework and published the above early-career research works. Graduating with department rank 1 required disciplined time management and a relentless focus on both academic and research excellence, preparing me for the intense demands of doctoral research. Further demonstrating my commitment to the research community, I have actively served as a reviewer for several top-tier machine learning conferences (ICML, CVPR, ICLR, ICCV, AAMAS, and WACV) while working at GDM.

PhD at UT Austin

The University of Texas at Austin offers an ideal environment for my goal of bridging the gap between resource-intensive foundation models and accessible, unified intelligence. I am particularly drawn to Prof. Kristen Grauman’s and Philipp Krähenbühl’s pioneering work in vision and efficient representation learning. Furthermore, I aim to leverage Prof. Chenfeng Xu’s expertise in efficient generative models, building upon my experience with adaptive compute and model compression. The collaborative ecosystem at UT Austin is the perfect place to drive the next generation of efficient, generalist models.

References from the submitted SOP

  1. Sahil Goyal et al. “Masked Generative Nested Transformers with Decode Time Scaling.” ICML 2025 and ICLR DeLTa Workshop 2025.
  2. Sahil Goyal*, Abhinav Mahajan*, Swasti Mishra, Prateksha Udhayanan, KJ Joseph, Balaji Vasan Srinivasan. “Design-o-meter: Towards Evaluating and Refining Graphic Designs.” WACV 2025.
  3. Sahil Goyal, Sarthak Bhagat, Shagun Uppal, Yi Yu, Yifang Yin, Rajiv Ratn Shah. “Emotionally Enhanced Talking Face Generation.” ICCV CVEU Workshop 2023 and ACM MM McGE Workshop 2023.
  4. Sahil Goyal, Sarthak Bhagat, Shagun Uppal, Dhroov Goel, Sakshat Mali, Yi Yu, Yifang Yin, Rajiv Ratn Shah. “Emotional Talking Faces: Making Videos More Expressive and Realistic.” Best Demo Paper, ACM MM Asia 2022.
  5. Thaddäus Wiedemer*, Yuxuan Li, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, Robert Geirhos*. “Video models are zero-shot learners and reasoners.” arXiv 2025.
  6. Chaorui Deng*, Deyao Zhu*, Kunchang Li*, Chenhui Gou*, Feng Li*, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan. “Emerging Properties in Unified Multimodal Pretraining.” arXiv 2025.
  7. Aditya Kusupati*, Gantavya Bhatt*, Aniket Rege*, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, Ali Farhadi. “Matryoshka Representation Learning.” NeurIPS 2022.
  8. Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, William T. Freeman. “MaskGIT: Masked Generative Image Transformer.” CVPR 2022.