Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST ( Adaptive Layer Selection for ViT Fine-Tuning ), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.