PEFT: Parameter-Efficient Fine-Tuning Guide (2026)
A practitioner's guide to PEFT (Parameter-Efficient Fine-Tuning) covering every major method, practical code examples, hardware requirements, and decision frameworks for LLM customization.

PEFT, or Parameter-Efficient Fine-Tuning, covers a family of techniques for adapting large pretrained models to specific tasks. The core idea: train only a small fraction of parameters (typically under 1%) and freeze everything else. Instead of retraining billions of weights at massive compute cost, you inject or select a tiny set of trainable parameters that get you close to full fine-tuning performance. This has become the default path for LLM customization. GPU memory requirements drop sharply, and checkpoint sizes go from tens of gigabytes to a few megabytes per task.
The term also refers to the Hugging Face PEFT library, an open-source Python package (Apache-2.0, 21k+ GitHub stars) implementing these techniques with integrations for Transformers, Diffusers, and Accelerate. As of September 2026, the latest stable release is v0.21.1.
Why Parameter Efficient Fine Tuning Matters
Full fine-tuning updates every parameter in a model. For a 70B-parameter LLM, that means storing optimizer states, gradients, and model copies across multiple high-end GPUs. That's a lot of compute for a single downstream task. Most teams can't justify the expense, especially when they need to customize models for several use cases.
PEFT methods fix this by freezing the pretrained weights and introducing a small number of trainable parameters. The original LoRA paper (Hu et al., 2021, Microsoft) demonstrated a 10,000x reduction in trainable parameters and a 3x reduction in GPU memory on GPT-3 175B compared to full fine-tuning with Adam. In practice, PEFT fine tuning retains the large majority of full fine-tuning quality at a fraction of the cost. LoRA's inventor Edward Hu noted that low-rank updates "perform almost as well as full fine tuning in large models like GPT-3."
The benefits stack up fast when you need multiple task-specific models. You keep one frozen base model and swap lightweight adapters (often just a few MB each) per task. No separate multi-gigabyte model copies to store and serve.
PEFT Methods: A Complete Overview
The Hugging Face PEFT library organizes methods into three categories: prompt-based methods, layer tuning, and adapter methods.
Adapter Methods
These add small trainable matrices to existing model layers. They're the most popular category, led by LoRA.
- LoRA (Low-Rank Adaptation) The most widely used PEFT method, accounting for 98.4% of Hugging Face Hub model cards that mention a PEFT technique. LoRA injects low-rank decomposition matrices into transformer layers, bringing trainable parameters down to under 0.5% of the total. If you're new to PEFT, LoRA is the recommended starting point. For a deeper look at how low-rank adapters work, see our guide to LoRA adapters.
- **QLoRA** Combines LoRA with 4-bit NormalFloat (NF4) quantization of the base model via bitsandbytes. This makes it possible to fine-tune 65B+ parameter models on a single GPU. For a side-by-side comparison, see our LoRA vs QLoRA breakdown.
- AdaLoRA Adaptively allocates rank budget across weight matrices based on importance scores, rather than using a fixed rank everywhere.
- IA3 (Infused Adapter by Inhibiting and Amplifying Inner Activations) Learns rescaling vectors for keys, values, and feedforward layers. Even fewer trainable parameters than LoRA, though narrower applicability.
- DoRA (Weight-Decomposed Low-Rank Adaptation) Decomposes weight updates into magnitude and direction components. In the PEFT library, you enable DoRA via the
use_doraflag inLoraConfig. - OFT / BOFT Orthogonal Fine-Tuning methods that preserve the pretrained model's geometric structure during adaptation.
- LoHa / LoKr Use Hadamard and Kronecker products respectively for more expressive low-rank decompositions.
- HiRA / GLoRA Added to the PEFT library in v0.20.0 (July 2026). HiRA (ICLR 2025) multiplies the low-rank product elementwise with frozen base weights for full-rank updates at LoRA parameter counts. GLoRA extends LoRA with configurable weight, activation, and bias adaptation.
Prompt-Based Methods
These prepend learnable tensors to model inputs without modifying model weights.
- Prompt Tuning Adds trainable "soft prompt" tokens to the input embedding. Only the prompt embeddings get updated during training.
- Prefix Tuning Prepends trainable vectors to the key and value states in every transformer layer. This gives the model more capacity than prompt tuning.
- P-Tuning Uses a small encoder network to generate soft prompt embeddings, bridging prompt tuning and prefix tuning.
Layer Tuning
- LayerNorm Tuning Trains only the LayerNorm parameters across the model.
- TrainableTokens Targets specific embedding tokens for training.

How to Set Up PEFT Fine Tuning: Step-by-Step
Installation
pip install peft transformers accelerate datasetsFor QLoRA, also install bitsandbytes:
pip install bitsandbytesRequires Python 3.10+.
Basic LoRA Training
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer
from peft import LoraConfig, get_peft_model, TaskType
# 1. Load the base model
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-1B")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-1B")
# 2. Create LoRA config
peft_config = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=8, # rank -- start with 8-16, scale to 32-64 for complex tasks
lora_alpha=16, # scaling factor, recommended ~2x rank
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05,
bias="none",
inference_mode=False,
)
# 3. Wrap the model
peft_model = get_peft_model(model, peft_config)
peft_model.print_trainable_parameters()
# Output: trainable params: ~4M || all params: ~1.2B || trainable%: 0.33%
# 4. Train with HF Trainer
training_args = TrainingArguments(
output_dir="my-lora-adapter",
learning_rate=1e-3,
per_device_train_batch_size=32,
num_train_epochs=2,
weight_decay=0.01,
save_strategy="epoch",
)
trainer = Trainer(
model=peft_model,
args=training_args,
train_dataset=tokenized_train,
data_collator=data_collator,
)
trainer.train()
# 5. Save -- only adapter weights (a few MB)
peft_model.save_pretrained("my-lora-adapter")QLoRA Setup (4-bit Quantized Training)
import torch
from transformers import BitsAndBytesConfig, AutoModelForCausalLM
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
# Quantization config
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
# Load quantized base model
model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-v0.1",
quantization_config=bnb_config,
device_map="auto",
)
model = prepare_model_for_kbit_training(model)
# Apply LoRA on top
config = LoraConfig(
r=16,
lora_alpha=8,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)This gets a 7B model fine-tuning on a consumer RTX 4090.
Loading and Inference
from peft import AutoPeftModelForCausalLM
# Load adapter directly
model = AutoPeftModelForCausalLM.from_pretrained("my-lora-adapter")
model.eval()
# Or merge into base model for zero-overhead inference
merged_model = model.merge_and_unload()
When to Use PEFT vs Full Fine-Tuning
Choose PEFT when:
- Hardware is limited. You've got one or a few GPUs, not a multi-node cluster.
- You need multiple task-specific models. Swap lightweight adapters instead of storing separate full models.
- The task is close to the pretraining domain. Classification, summarization, Q&A, or instruction following over familiar content.
- You need fast iteration. Training runs finish in hours instead of days.
Choose full fine-tuning when:
- Maximum accuracy is non-negotiable. Some highly specialized domains (rare languages, novel scientific notation) still benefit from updating all parameters.
- You have the compute budget. And the task justifies it.
- The domain is far from pretraining data. Extremely niche tasks where the pretrained representations need deep restructuring.
For most production use cases in 2026, PEFT is the right starting point. If your PEFT-adapted model falls short on benchmarks, you can always scale up rank or switch to full fine-tuning.
Choosing the Right PEFT Method
| Factor | Recommended Method | |---|---| | General-purpose starting point | LoRA (r=8-16) | | Severely constrained GPU memory | QLoRA (4-bit base + LoRA) | | Multiple tasks on one base model | LoRA with adapter swapping | | Minimal parameter count needed | IA3 | | Preserving model geometry matters | OFT / BOFT | | Need full-rank expressiveness | HiRA | | Non-transformer / diffusion models | LoRA (supported via Diffusers) |
Key LoraConfig parameters to tune: r (rank: 8 for simple tasks, 32+ for complex ones), lora_alpha (set to roughly 2x rank), target_modules (start with attention projections, expand to "all-linear" for QLoRA-style coverage), and lora_dropout (0.05-0.1).
From Fine-Tuned Model to Production Agent
Most PEFT guides stop at "save your adapter." Real production is different. A fine-tuned model needs to connect to actual tools and run on its own. PEFT-tuned models trained on tool-calling schemas can show strong improvements in structured output and function-call accuracy, even with modest training data.
Multi-adapter serving is becoming standard. Teams load multiple LoRA adapters on a single base model and route requests to the right adapter per task. Databricks reports up to 1.5x throughput improvement for LoRA workloads in their custom inference runtime.
For teams that need their fine-tuned models running as persistent agents (invoking tools, managing workflows, operating around the clock), platforms like Gamut bridge the gap. Gamut deploys AI agents that connect to any tool via MCP servers, letting a PEFT-tuned model go from adapter checkpoint to always-on production agent without building custom serving infrastructure.
Security Considerations
PEFT adapters are small and easy to distribute. That's a feature, but it's also a risk. Research from Princeton, Stanford, Virginia Tech, and IBM Research (published at ICLR 2024) showed that safety alignment can be stripped from models with as few as 10 adversarial fine-tuning examples costing under $0.20. The UK AI Security Institute confirmed that current defenses against malicious fine-tuning remain insufficient.
On a practical level: always use the safetensors checkpoint format (the PEFT library default) when loading adapters from untrusted sources. It prevents arbitrary code execution during deserialization. Researchers have built PADBench, a benchmark of 13,300 benign and backdoored adapters, to study backdoor detection, but few reliable detection tools exist yet. Audit the provenance of any community adapter before deploying it.
FAQ
What is the difference between PEFT and LoRA?
PEFT is the umbrella term for all parameter-efficient fine-tuning techniques. LoRA (Low-Rank Adaptation) is one specific method under that umbrella, and by far the most popular one. It accounts for over 98% of PEFT-related model cards on Hugging Face Hub. Other PEFT methods include QLoRA, prefix tuning, prompt tuning, IA3, and more.
How does PEFT compare to full fine-tuning in performance?
PEFT typically gets close to full fine-tuning performance while using substantially less compute. The gap narrows as model size increases. LoRA's inventor Edward Hu noted that low-rank updates "perform almost as well as full fine tuning in large models like GPT-3."
What hardware do I need to fine-tune with PEFT?
With QLoRA, you can fine-tune a 7B model on a consumer GPU like an RTX 4090 (24GB VRAM). A 13B model fits on a single A100 40GB. Standard LoRA (without quantization) needs roughly 2-3x more memory than QLoRA. Full fine-tuning of a 7B model typically requires 4-8 A100 80GB GPUs.
Can PEFT be used with diffusion models?
Yes. The PEFT library integrates with Hugging Face Diffusers. LoRA training of Stable Diffusion produces a compact adapter checkpoint (typically a few MB), making it practical to train custom styles and concepts on consumer hardware.
Which PEFT method should I start with?
Start with LoRA. It's the recommended starting point in the official docs, has the broadest tooling support, and works well across model architectures. If GPU memory is tight, add 4-bit quantization (QLoRA). Only explore other peft methods like IA3 or prefix tuning if you have a specific constraint that LoRA doesn't address.
Deploy Your Fine-Tuned Models as Production Agents
Gamut connects your PEFT-tuned models to any tool via MCP servers and runs them as persistent, always-on AI agents -- no custom serving infrastructure required.