Training Configuration
Hyperparameters, LoRA adapters, mixed precision, and advanced options. How to configure a training job for your use case.
Training config structure
When you create a training job from the dashboard, the form exposes all these parameters through input fields, dropdowns, and checkboxes. The basic fields (learning rate, batch size, epochs, max length) are visible by default. Advanced options (gradient accumulation, LoRA, evaluation, checkpointing) are in a collapsible section.
Under the hood, these settings produce a config object like this:
{
"epochs": 3,
"batch_size": 4,
"learning_rate": 5e-5,
"weight_decay": 0.01,
"warmup_ratio": 0.1,
"max_length": 512,
"gradient_accumulation_steps": 4,
"mixed_precision": "bf16",
"gradient_checkpointing": true,
"lr_scheduler": "cosine",
"use_peft": true,
"peft_method": "lora",
"lora_r": 16,
"lora_alpha": 32,
"lora_dropout": 0.05,
"save_strategy": "steps",
"save_steps": 500,
"eval_strategy": "steps",
"eval_steps": 250,
"logging_steps": 10,
"seed": 42
}Core parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
epochs | int | 3 | Number of full passes over the training data. |
batch_size | int | 4 | Examples per forward/backward pass. Limited by GPU memory. |
learning_rate | float | 5e-5 | Peak learning rate. Typical range for fine-tuning: 1e-5 to 5e-4. |
weight_decay | float | 0.01 | L2 regularization. Helps prevent overfitting. |
warmup_ratio | float | 0.0 | Fraction of total steps for linear warmup (e.g., 0.1 = 10% warmup). |
max_length | int | 512 | Maximum token sequence length. Longer sequences are truncated. |
seed | int | 42 | Random seed for reproducibility. |
Batch size and gradient accumulation
Your effective batch size is batch_size * gradient_accumulation_steps. If your GPU can't fit a large batch, reduce batch_size and increase gradient_accumulation_steps:
// Effective batch size = 4 * 8 = 32
{
"batch_size": 4,
"gradient_accumulation_steps": 8
}Gradient accumulation uses less memory per step but is slower overall (more forward passes before one weight update). This is the standard approach when VRAM is tight.
Mixed precision
The mixed_precision field controls floating-point precision during training:
| Value | When to use |
|---|---|
"bf16" | Ampere GPUs and newer (A100, RTX 3090/4090). Best stability. Recommended default. |
"fp16" | Older GPUs with Tensor Cores (V100, RTX 2000 series). May need loss scaling. |
"fp32" | Full precision. More memory, slower, but maximally stable. Fallback if fp16/bf16 causes NaN. |
"auto" | Let the worker decide based on its GPU capabilities. |
When using Parsyn Workers, "auto" is a safe choice. The worker will pick bf16 on A100s and fp16 on older hardware.
LoRA (Low-Rank Adaptation)
LoRA adds small trainable matrices to specific layers instead of updating all model weights. This cuts memory usage dramatically and produces a compact adapter file (10-100 MB) rather than a full model copy. In the training job form, enable it with the PEFT checkbox in the advanced section, then choose between LoRA and QLoRA.
When to use LoRA
- Limited GPU memory (less than 24 GB VRAM for 7B+ models)
- Fast training iterations
- Multiple fine-tuned versions without storing full copies
- Domain adaptation or style transfer on an already-capable model
LoRA parameters
| Parameter | Default | Description |
|---|---|---|
use_peft | false | Enable PEFT (Parameter-Efficient Fine-Tuning). |
peft_method | "lora" | lora or qlora (quantized LoRA for even less memory). |
lora_r | 8 | Rank of low-rank matrices. Higher = more capacity. Common: 4, 8, 16, 32, 64. |
lora_alpha | 16 | Scaling factor. Typically set to 2 * lora_r. |
lora_dropout | 0.05 | Dropout on LoRA layers. Set to 0 for very small datasets. |
quantization | "none" | For QLoRA: "4bit" or "8bit". Loads base model quantized. |
LoRA vs full fine-tuning
| LoRA | Full fine-tuning | |
|---|---|---|
| Memory (7B model) | ~6 GB VRAM | ~28 GB VRAM (fp16) |
| Speed | Fast | Slower |
| Output size | ~50 MB adapter | ~14 GB full model |
| Quality ceiling | Slightly lower | Maximum |
| Best for | Domain adaptation, style, instructions | Significant capability changes, new knowledge |
Learning rate schedulers
The lr_scheduler field controls how the learning rate changes during training:
"linear": Linear decay from peak to zero."cosine": Cosine decay (smooth, widely used). Good default."polynomial": Polynomial decay with configurable power."constant": Fixed learning rate throughout training."constant_with_warmup": Warmup phase, then constant.
All schedulers support a warmup phase controlled by warmup_ratio. During warmup, the learning rate increases linearly from 0 to the configured value.
Evaluation during training
If you have a validation dataset, configure in-training evaluation in the Evaluation section of the advanced options. Set the evaluation strategy to "steps" or "epoch", select a validation dataset, and optionally enable early stopping.
The equivalent config:
{
"eval_strategy": "steps",
"eval_steps": 250,
"validation_dataset_id": 43,
"load_best_model_at_end": true,
"metric_for_best_model": "eval_loss",
"early_stopping_patience": 3,
"early_stopping_threshold": 0.001
}The worker pauses training at each evaluation checkpoint, runs the validation set, and reports validation loss. With early stopping, training halts automatically when validation loss stops improving, preventing overfitting.
Checkpointing
Control checkpoint saving with save_strategy:
"steps": Save everysave_stepssteps."epoch": Save at the end of each epoch."no": No intermediate checkpoints (only save the final model).
Each checkpoint is uploaded to S3 and includes model weights, optimizer state, and metrics at that step. If training is interrupted, you can resume from the last checkpoint.
Recipes
Instruction tuning (LoRA, 7B model, RTX 4090 or equivalent)
{
"epochs": 3,
"batch_size": 2,
"gradient_accumulation_steps": 8,
"learning_rate": 2e-4,
"warmup_ratio": 0.1,
"max_length": 1024,
"mixed_precision": "bf16",
"gradient_checkpointing": true,
"lr_scheduler": "cosine",
"use_peft": true,
"peft_method": "lora",
"lora_r": 16,
"lora_alpha": 32,
"lora_dropout": 0.05,
"save_strategy": "steps",
"save_steps": 200,
"eval_strategy": "steps",
"eval_steps": 100,
"logging_steps": 5
}Full fine-tuning (small model, A100 80GB)
{
"epochs": 5,
"batch_size": 16,
"gradient_accumulation_steps": 2,
"learning_rate": 5e-5,
"weight_decay": 0.01,
"warmup_ratio": 0.05,
"max_length": 2048,
"mixed_precision": "bf16",
"lr_scheduler": "cosine",
"use_peft": false,
"save_strategy": "steps",
"save_steps": 1000,
"eval_strategy": "steps",
"eval_steps": 500,
"logging_steps": 10
}Quick experiment (GPT-2, fast iteration)
{
"epochs": 1,
"batch_size": 8,
"learning_rate": 5e-5,
"max_length": 256,
"mixed_precision": "auto",
"use_peft": false,
"save_strategy": "no",
"logging_steps": 1
}Troubleshooting
Loss is NaN
- Lower the learning rate (e.g., from 5e-5 to 1e-5).
- Switch from fp16 to bf16, or use fp32.
- Check dataset for empty or malformed examples.
- Enable gradient checkpointing to stabilize memory.
Out of memory (OOM)
- Reduce
batch_size, compensate withgradient_accumulation_steps. - Reduce
max_length. - Enable LoRA or QLoRA.
- Enable gradient checkpointing.
- Use mixed precision if you're in fp32.
Training is slow
- Enable mixed precision.
- Increase batch size if VRAM allows.
- Reduce
max_lengthif your data is shorter. - Check
logging_steps: logging every step adds overhead.
Model quality is poor
- Train for more epochs (watch validation loss for overfitting).
- Increase LoRA rank and alpha.
- Review dataset quality (consistency, variety, accuracy).
- Try different learning rates. The optimum varies by model and dataset.