P
Parsyn
/Docs
Back to Home

Training Configuration

Hyperparameters, LoRA adapters, mixed precision, and advanced options. How to configure a training job for your use case.

Training config structure

When you create a training job from the dashboard, the form exposes all these parameters through input fields, dropdowns, and checkboxes. The basic fields (learning rate, batch size, epochs, max length) are visible by default. Advanced options (gradient accumulation, LoRA, evaluation, checkpointing) are in a collapsible section.

Under the hood, these settings produce a config object like this:

{
  "epochs": 3,
  "batch_size": 4,
  "learning_rate": 5e-5,
  "weight_decay": 0.01,
  "warmup_ratio": 0.1,
  "max_length": 512,
  "gradient_accumulation_steps": 4,
  "mixed_precision": "bf16",
  "gradient_checkpointing": true,
  "lr_scheduler": "cosine",
  "use_peft": true,
  "peft_method": "lora",
  "lora_r": 16,
  "lora_alpha": 32,
  "lora_dropout": 0.05,
  "save_strategy": "steps",
  "save_steps": 500,
  "eval_strategy": "steps",
  "eval_steps": 250,
  "logging_steps": 10,
  "seed": 42
}

Core parameters

ParameterTypeDefaultDescription
epochsint3Number of full passes over the training data.
batch_sizeint4Examples per forward/backward pass. Limited by GPU memory.
learning_ratefloat5e-5Peak learning rate. Typical range for fine-tuning: 1e-5 to 5e-4.
weight_decayfloat0.01L2 regularization. Helps prevent overfitting.
warmup_ratiofloat0.0Fraction of total steps for linear warmup (e.g., 0.1 = 10% warmup).
max_lengthint512Maximum token sequence length. Longer sequences are truncated.
seedint42Random seed for reproducibility.

Batch size and gradient accumulation

Your effective batch size is batch_size * gradient_accumulation_steps. If your GPU can't fit a large batch, reduce batch_size and increase gradient_accumulation_steps:

// Effective batch size = 4 * 8 = 32
{
  "batch_size": 4,
  "gradient_accumulation_steps": 8
}

Gradient accumulation uses less memory per step but is slower overall (more forward passes before one weight update). This is the standard approach when VRAM is tight.

Mixed precision

The mixed_precision field controls floating-point precision during training:

ValueWhen to use
"bf16"Ampere GPUs and newer (A100, RTX 3090/4090). Best stability. Recommended default.
"fp16"Older GPUs with Tensor Cores (V100, RTX 2000 series). May need loss scaling.
"fp32"Full precision. More memory, slower, but maximally stable. Fallback if fp16/bf16 causes NaN.
"auto"Let the worker decide based on its GPU capabilities.

When using Parsyn Workers, "auto" is a safe choice. The worker will pick bf16 on A100s and fp16 on older hardware.

LoRA (Low-Rank Adaptation)

LoRA adds small trainable matrices to specific layers instead of updating all model weights. This cuts memory usage dramatically and produces a compact adapter file (10-100 MB) rather than a full model copy. In the training job form, enable it with the PEFT checkbox in the advanced section, then choose between LoRA and QLoRA.

When to use LoRA

  • Limited GPU memory (less than 24 GB VRAM for 7B+ models)
  • Fast training iterations
  • Multiple fine-tuned versions without storing full copies
  • Domain adaptation or style transfer on an already-capable model

LoRA parameters

ParameterDefaultDescription
use_peftfalseEnable PEFT (Parameter-Efficient Fine-Tuning).
peft_method"lora"lora or qlora (quantized LoRA for even less memory).
lora_r8Rank of low-rank matrices. Higher = more capacity. Common: 4, 8, 16, 32, 64.
lora_alpha16Scaling factor. Typically set to 2 * lora_r.
lora_dropout0.05Dropout on LoRA layers. Set to 0 for very small datasets.
quantization"none"For QLoRA: "4bit" or "8bit". Loads base model quantized.

LoRA vs full fine-tuning

LoRAFull fine-tuning
Memory (7B model)~6 GB VRAM~28 GB VRAM (fp16)
SpeedFastSlower
Output size~50 MB adapter~14 GB full model
Quality ceilingSlightly lowerMaximum
Best forDomain adaptation, style, instructionsSignificant capability changes, new knowledge

Learning rate schedulers

The lr_scheduler field controls how the learning rate changes during training:

  • "linear": Linear decay from peak to zero.
  • "cosine": Cosine decay (smooth, widely used). Good default.
  • "polynomial": Polynomial decay with configurable power.
  • "constant": Fixed learning rate throughout training.
  • "constant_with_warmup": Warmup phase, then constant.

All schedulers support a warmup phase controlled by warmup_ratio. During warmup, the learning rate increases linearly from 0 to the configured value.

Evaluation during training

If you have a validation dataset, configure in-training evaluation in the Evaluation section of the advanced options. Set the evaluation strategy to "steps" or "epoch", select a validation dataset, and optionally enable early stopping.

The equivalent config:

{
  "eval_strategy": "steps",
  "eval_steps": 250,
  "validation_dataset_id": 43,
  "load_best_model_at_end": true,
  "metric_for_best_model": "eval_loss",
  "early_stopping_patience": 3,
  "early_stopping_threshold": 0.001
}

The worker pauses training at each evaluation checkpoint, runs the validation set, and reports validation loss. With early stopping, training halts automatically when validation loss stops improving, preventing overfitting.

Checkpointing

Control checkpoint saving with save_strategy:

  • "steps": Save every save_steps steps.
  • "epoch": Save at the end of each epoch.
  • "no": No intermediate checkpoints (only save the final model).

Each checkpoint is uploaded to S3 and includes model weights, optimizer state, and metrics at that step. If training is interrupted, you can resume from the last checkpoint.

Recipes

Instruction tuning (LoRA, 7B model, RTX 4090 or equivalent)

{
  "epochs": 3,
  "batch_size": 2,
  "gradient_accumulation_steps": 8,
  "learning_rate": 2e-4,
  "warmup_ratio": 0.1,
  "max_length": 1024,
  "mixed_precision": "bf16",
  "gradient_checkpointing": true,
  "lr_scheduler": "cosine",
  "use_peft": true,
  "peft_method": "lora",
  "lora_r": 16,
  "lora_alpha": 32,
  "lora_dropout": 0.05,
  "save_strategy": "steps",
  "save_steps": 200,
  "eval_strategy": "steps",
  "eval_steps": 100,
  "logging_steps": 5
}

Full fine-tuning (small model, A100 80GB)

{
  "epochs": 5,
  "batch_size": 16,
  "gradient_accumulation_steps": 2,
  "learning_rate": 5e-5,
  "weight_decay": 0.01,
  "warmup_ratio": 0.05,
  "max_length": 2048,
  "mixed_precision": "bf16",
  "lr_scheduler": "cosine",
  "use_peft": false,
  "save_strategy": "steps",
  "save_steps": 1000,
  "eval_strategy": "steps",
  "eval_steps": 500,
  "logging_steps": 10
}

Quick experiment (GPT-2, fast iteration)

{
  "epochs": 1,
  "batch_size": 8,
  "learning_rate": 5e-5,
  "max_length": 256,
  "mixed_precision": "auto",
  "use_peft": false,
  "save_strategy": "no",
  "logging_steps": 1
}

Troubleshooting

Loss is NaN

  • Lower the learning rate (e.g., from 5e-5 to 1e-5).
  • Switch from fp16 to bf16, or use fp32.
  • Check dataset for empty or malformed examples.
  • Enable gradient checkpointing to stabilize memory.

Out of memory (OOM)

  • Reduce batch_size, compensate with gradient_accumulation_steps.
  • Reduce max_length.
  • Enable LoRA or QLoRA.
  • Enable gradient checkpointing.
  • Use mixed precision if you're in fp32.

Training is slow

  • Enable mixed precision.
  • Increase batch size if VRAM allows.
  • Reduce max_length if your data is shorter.
  • Check logging_steps: logging every step adds overhead.

Model quality is poor

  • Train for more epochs (watch validation loss for overfitting).
  • Increase LoRA rank and alpha.
  • Review dataset quality (consistency, variety, accuracy).
  • Try different learning rates. The optimum varies by model and dataset.