Distributed Training
Running training jobs across multiple workers with federated learning. How rounds, weight aggregation, and multi-GPU coordination work.
Overview
Parsyn supports distributed training through federated learning. Each worker trains on a subset of the data, and the platform periodically aggregates the results into a global model. This approach works across machines on different networks without requiring high-bandwidth interconnects.
You can mix worker types: use your own GPUs for some of the computation and Parsyn Workers for the rest. The platform handles data distribution, round coordination, and weight aggregation.
How it works
Round-based training
Federated training runs in rounds:
- Distribution: The platform sends the current global model weights to all participating workers.
- Local training: Each worker trains on its shard of the data for a configured number of steps.
- Upload: Each worker uploads its updated weights to S3.
- Aggregation: The platform aggregates the weights using federated averaging (FedAvg), weighted by each worker's sample count.
- Next round: The aggregated weights become the new global model. Repeat.
Round 1:
Platform ──[global weights]──> Worker A (your GPU)
Platform ──[global weights]──> Worker B (Parsyn A100)
Platform ──[global weights]──> Worker C (your GPU)
Workers train locally, upload weights
Platform aggregates ──> new global weights
Round 2:
Platform ──[new weights]──> Worker A, B, C
...Configuration
Enable distributed training when creating a job. In the dashboard's training job form, expand the advanced section to access distributed settings. You can also configure it via the API:
POST /api/v1/training-jobs
{
"name": "distributed-llama-finetune",
"dataset_id": 42,
"model_id": 7,
"config": {
"epochs": 3,
"batch_size": 4,
"learning_rate": 2e-4,
"mixed_precision": "bf16",
"use_peft": true,
"lora_r": 16,
"lora_alpha": 32
},
"distributed": {
"enabled": true,
"strategy": "federated_averaging",
"num_rounds": 10,
"steps_per_round": 500,
"min_workers": 2,
"max_workers": 5,
"aggregation_timeout": 600,
"fault_tolerance": 1
},
"worker_selection": {
"mode": "mixed",
"worker_ids": [1, 3],
"parsyn_worker_ids": [5]
}
}Distributed parameters
| Parameter | Default | Description |
|---|---|---|
strategy | federated_averaging | Aggregation strategy. FedAvg is currently the supported method. |
num_rounds | 10 | Total aggregation rounds. |
steps_per_round | 500 | Training steps each worker performs per round. |
min_workers | 2 | Minimum workers required to start a round. |
max_workers | 5 | Maximum workers to include. |
aggregation_timeout | 600 | Seconds to wait for all workers before aggregating with available results. |
fault_tolerance | 1 | Number of consecutive failed rounds before aborting the job. |
Data distribution
The platform splits the dataset across participating workers. The split is roughly equal and random (seeded for reproducibility). Each worker downloads only its shard.
The validation set is not split. Every worker receives the full validation set so evaluation metrics are comparable across workers.
Tracking progress
Open the training job from the Training page to see distributed training status in the detail view:
- Current round and total rounds
- Per-worker progress within the current round
- Which workers have reported and which are still training
- Per-worker loss curves
- Aggregated loss after each round
- Worker GPU utilization and throughput
Failure handling
Worker disconnects mid-round
The platform waits for the remaining workers. As long as min_workers complete the round, aggregation proceeds. The disconnected worker can rejoin in a later round with the latest global weights.
Not enough workers
If fewer than min_workers complete a round within the aggregation_timeout, the round fails. After fault_tolerance consecutive failures, the entire job is marked as failed.
Worker rejoins
When a worker reconnects, the platform sends it the latest aggregated weights. It picks up from the current round, not from where it left off.
When to use distributed training
Good fit
- Multiple GPU machines in different locations or cloud accounts.
- Large dataset that takes too long on a single GPU.
- Data locality requirements (workers keep data on-premise, only share model weights).
- Redundancy: if one machine goes down, the others continue.
Not ideal
- Small datasets (<1000 examples). Single GPU will finish faster than the coordination overhead.
- Single machine with multiple GPUs. Use separate worker instances per GPU instead of federated learning.
- Maximum training quality is critical. FedAvg introduces some convergence gap compared to centralized training.
Practical tips
- Use LoRA for distributed training. Adapter weights are small (tens of MB), making aggregation fast. Full model weight aggregation moves gigabytes per round.
- More rounds with fewer steps per round gives better convergence but more communication overhead. Start with 10 rounds / 500 steps and adjust.
- Monitor validation loss per round. If it plateaus after 5-6 rounds, you can stop early.
- Keep data balanced. If one worker has 90% of the data, the weighted average will be dominated by that worker.