Worker Setup
Connecting your own GPU machines to the Parsyn platform. Installation, configuration, fleet management, and troubleshooting.
Why connect your own workers
Parsyn Workers (managed GPUs) are the fastest way to start training, but connecting your own machines gives you:
- No credit cost: You pay for hardware and electricity, not per-hour credits.
- Inference capability: Parsyn Workers handle training only. Chatting with your models requires your own workers.
- Data control: Your GPU machines can be on your own network. Model weights pass through S3 storage, but compute stays on your hardware.
- Always available: No fleet capacity limits. Your machines are yours.
Requirements
- An NVIDIA GPU with its drivers installed
- 8 GB+ VRAM (16 GB+ recommended for 7B models with LoRA)
- Linux, or Windows with WSL2
- Internet access to
parsyn.progatis.com(WebSocket + HTTPS)
No Python and no CUDA toolkit to install: the worker runs from a container image that carries its own runtime. Drivers are the one thing the installer will not touch — check yours answer:
nvidia-smiInstallation
One command on the machine with the GPU:
curl -fsSL https://get.parsyn.progatis.com/worker | sh It installs Docker and the NVIDIA container toolkit if they are missing, then the parsyn command, then pulls the worker image. Then:
parsyn init # asks two questions
parsyn start # connects and enrolls
parsyn logs -f # follow it
parsyn doctor # diagnose, if something looks wrongparsyn init asks for your enrollment key at a prompt rather than taking it as an argument, so it never reaches your shell history. Create the key from Workers → Add a machine in the dashboard, which shows these same three commands.
Unattended installs
For cloud-init, Ansible or an image build, parsyn init takes the same configuration without prompting. The key comes in on standard input for the same reason it is not a flag — a command line is visible in ps while it runs:
curl -fsSL https://get.parsyn.progatis.com/worker | sh
printf '%s' "$PARSYN_KEY" | parsyn init --key-stdin --non-interactive --name gpu-box-1
parsyn startparsyn init --help lists every option, including --url for a self-hosted platform and --force to replace an existing configuration.
Registering a worker
Before a worker can connect, it needs to be registered on the platform. Two options:
Enrollment key
Create an enrollment key from the Enrollment Keys page in the dashboard:
curl -X POST https://parsyn.progatis.com/api/v1/master-keys \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "datacenter-east-fleet",
"description": "Auto-enrollment for east datacenter",
"auto_name_prefix": "east-worker",
"default_worker_type": "both",
"max_enrollments": 50
}' The enrollment key (format: enroll_...) can be shared with any number of machines. On first connection, each worker registers itself automatically and receives its own individual key (format: wkr_...). The worker writes it back into its .env file under WORKER_API_KEY and reuses it on subsequent starts, skipping the enrollment handshake.
Bake the enrollment key into your provisioning script or machine image, and new machines register themselves.
Configuration
Workers are configured through environment variables:
# Connection (required)
PLATFORM_URL=wss://parsyn.progatis.com/ws/worker
WORKER_ENROLLMENT_KEY=enroll_xyz789... # at first install
WORKER_API_KEY=wkr_abc123... # written automatically after enrollment
# GPU
GPU_DEVICE=cuda:0
MAX_BATCH_SIZE=32
# Worker behavior
WORKER_TYPE=both # fine_tuner, prompter, or both
HEARTBEAT_INTERVAL=30 # Seconds between heartbeats
RECONNECT_DELAY=5 # Initial reconnect delay
# Performance
QUANTIZATION=none # none, 4bit, or 8bit for inference
TRUST_REMOTE_CODE=false # Allow HF models with custom code
# Storage
CACHE_DIR=~/.cache/parsyn
LOG_LEVEL=INFOFull configuration reference
| Variable | Required | Default | Description |
|---|---|---|---|
PLATFORM_URL | Yes | WebSocket URL. Use wss://parsyn.progatis.com/ws/worker for the hosted platform. | |
WORKER_API_KEY | Yes* | Individual API key from manual registration. | |
WORKER_ENROLLMENT_KEY | Yes* | Master key for auto-enrollment on first connection. | |
GPU_DEVICE | No | cuda:0 | CUDA device. cuda:1 for second GPU, etc. |
MAX_BATCH_SIZE | No | 32 | Upper limit on batch size the worker will accept. |
WORKER_TYPE | No | both | fine_tuner (training only), prompter (inference only), or both. |
QUANTIZATION | No | none | Load models quantized for inference. Reduces VRAM at some quality cost. |
TRUST_REMOTE_CODE | No | false | Allow models with custom Python code from HuggingFace. Security risk, enable only for trusted models. |
CACHE_DIR | No | ~/.cache/parsyn | Local cache for downloaded models and datasets. |
TRAINING_TIMEOUT_SECONDS | No | 86400 | Maximum training time per job (24 hours). |
INFERENCE_TIMEOUT_SECONDS | No | 300 | Maximum inference time per request (5 minutes). |
MAX_INFERENCE_CACHED_MODELS | No | 2 | Number of models kept in GPU memory for fast inference switching. |
* One of WORKER_API_KEY or WORKER_ENROLLMENT_KEY is required.
Starting the worker
parsyn-worker startOn startup, the worker:
- Resolves its API key (individual key, cached credentials from previous enrollment, or enrollment key)
- Connects to the platform via WebSocket
- Authenticates and (if enrolling) receives its individual API key
- Runs a GPU benchmark (cached for 24 hours)
- Sends hardware info and benchmark results to the platform
- Starts the heartbeat loop (every 30s) and metrics loop (every 5s)
- Waits for tasks
Expected output:
INFO Worker abc12345 connecting to wss://parsyn.progatis.com/ws/worker
INFO Connected to platform
INFO Running GPU benchmark...
INFO Benchmark complete: score=85.2, TFLOPS=82.6, bandwidth=1008 GB/s
INFO Registered: NVIDIA A100-SXM4-80GB, 80 GB VRAM, CUDA 12.1
INFO Worker ready, waiting for tasks...GPU benchmark
On first connection (and every 24 hours), the worker runs an automated GPU benchmark. The platform uses benchmark scores for intelligent job assignment in auto mode, matching jobs to the most capable available worker.
The benchmark measures:
- TFLOPS: FP16 matrix multiplication throughput
- Memory bandwidth: Large tensor transfer speed (GB/s)
- Max batch size: Binary search for largest batch that fits in memory
- Composite score: Weighted combination (50% TFLOPS, 30% bandwidth, 20% VRAM), normalized against an RTX 4090 baseline
If the benchmark times out (60 seconds), the worker uses conservative fallback values and reports to the platform anyway.
Running as a system service
For production, use systemd so the worker starts on boot and restarts on failure:
[Unit]
Description=Parsyn Training Worker
After=network.target
[Service]
Type=simple
User=parsyn
WorkingDirectory=/opt/parsyn-worker
EnvironmentFile=/opt/parsyn-worker/.env
ExecStart=/opt/parsyn-worker/venv/bin/parsyn-worker start
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable parsyn-worker
sudo systemctl start parsyn-worker
sudo journalctl -u parsyn-worker -fMultiple GPUs on one machine
Each worker process uses one GPU. Run multiple instances with different GPU_DEVICE values:
# Terminal 1
GPU_DEVICE=cuda:0 parsyn-worker start
# Terminal 2
GPU_DEVICE=cuda:1 parsyn-worker startWith systemd, use a template unit:
# /etc/systemd/system/[email protected]
[Service]
EnvironmentFile=/opt/parsyn-worker/.env.%i
ExecStart=/opt/parsyn-worker/venv/bin/parsyn-worker start
# Enable per GPU
sudo systemctl enable parsyn-worker@gpu0 parsyn-worker@gpu1If using master key enrollment, each GPU instance enrolls as a separate worker and receives its own API key.
Fleet management
Master key lifecycle
Master keys support limits and expiration:
max_enrollments: cap the number of workers that can enroll (null for unlimited)expires_at: expiration date after which new enrollments are rejectedauto_name_prefix: workers are named sequentially (e.g., "east-worker-1", "east-worker-2")
Revoking a master key
DELETE /api/v1/master-keys/{id}Revocation prevents new enrollments but does not disconnect existing workers. Those workers already have their own individual API keys.
Monitoring your fleet
The Workers page shows summary cards at the top (total, online, training, idle, average GPU utilization) and a live table of all workers. A green pulsing dot indicates the real-time WebSocket connection is active.
Each row in the table shows:
- Connection status (green dot for idle, red pulsing for training, gray for offline)
- GPU utilization percentage with a progress bar, plus VRAM used
- CPU and RAM usage bars
- Current job ID with a progress bar and percentage (if training)
- Worker type badge (Both, Training only, Inference only)
- ETA for the current job
Click a row to open a detail modal. Use the action icons to view details or delete the worker.
Individual key persistence
When a worker enrolls via an enrollment key, it receives an individual key (wkr_...) that the worker process writes directly into its .env file under WORKER_API_KEY. On subsequent starts, the worker uses that key and skips the enrollment handshake.
If you need to re-enroll a worker (e.g., after losing the .env), start it again with the enrollment key.
Troubleshooting
Worker can't connect
- Verify network access:
curl https://parsyn.progatis.com/health - Check that the enrollment key is still active in the dashboard.
- If behind a corporate firewall, ensure WebSocket (wss://) traffic on port 443 is allowed.
Worker connects but disconnects immediately
- Check the worker logs for authentication errors.
- Verify the key prefix:
wkr_for individual keys,enroll_for enrollment keys. - Check that the enrollment key hasn't been revoked or reached its enrollment limit.
GPU not detected
- Run
nvidia-smito confirm the GPU is visible. - Check PyTorch:
python -c "import torch; print(torch.cuda.is_available())" - If using Docker, use the NVIDIA runtime:
docker run --gpus all ...
Training crashes with OOM
MAX_BATCH_SIZEcaps the batch size the worker accepts. Lower it if needed.- Enable quantization:
QUANTIZATION=4bit - Make sure no other processes are using the GPU.
Model download fails
- Check internet access to HuggingFace:
curl https://huggingface.co - For gated models (Llama, Gemma, etc.), set
HF_TOKENwith a token that has access. - Verify
CACHE_DIRhas enough disk space. 7B models need 15+ GB.