P
Parsyn
/Docs
Back to Home

Worker Setup

Connecting your own GPU machines to the Parsyn platform. Installation, configuration, fleet management, and troubleshooting.

Why connect your own workers

Parsyn Workers (managed GPUs) are the fastest way to start training, but connecting your own machines gives you:

  • No credit cost: You pay for hardware and electricity, not per-hour credits.
  • Inference capability: Parsyn Workers handle training only. Chatting with your models requires your own workers.
  • Data control: Your GPU machines can be on your own network. Model weights pass through S3 storage, but compute stays on your hardware.
  • Always available: No fleet capacity limits. Your machines are yours.

Requirements

  • An NVIDIA GPU with its drivers installed
  • 8 GB+ VRAM (16 GB+ recommended for 7B models with LoRA)
  • Linux, or Windows with WSL2
  • Internet access to parsyn.progatis.com (WebSocket + HTTPS)

No Python and no CUDA toolkit to install: the worker runs from a container image that carries its own runtime. Drivers are the one thing the installer will not touch — check yours answer:

nvidia-smi

Installation

One command on the machine with the GPU:

curl -fsSL https://get.parsyn.progatis.com/worker | sh

It installs Docker and the NVIDIA container toolkit if they are missing, then the parsyn command, then pulls the worker image. Then:

parsyn init     # asks two questions
parsyn start    # connects and enrolls
parsyn logs -f  # follow it
parsyn doctor   # diagnose, if something looks wrong

parsyn init asks for your enrollment key at a prompt rather than taking it as an argument, so it never reaches your shell history. Create the key from Workers → Add a machine in the dashboard, which shows these same three commands.

Unattended installs

For cloud-init, Ansible or an image build, parsyn init takes the same configuration without prompting. The key comes in on standard input for the same reason it is not a flag — a command line is visible in ps while it runs:

curl -fsSL https://get.parsyn.progatis.com/worker | sh
printf '%s' "$PARSYN_KEY" | parsyn init --key-stdin --non-interactive   --name gpu-box-1
parsyn start

parsyn init --help lists every option, including --url for a self-hosted platform and --force to replace an existing configuration.

Registering a worker

Before a worker can connect, it needs to be registered on the platform. Two options:

Enrollment key

Create an enrollment key from the Enrollment Keys page in the dashboard:

curl -X POST https://parsyn.progatis.com/api/v1/master-keys \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "datacenter-east-fleet",
    "description": "Auto-enrollment for east datacenter",
    "auto_name_prefix": "east-worker",
    "default_worker_type": "both",
    "max_enrollments": 50
  }'

The enrollment key (format: enroll_...) can be shared with any number of machines. On first connection, each worker registers itself automatically and receives its own individual key (format: wkr_...). The worker writes it back into its .env file under WORKER_API_KEY and reuses it on subsequent starts, skipping the enrollment handshake.

Bake the enrollment key into your provisioning script or machine image, and new machines register themselves.

Configuration

Workers are configured through environment variables:

# Connection (required)
PLATFORM_URL=wss://parsyn.progatis.com/ws/worker
WORKER_ENROLLMENT_KEY=enroll_xyz789...    # at first install
WORKER_API_KEY=wkr_abc123...              # written automatically after enrollment

# GPU
GPU_DEVICE=cuda:0
MAX_BATCH_SIZE=32

# Worker behavior
WORKER_TYPE=both          # fine_tuner, prompter, or both
HEARTBEAT_INTERVAL=30     # Seconds between heartbeats
RECONNECT_DELAY=5         # Initial reconnect delay

# Performance
QUANTIZATION=none         # none, 4bit, or 8bit for inference
TRUST_REMOTE_CODE=false   # Allow HF models with custom code

# Storage
CACHE_DIR=~/.cache/parsyn
LOG_LEVEL=INFO

Full configuration reference

VariableRequiredDefaultDescription
PLATFORM_URLYesWebSocket URL. Use wss://parsyn.progatis.com/ws/worker for the hosted platform.
WORKER_API_KEYYes*Individual API key from manual registration.
WORKER_ENROLLMENT_KEYYes*Master key for auto-enrollment on first connection.
GPU_DEVICENocuda:0CUDA device. cuda:1 for second GPU, etc.
MAX_BATCH_SIZENo32Upper limit on batch size the worker will accept.
WORKER_TYPENobothfine_tuner (training only), prompter (inference only), or both.
QUANTIZATIONNononeLoad models quantized for inference. Reduces VRAM at some quality cost.
TRUST_REMOTE_CODENofalseAllow models with custom Python code from HuggingFace. Security risk, enable only for trusted models.
CACHE_DIRNo~/.cache/parsynLocal cache for downloaded models and datasets.
TRAINING_TIMEOUT_SECONDSNo86400Maximum training time per job (24 hours).
INFERENCE_TIMEOUT_SECONDSNo300Maximum inference time per request (5 minutes).
MAX_INFERENCE_CACHED_MODELSNo2Number of models kept in GPU memory for fast inference switching.

* One of WORKER_API_KEY or WORKER_ENROLLMENT_KEY is required.

Starting the worker

parsyn-worker start

On startup, the worker:

  1. Resolves its API key (individual key, cached credentials from previous enrollment, or enrollment key)
  2. Connects to the platform via WebSocket
  3. Authenticates and (if enrolling) receives its individual API key
  4. Runs a GPU benchmark (cached for 24 hours)
  5. Sends hardware info and benchmark results to the platform
  6. Starts the heartbeat loop (every 30s) and metrics loop (every 5s)
  7. Waits for tasks

Expected output:

INFO  Worker abc12345 connecting to wss://parsyn.progatis.com/ws/worker
INFO  Connected to platform
INFO  Running GPU benchmark...
INFO  Benchmark complete: score=85.2, TFLOPS=82.6, bandwidth=1008 GB/s
INFO  Registered: NVIDIA A100-SXM4-80GB, 80 GB VRAM, CUDA 12.1
INFO  Worker ready, waiting for tasks...

GPU benchmark

On first connection (and every 24 hours), the worker runs an automated GPU benchmark. The platform uses benchmark scores for intelligent job assignment in auto mode, matching jobs to the most capable available worker.

The benchmark measures:

  • TFLOPS: FP16 matrix multiplication throughput
  • Memory bandwidth: Large tensor transfer speed (GB/s)
  • Max batch size: Binary search for largest batch that fits in memory
  • Composite score: Weighted combination (50% TFLOPS, 30% bandwidth, 20% VRAM), normalized against an RTX 4090 baseline

If the benchmark times out (60 seconds), the worker uses conservative fallback values and reports to the platform anyway.

Running as a system service

For production, use systemd so the worker starts on boot and restarts on failure:

[Unit]
Description=Parsyn Training Worker
After=network.target

[Service]
Type=simple
User=parsyn
WorkingDirectory=/opt/parsyn-worker
EnvironmentFile=/opt/parsyn-worker/.env
ExecStart=/opt/parsyn-worker/venv/bin/parsyn-worker start
Restart=always
RestartSec=10

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable parsyn-worker
sudo systemctl start parsyn-worker
sudo journalctl -u parsyn-worker -f

Multiple GPUs on one machine

Each worker process uses one GPU. Run multiple instances with different GPU_DEVICE values:

# Terminal 1
GPU_DEVICE=cuda:0 parsyn-worker start

# Terminal 2
GPU_DEVICE=cuda:1 parsyn-worker start

With systemd, use a template unit:

# /etc/systemd/system/[email protected]
[Service]
EnvironmentFile=/opt/parsyn-worker/.env.%i
ExecStart=/opt/parsyn-worker/venv/bin/parsyn-worker start

# Enable per GPU
sudo systemctl enable parsyn-worker@gpu0 parsyn-worker@gpu1

If using master key enrollment, each GPU instance enrolls as a separate worker and receives its own API key.

Fleet management

Master key lifecycle

Master keys support limits and expiration:

  • max_enrollments: cap the number of workers that can enroll (null for unlimited)
  • expires_at: expiration date after which new enrollments are rejected
  • auto_name_prefix: workers are named sequentially (e.g., "east-worker-1", "east-worker-2")

Revoking a master key

DELETE /api/v1/master-keys/{id}

Revocation prevents new enrollments but does not disconnect existing workers. Those workers already have their own individual API keys.

Monitoring your fleet

The Workers page shows summary cards at the top (total, online, training, idle, average GPU utilization) and a live table of all workers. A green pulsing dot indicates the real-time WebSocket connection is active.

Each row in the table shows:

  • Connection status (green dot for idle, red pulsing for training, gray for offline)
  • GPU utilization percentage with a progress bar, plus VRAM used
  • CPU and RAM usage bars
  • Current job ID with a progress bar and percentage (if training)
  • Worker type badge (Both, Training only, Inference only)
  • ETA for the current job

Click a row to open a detail modal. Use the action icons to view details or delete the worker.

Individual key persistence

When a worker enrolls via an enrollment key, it receives an individual key (wkr_...) that the worker process writes directly into its .env file under WORKER_API_KEY. On subsequent starts, the worker uses that key and skips the enrollment handshake.

If you need to re-enroll a worker (e.g., after losing the .env), start it again with the enrollment key.

Troubleshooting

Worker can't connect

  • Verify network access: curl https://parsyn.progatis.com/health
  • Check that the enrollment key is still active in the dashboard.
  • If behind a corporate firewall, ensure WebSocket (wss://) traffic on port 443 is allowed.

Worker connects but disconnects immediately

  • Check the worker logs for authentication errors.
  • Verify the key prefix: wkr_ for individual keys, enroll_ for enrollment keys.
  • Check that the enrollment key hasn't been revoked or reached its enrollment limit.

GPU not detected

  • Run nvidia-smi to confirm the GPU is visible.
  • Check PyTorch: python -c "import torch; print(torch.cuda.is_available())"
  • If using Docker, use the NVIDIA runtime: docker run --gpus all ...

Training crashes with OOM

  • MAX_BATCH_SIZE caps the batch size the worker accepts. Lower it if needed.
  • Enable quantization: QUANTIZATION=4bit
  • Make sure no other processes are using the GPU.

Model download fails

  • Check internet access to HuggingFace: curl https://huggingface.co
  • For gated models (Llama, Gemma, etc.), set HF_TOKEN with a token that has access.
  • Verify CACHE_DIR has enough disk space. 7B models need 15+ GB.