Core Concepts
The building blocks of Parsyn: how the platform, workers, datasets, models, and credits fit together.
The platform
The Parsyn platform is the central hub. It runs at parsyn.progatis.com and handles everything except the actual GPU compute: user management, dataset storage, job orchestration, metrics collection, billing, and the web dashboard.
All state lives in the platform. Workers are stateless, they connect, receive tasks, do the work, and report back. This means you can add, remove, or replace workers at any time without losing data or progress.
Workers
Workers are the compute layer. A worker is a process running on a GPU machine that connects to the platform via WebSocket and executes tasks: training, inference, evaluation, or post-training operations.
Parsyn supports two types of workers:
Parsyn Workers (managed)
GPU machines managed by the Parsyn team. You don't install anything. Just select the GPU type you want when creating a training job, and the platform assigns one from its fleet.
- Billed in credits per hour (price varies by GPU type: A100, RTX 4090, etc.)
- Available for training only (not inference)
- Managed, updated, and monitored by the Parsyn team
- Subject to fleet availability
User Workers (bring your own GPU)
Your own GPU machines. You install the parsyn-worker client, configure it with an API key, and it connects to the platform. From that point, you manage the hardware; the platform manages the workload.
- No cost beyond your own hardware and electricity
- Supports training, inference, and evaluation
- You control uptime, maintenance, and physical security
- Configurable as
fine_tuner(training only),prompter(inference only), orboth
Worker lifecycle
When a worker connects to the platform, it:
- Authenticates with its API key (or enrolls via a master key on first connection)
- Runs a GPU benchmark and sends hardware info (GPU model, VRAM, CUDA version, benchmark score)
- Enters
idlestate and waits for tasks - Sends a heartbeat every 30 seconds
- Reports system metrics (GPU utilization, memory, CPU, RAM) every 5 seconds
If the platform doesn't receive a heartbeat for 90 seconds, the worker is marked offline. Workers automatically reconnect on disconnection with exponential backoff (5s to 60s).
Datasets
A dataset is a file stored in S3-compatible storage with metadata tracked in the platform's database. Supported formats:
| Format | Best for |
|---|---|
| JSONL (recommended) | Instruction tuning, chat data. One JSON object per line. Supports nested structures like chat messages. |
| CSV | Simple prompt/completion pairs, tabular data. |
| Parquet | Large datasets (1 GB+). Efficient compression, fast loading. |
Dataset lifecycle
- Upload: Files go through a multipart upload pipeline. Large files are chunked and reassembled server-side. The file is stored in S3.
- Validation: The platform checks format, schema (expected fields), encoding (UTF-8), and data integrity.
- Statistics: An async job computes stats: example count, field distributions, token counts, duplicate detection. Runs in the background via Celery.
- Ready: The dataset can now be used in training jobs. You can also run operations: split, deduplicate, convert format, or normalize the schema.
Models
A model in Parsyn comes from one of two sources:
- HuggingFace Hub: Provide a model ID (e.g.,
meta-llama/Llama-3.1-8B). The worker downloads it at training time. - Custom upload: Upload your own model files (safetensors, PyTorch checkpoints) to the platform's storage.
Model states
| State | Meaning |
|---|---|
registered | Metadata created, ready to use as a base for training. |
training | A training job is actively using this model. |
ready | Training complete, model can be used for inference. |
failed | Training failed. Check the job logs. |
Fine-tuning creates a new model that references the original base. You get a full lineage: base model, training config, dataset, and all checkpoints.
Training jobs
A training job ties together a dataset, a base model, a training configuration, and compute resources.
From the Training page, click New Training Job to open the creation form. You select a dataset, a base model, configure hyperparameters (learning rate, batch size, epochs, LoRA settings, etc.), and choose your worker selection mode. After creating the job, start it with the play button.
The lifecycle:
- Create: Define the job through the form. Status:
pending. - Start: The platform queues the job and assigns it to an available worker. Status:
queued. - Dispatch: The worker receives the job via WebSocket, downloads the dataset and model, and begins training. Status:
running. - Progress: The worker streams metrics back every few seconds. The training detail page shows live charts: loss curve, learning rate, GPU utilization, throughput, and ETA.
- Completion: The worker uploads the fine-tuned model to S3 and reports success. Status:
completed.
Job states
| State | Description |
|---|---|
pending | Created, not yet started. |
queued | Started, waiting for a worker to pick it up. |
running | Worker is training. Metrics are streaming. |
completed | Finished. Fine-tuned model is available. |
failed | Something went wrong. Check the error in the job details. |
cancelled | Stopped by the user. A checkpoint may be available. |
Worker selection modes
When creating a job, you choose how the platform assigns compute:
| Mode | Behavior |
|---|---|
auto | The platform picks the best available worker, using GPU benchmark scores to match the job. Includes both Parsyn and your workers. |
specific_user_workers | You choose which of your connected workers to use. |
specific_parsyn_workers | You choose which GPU type from Parsyn's fleet. Billed in credits. |
mixed | You select from both your workers and Parsyn's fleet. |
Checkpoints
During training, checkpoints are saved to S3 at regular intervals. Each checkpoint includes model weights (or LoRA adapter), optimizer state, and metrics. If training is interrupted, you can resume from the last checkpoint.
Credits and billing
Parsyn uses a credit-based system for Parsyn Worker usage. Credits are prepaid: you purchase a credit pack, and credits are deducted as you use managed GPUs.
The Credits page in the dashboard has three tabs:
- History: Every credit transaction (purchases, consumption, refunds, bonuses) with filters and pagination. Each entry shows the amount, resulting balance, and a description.
- Buy: Available credit packs with prices. Click Purchase to complete payment through Stripe.
- Pricing: Credits-per-hour rate for each GPU type in Parsyn's fleet.
Your current balance, lifetime purchased, and lifetime consumed are always visible at the top of the page. Using your own workers does not consume credits. Credits only apply to Parsyn-managed GPUs.
Real-time communication
The platform uses WebSocket for two channels:
- Platform to worker: Sends commands (start training, stop, load weights). Receives progress, metrics, status updates, and errors.
- Platform to dashboard: Pushes live updates to the web UI. When a worker reports metrics, the platform broadcasts them to all users watching that job.
Metrics are forwarded to the dashboard in real time regardless of database write throttling. The dashboard charts update live with loss curves, GPU utilization, throughput, and ETA.
Data isolation
Every resource (datasets, models, training jobs, workers) is scoped to a user. The backend enforces this at the query level: every database query includes a user_id filter. There is no way for one user to access another user's data through the API.
Parsyn Workers are shared infrastructure, but each job runs in isolation: a Parsyn Worker is assigned to one user's job at a time, and the worker is released after the job completes.
Inference
Once a model is fine-tuned, you can test it through the built-in chat interface. The platform routes inference requests to one of your connected workers that has prompter or both capability. The worker loads the model (with LRU caching for fast switching), generates tokens, and streams them back to the dashboard.
Inference features include adjustable parameters (temperature, top_p, top_k, max tokens), conversation history, system prompts, and A/B model comparison.
Inference requires your own workers. Parsyn Workers handle training only. You need at least one connected worker with prompter or both type to use the chat feature.
Training pipelines
A pipeline automates a sequence of steps that would otherwise require manual intervention between them. Instead of starting a training job, waiting for it to finish, then manually triggering an evaluation, and then an export, you define the full sequence upfront and the platform executes it end-to-end.
Steps in a pipeline can include:
- train: Run a training job with a given dataset, model, and config.
- evaluate: Run an evaluation (or evaluation suite) against the output of the previous training step.
- export: Export the trained model to a deployment format (GGUF, ONNX, GPTQ, AWQ).
Each step depends on the result of the previous one. If a step fails, the pipeline stops and reports the error — the subsequent steps are not executed. You can monitor the pipeline's progress from the dashboard, which shows the state of each step individually.
Pipelines are particularly useful for reproducible experiment loops: train a model, evaluate it against a fixed suite, export it if evaluation passes a threshold.
Teams and organizations
Organizations allow teams to collaborate on the platform while keeping resources structured and isolated.
The hierarchy is:
- Organization: Top-level entity, typically a company or research group. Has a name, members, and billing configuration.
- Teams: Sub-groups within the organization. A user can belong to multiple teams, each with its own role (e.g., admin, member).
- Projects: A project belongs to an organization and scopes a set of shared resources — datasets, models, and training jobs — to the members of the associated team.
Within a project, team members can see and work with shared resources. Resources in one project are not accessible from another, even within the same organization. This gives you per-team isolation while still allowing organization-level oversight.
Personal resources (datasets and models created outside any organization) remain private to the user who created them.
Evaluation suites
An evaluation suite is a reusable set of prompts with expected outputs or scoring criteria. Once defined, you can run the suite against any model to produce a scored result. Running the same suite against multiple models gives you a direct, apples-to-apples comparison.
This is different from the simpler Evaluations feature, which measures a single metric (e.g., perplexity) on a dataset. Evaluation suites are designed for structured, repeatable quality assessment:
- Define prompt sets once, reuse them across every training run.
- Each run produces a scored result stored in the platform.
- Compare results across model versions, datasets, or hyperparameter configurations.
- Use suites as a gate in a training pipeline — only proceed to export if evaluation passes.
Suite runs are tracked historically, so you can chart how quality evolves across training iterations.