P
Parsyn
/Docs
Back to Home

Dataset Management

Uploading, validating, and preparing datasets for fine-tuning. Covers supported formats, the upload pipeline, and built-in data operations.

Supported formats

JSONL (recommended)

One JSON object per line. The preferred format for instruction tuning because it naturally represents prompt/completion pairs and supports nested structures.

{"prompt": "Translate to French: Hello, how are you?", "completion": "Bonjour, comment allez-vous ?"}
{"prompt": "Summarize: The cat sat on the mat.", "completion": "A cat rests on a mat."}
{"prompt": "Extract the entity: John works at Google.", "completion": "Person: John, Organization: Google"}

For chat-style fine-tuning, use a messages field with role-based formatting:

{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is Python?"}, {"role": "assistant", "content": "Python is a high-level programming language."}]}
{"messages": [{"role": "user", "content": "Explain recursion."}, {"role": "assistant", "content": "Recursion is when a function calls itself to solve a problem by breaking it into smaller instances of the same problem."}]}

CSV

Standard comma-separated values. Column headers are required.

prompt,completion
"What is a neural network?","A neural network is a computational model inspired by biological neurons."
"Define overfitting.","Overfitting is when a model memorizes training data instead of learning general patterns."

CSV fields containing commas, newlines, or quotes must be properly escaped. Parsyn validates this during the validation step, but malformed CSVs will be rejected.

Parquet

Columnar format with efficient compression. Recommended for datasets larger than 1 GB. Same column structure: prompt/completion or messages.

Uploading a dataset

Go to the Datasets page and click Upload Dataset. In the dialog:

  1. Enter a name for your dataset
  2. Add an optional description
  3. Select your file (.jsonl, .json, .csv, or .parquet) or drag and drop it into the file area. Maximum file size is 10 GB.
  4. Click Upload

The file is uploaded through a multipart pipeline. Large files are automatically chunked, reassembled server-side, and stored in S3. If an upload is interrupted, you can retry without losing progress.

Once uploaded, the dataset appears in your list with its name, format badge, size, and row count.

Dataset detail view

Click any dataset in the list to open its detail page. The detail view has several tabs:

  • Overview: Name, description, format, size, row count, creation date.
  • Schema: Detected field names and types.
  • Distributions: Charts showing token count distributions and field value distributions.
  • Validation: Data quality report with any issues found.
  • Preview: Sample rows from the dataset in a table.

Statistics

In the dataset detail page, click Compute Statistics to trigger an analysis. This runs asynchronously in the background. The page polls for progress and shows results when done.

Statistics include:

  • Total example count
  • Average, min, and max prompt/completion length (characters and tokens)
  • Field distributions for categorical data
  • Duplicate count

Results appear in the Distributions tab with visual charts.

Data operations

From the dataset detail page, open the Preprocessing Wizard to run operations on your data. Each operation creates a new dataset, keeping the original intact. You'll be asked to name the output dataset.

Deduplication

Removes duplicate examples. Two modes:

  • Exact: Removes rows that are character-for-character identical.
  • Fuzzy: Removes rows above a similarity threshold (default 90%). Useful for catching near-duplicates with minor variations.

Normalization

Cleans up text formatting. Options include:

  • Strip leading/trailing whitespace
  • Collapse consecutive whitespace
  • Convert to lowercase
  • Remove URLs
  • Remove HTML tags

Train/validation split

Splits the dataset into separate training and validation sets (and optionally a test set). Configure the ratios (must sum to 100%), toggle shuffling, and set a seed for reproducibility.

Format conversion

Converts between JSONL, CSV, and Parquet. All six directions are supported.

Sampling

Creates a smaller dataset from a random sample. Set either a fixed number of examples or a percentage of the original data. Useful for quick experiments before committing to a full training run.

Schema normalization

Maps non-standard field names (input/output, instruction/response, etc.) to the standard prompt/completion schema expected by the training pipeline.

Run deduplication and normalization before splitting. This way, your validation set is clean and doesn't contain duplicates from the training set.

Best practices

Data quality

  • Be consistent. Same format and style across all examples. If some completions end with a period and some don't, pick one convention.
  • Include variety. 100 diverse examples often outperform 1000 similar ones for instruction tuning.
  • Check for contamination. If your evaluation set overlaps with training data, your metrics will be misleading.

Sizing

  • LoRA instruction tuning: 500 to 5,000 examples is a good starting range.
  • Full fine-tuning: 10,000+ examples to see meaningful improvement.
  • Quality matters more than quantity. A small, clean dataset beats a large, noisy one.

File format

  • Use JSONL for datasets under 1 GB. Human-readable, easy to debug.
  • Use Parquet for large datasets. Better compression, faster loading.
  • Avoid CSV for complex data (nested objects, multi-line text).