Preprocessing Training Data

Convert and clean raw data (XML, HTML, JSON, TXT) to fit the training purpose (PT/SFT), and connect it directly to the Training tab.

Preprocessing training data


Specifying the Source Folder

Specify the path of the source folder containing archive files (.zip/.tar.gz) or raw files (XML, HTML, JSON, TXT, etc.). The default path is /@team/.AppData/fine_tuning/raw.

Selecting the Target Training Stage

Select the purpose for which the data will be converted.

StageDescription
PT — Pre-training raw corpusUses the entire set of converted MD files as a domain corpus for pre-training
SFT — Instruction dataset (Alpaca JSONL)Converts data into an instruction dataset in instruction/input/output format
DPONot currently available

In PT mode, no separate output file is required. Once conversion is complete, .md files are created inside the source folder, and the source folder path is connected directly to the Training tab.

Raw Formats Contained in the Source Folder

For reference, this shows which formats of raw files are mixed in the source folder. You can choose among KIPRIS XML (Korean patents), HTML (web documents), JSON (structured data), and TXT/MD (plain text); this selection only affects the pipeline description below. The actual server-side processing batch-converts every file in the folder.

Running the Pipeline

The pipeline consists of two stages, which can be run individually or all at once in a batch.

  1. Automatic archive extraction: Automatically unpacks .zip/.tar.gz files in the source folder and lists the files.
  2. Document → Markdown conversion: Batch-converts document files (XML, HTML, JSON, etc.) in the source folder to Markdown. The converted .md files are created in the source folder.

Clicking the Run Full Pipeline in Batch button runs both stages in sequence.