Preprocessing Training Data
Convert and clean raw data (XML, HTML, JSON, TXT) to fit the training purpose (PT/SFT), and connect it directly to the Training tab.

Specifying the Source Folder
Specify the path of the source folder containing archive files (.zip/.tar.gz) or raw files (XML, HTML, JSON, TXT, etc.). The default path is /@team/.AppData/fine_tuning/raw.
Selecting the Target Training Stage
Select the purpose for which the data will be converted.
| Stage | Description |
|---|---|
| PT — Pre-training raw corpus | Uses the entire set of converted MD files as a domain corpus for pre-training |
| SFT — Instruction dataset (Alpaca JSONL) | Converts data into an instruction dataset in instruction/input/output format |
| DPO | Not currently available |
In PT mode, no separate output file is required. Once conversion is complete, .md files are created inside the source folder, and the source folder path is connected directly to the Training tab.
Raw Formats Contained in the Source Folder
For reference, this shows which formats of raw files are mixed in the source folder. You can choose among KIPRIS XML (Korean patents), HTML (web documents), JSON (structured data), and TXT/MD (plain text); this selection only affects the pipeline description below. The actual server-side processing batch-converts every file in the folder.
Running the Pipeline
The pipeline consists of two stages, which can be run individually or all at once in a batch.
- Automatic archive extraction: Automatically unpacks .zip/.tar.gz files in the source folder and lists the files.
- Document → Markdown conversion: Batch-converts document files (XML, HTML, JSON, etc.) in the source folder to Markdown. The converted .md files are created in the source folder.
Clicking the Run Full Pipeline in Batch button runs both stages in sequence.