Training

Run fine-tuning by specifying a preprocessed dataset, a base model, and a training method (STAGE).


1. Selecting a Training Dataset

Select a dataset folder that has already been prepared for training, not the original source folder. Choose among Team AppData Training Set, Personal AppData Training Set, or enter a Custom path directly. Once you select a folder, it's immediately validated below and shows the number of files available for training.

Selecting a training dataset

Data Format

Choose from Auto-detect, Alpaca, ShareGPT, Pairwise, and Plain Text.

  • SFT uses JSON/JSONL based on instruction/input/output.
  • DPO/ORPO uses JSON/JSONL based on chosen/rejected.
  • .txt/.md raw text is exclusive to Plain Text (PT pre-training); selecting it automatically switches the stage to PT.

Training File Extensions and Output Directory

Select the file extensions to include in training (.md, .txt, .json, .jsonl), and specify the output directory where training results will be saved. Results are automatically created in subfolders by run number, keeping each run's results separate.

2. Selecting a Model

Enter a HUGGINGFACE MODEL ID directly, or search for a model using the search bar. You can also quickly choose from Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Gemma-2-9B-It, Llama-3.1-8B-Instruct, DeepSeek-R1-14B, and DeepSeek-R1-8B.

By default, CHAT TEMPLATE uses automatic detection based on the model name (recommended). Select it manually only if it isn't working correctly.

Selecting a model and training method

3. Selecting a Training Method (STAGE) and Fine-tuning Approach

STAGEDescription
SFT (Supervised Fine-Tuning)Teaches a specific task using instruction/input/output data
DPOTrains preferences using chosen/rejected data
ORPOA preference-based alignment method similar to DPO
PT (Pre-Training)Continued Pre-Training. Trains domain knowledge using plain text

For FINETUNING TYPE, choose from LoRA, QLoRA (4-bit quantization), Full (full fine-tuning), and Freeze (layer freezing). You can also adjust detailed hyperparameters such as LoRA RANK, LORA ALPHA, and DROPOUT.

4. Starting Training

Click the Start Training button to begin training. In the right panel, you can monitor status, model, RUN ID, result model path, and progress in real time, along with output logs. To stop training, click the Stop button.

Good to Know

  • If the data format doesn't match the target stage, training may not start.
  • You can't change the settings of the same run while it's training.