AI Model Settings

If you have a GPU, you can easily use on-premises (local) AI through the SetFN workspace. You can also use sLLM (small language model)-based AI without issue in air-gapped or network-isolated environments. This page guides you through selecting a local AI runtime and configuring the inference endpoint used by SetFN.

AI model settings


Local AI

You can connect the SetFN workspace to an inference runtime running in your local environment to enable AI features.

Runtime Settings

You can choose one of the following inference runtimes.

RuntimeDescriptionDefault Endpoint Example
OllamaUses Ollama models installed locally.http://localhost:11434
LM StudioUses models run via LM Studio's local server feature.http://localhost:1234
Docker Model RunnerUses a model runtime running in a Docker environment.http://localhost:12434

Endpoint Settings

Enter the server address of the selected runtime.

  • If the runtime runs on the same machine, use a localhost address. (e.g., http://localhost:11434)
  • If the runtime runs on a different server, enter that server's IP or domain and port.
  • The endpoint you enter must be reachable from the SetFN server.

API Token Settings

The API Token is optional.

Enter an API Token only if authentication is enabled on the selected runtime. If your local runtime doesn't use authentication, you can leave this blank and save.

Installing a New Model

Enter the name of the model you want to use and click Install. The selected AI runtime will pull the model. If you use Ollama, you can find available model names and tags in the official Ollama model library.

For example, entering the following model name automatically downloads and installs that model:

gemma4:31b

Model files can be large. Before installation, check the server's available storage, memory, and GPU specifications.

Understanding Model Capabilities

Local AI runtimes such as Ollama display capability tags for the features each model supports. Because models are designed for different tasks, check that a model has the capabilities required for your intended use.

CapabilityMeaningCommon Uses
completionA general text-generation model that continues from the provided context and answers questions.Chat, document writing, summarization
thinkingSupports reasoning through a problem in multiple steps before producing an answer. It can improve results on complex tasks, but may take longer to respond.Complex questions, analysis, planning
tool useCan select external tools or functions and construct the arguments needed to call them. Whether a tool is actually executed depends on connected features and permissions.Agent tasks, search, external function calls
visionA VL (Vision-Language) model that can understand image input together with text.Image descriptions, scanned document and screen analysis
embeddingAn embedding-only model that converts text meaning into numeric vectors. It is not intended to generate regular chat responses.Knowledge indexing, similar-document search, RAG
cloudRuns in Ollama Cloud instead of entirely on local hardware. It is available after signing up for Ollama Cloud and authenticating the runtime.Environments with limited local resources, running larger models

A model can support multiple capabilities. For example, a model with both vision and completion can understand an image and respond in text. Use an embedding model as the search model under AI DB Settings.

cloud models run remotely without downloading the entire model file to the local machine. They require an internet connection and Ollama Cloud authentication. Available models and usage limits may vary according to the subscribed service policy.

Installed Models

Check the list of models available on the selected local AI runtime.

The installed model list shows each model's name, parameter count, size, and installation date. Models you no longer use can be removed with the Delete button.

Vision Model Priority

AI model settings

Select the VL (Vision-Language) models that can understand images and documents together, and set their response priority.

  • Select which enabled models to use for vision input.
  • Selected models are called in priority order.
  • Use the up/down buttons to adjust the call order.

AI Summary Model Priority

Select the models used to synthesize knowledge search results and create summaries (digests) for public chats, then set their call order.

Available Models

Choose from enabled models and add the ones you want to use for knowledge search and summarization to the priority list. Vision-only VL models and embedding-only models are excluded.

Selected Priority

Use the up/down buttons to change the call order. If a higher-priority model becomes unavailable—for example, because the runtime changed or its connection failed—the next model in the list is used automatically.

Automatic Selection Policy

Choose how a model is selected when the priority list is empty or none of the listed models are available.

PolicyDescription
Response speed prioritySelects a model that can respond quickly.
BalancedConsiders both response speed and output quality.
Quality priorityPrioritizes models that produce higher-quality summaries. Recommended when summary quality is important.

Editor AI Autocomplete

Configure AI to suggest the next sentence or continuation while you write in the Markdown editor.

Editor AI autocomplete

When enabled, AI suggests text based on the cursor position and surrounding context while you edit a .md document. This is useful when continuing meeting minutes, reports, memos, or technical documentation.

To see how autocomplete appears in the editor, refer to the AI Text Suggestions section in Viewing and Editing Text Files.

Model Selection Method

The model used for editor AI autocomplete can be selected in two ways.

Selection MethodDescription
Automatic selectionAutomatically selects an available model that fits the policy.
Specific modelUses the local AI model designated by the administrator.

To use a specific model, first install or connect it under Installed Models on this page.

Automatic Selection Policy

When using automatic selection, choose a model-selection policy based on the following criteria.

PolicyDescriptionSuitable Situations
Response speed priorityPrioritizes models that can respond quickly.Short notes, real-time writing assistance
BalancedConsiders both speed and quality.General document writing, common team defaults
Quality priorityPrioritizes models that produce higher-quality responses.Reports, proposals, long-document drafts

Saving Settings

After confirming the runtime, endpoint, API Token, and model settings, save the configuration. Once saved, AI features such as AI chat, document-based Q&A, and knowledge indexing can use the configured local runtime and model.

Good to Know

  • The models used by the runtime must first be downloaded or loaded in Ollama, LM Studio, or Docker Model Runner.
  • If the runtime is not running or the endpoint address is incorrect, SetFN cannot call the model.
  • Check that port access isn't blocked by a firewall or network policy.