Knowledge Indexing

Check whether the documents users have uploaded are ready to be used in AI search. The search data created here is used for RAG, where the AI locates relevant documents to answer questions.

Knowledge indexing


Indexing Summary

Check the current indexing status at a glance.

ItemDescription
Pending updateNewly created or modified files waiting to be reflected in AI search
Pending deletionSearch data for deleted files waiting to be cleaned up
Pending embeddingDocuments whose text has been read and are waiting to be turned into AI search data
Embedding countThe total number of embeddings completed so far
Indexing modeThe scope of files targeted for AI search data creation (e.g., full indexing)
Embedding providerThe AI embedding model in use (e.g., Ollama bge-m3:latest / 1024d)
Included execution rootThe path targeted for indexing
Team checkpointThe last time team folders were checked
Client checkpointThe last time personal folders were checked
Last refreshThe last time the status was refreshed

You can leave Auto Refresh on, or manually refresh the status using the Refresh button.

Check Settings Before You Start

Before starting indexing, review the required settings below. Each link takes you to the relevant setting.

Required SettingWhat to CheckGo to Setting
AI runtimeSelect an AI runtime such as Ollama or LM Studio for indexing, and confirm that it is connected.Configure AI runtime
Vision model (VL)Select a vision model that reads the contents of PDFs and images.Configure vision model
AI vector DBSelect vector storage such as Local (SQLite), Milvus, or Qdrant, and configure its connection.Configure AI vector DB
Embedding modelSelect a model that converts documents into search vectors. (e.g., bge-m3)Configure embedding model

Selecting a vision or embedding model does not install it automatically. Download each required model to the configured AI runtime before starting indexing.

Collection and Processing Rules

Define when to create search data and which folders' documents are used for AI search.

Enable Knowledge Indexing

Turns the entire knowledge indexing feature on or off.

Indexing Mode Settings

ItemDescriptionDefault
Indexing modeSelect the indexing scope and methodFull indexing
Wait time (minutes)Time to wait after detecting a file change before starting indexing20
Maximum concurrent jobsNumber of indexing jobs that can run at the same time1
Allow worker executionIf turned off, changed files continue to be detected, but no AI search data is created.Enabled
Auto-pause under high loadPauses indexing temporarily when server usage is high, and resumes once the server stabilizes.Enabled
Boost modeStarts indexing within 5 seconds of a file change and immediately processes any backlog. Enable this when you want changes reflected quickly.Disabled
Allowed time windowSpecifies the time window during which indexing is allowed. If left blank, indexing can run at any time. (e.g., 01:00-06:00)No restriction

Included Roots

Specify the folder paths to index. If left blank, the entire path is targeted.

Example:

/@team/docs
/@client/projects

AI Embedding Model Settings

Select the AI model used to convert documents into search data.

ProviderDescription
OllamaUses Ollama models installed locally. (e.g., bge-m3:latest)
ONNX Local CPUUses a locally run ONNX model based on CPU.
ONNX Local GPUUses a locally run ONNX model with GPU acceleration.
Remote APIUses an embedding model via an external API.

Providers with no installed model are shown as No models available.

ItemDescription
Dimension sizeThe size of the search data used by the AI model. Set automatically based on the selected model. (e.g., 1024)
PDF/image extraction strategySelect the method used to read text from PDFs and images. VL first, fall back to OCR means the vision-understanding AI is used first, and character recognition (OCR) is used only if that fails.

PDF/Image Include Patterns

Enter file extensions or path rules for the PDFs and images from which text should be read.

Example:

*.pdf, *.png, image/**/*

PDF/Image Exclude Patterns

Enter file extensions or path rules for PDFs and images that should not be processed. If left blank, all files matched by the include patterns above are processed.

Example:

thumbs.db, *.gif, cache/**

Progress Status by Storage Space

Check how much material in each team or personal storage space is waiting to be delivered to the AI.

You can search by space (team/client) and user, and sort by the number of pending updates.

ItemDescription
SpaceTeam space or personal (client) space
UserThe relevant user, for personal space
Pending updateNewly created or modified files waiting to be reflected in AI search
Pending deletionSearch data for deleted files waiting to be cleaned up
Pending embeddingDocuments whose text has been read and are waiting to be turned into AI search data
Embedding countThe total number of embeddings completed

AI Training Data Status

Check how much AI search data has been generated per model. This is not a feature for training AI models from scratch using the original documents.

ItemDescription
ProviderThe model provider used for embedding (e.g., ollama)
ModelThe embedding model used (e.g., bge-m3:latest)
CountThe total number of embeddings generated by that model

You can sort by count to compare the scale of search data by model.