Skip to main content
User guide

AI settings

The LLM providers the gateway uses, reasoning effort, failover order, the knowledge base embedding model, field masking before data reaches the model, and failed cases. Available in every edition; multi-provider failover needs Enterprise.

AI settings is a tab of the Settings page. It manages the LLMs the gateway calls; only administrators can view or change it. Each provider is an OpenAI-compatible endpoint or an Azure OpenAI deployment. You bring your own model, and the gateway does not limit the number of calls.

The provider list is stored in llm_providers.yml on the gateway (RST_LLM_CONFIG). Without that file, the gateway uses LLM_API_KEY, LLM_BASE_URL and LLM_MODEL from .env as its only provider.

AI settings

Add or replace a provider

Prerequisites

  • The administrator role.

Steps

  1. Under Model routing, select Add a provider.
  2. Select Expand the settings, pick the Kind, and fill in the Model, base URL and API key. Set Tags, Timeout (seconds) and Reasoning as needed.
  3. Check that the provider's enable switch is on (it is on for a new provider).
  4. When replacing, select Remove on the old provider.
  5. Select Save and reload.

Saving an empty list makes every model call fail, so when replacing a provider, add the new one first, remove the old one, then save.

FieldDescription
Kindopenai (and compatible gateways) covers Volcengine Ark, DeepSeek, Qwen, Ollama, vLLM and the like. Azure needs the API version and the Deployment name (not the model name)
ModelFor example ark-code-latest, or the endpoint id for a self-built endpoint
API keyPaste a new key. Once one is saved, leave it empty to keep it; the page shows its last 4 characters
TagsFor example primary,cloud; labels only
Timeout (seconds)Limit for one call, 30 by default for a new provider. Fault investigation and alert correlation can take minutes per call on a reasoning model, so raise it for slow models. Calls at high reasoning wait at least RST_LLM_THINKING_TIMEOUT_S seconds (600 by default)
ReasoningSee below

Set the reasoning effort

Reasoning models (Doubao, GPT-5, Claude, Qwen3 and others) think before they answer, which adds tens of seconds per call. Reasoning is set per provider:

SettingBehavior
Auto (recommended)Decided per task: no thinking for titles or follow-up suggestions, thinking for fault investigation
OffNo task thinks. Use it when the model is too slow
Low, HighEvery task uses that level. Pick High when accuracy is all that matters

The gateway recognizes the vendor from the base URL and model name and translates the setting into that vendor's parameters. If a request with reasoning parameters is rejected (HTTP 400), it is retried once without them.

Failover

Failover chain lists the enabled providers in the order of the Model routing list. A new provider goes to the end of the list.

  • Enterprise (including a trial): when a call fails, the same request moves on to the next provider in turn.
  • Community and Professional: only the first enabled provider is used and there is no switch when it fails; the others show Standby · no switch.

Check model health

MetricMeaning
Model availabilityHow many enabled providers are healthy. The counts reset when the gateway restarts
Failover chainHow many providers calls can currently fail over through
Calls (24h)Calls, failures and p95 latency from the audit log. Not shown when auditing is off
Failures since startFailures since the gateway started, whether a provider is failing repeatedly right now, and the last failure

Failover chain and call activity shows the provider in use, each provider's last success and consecutive failures, and the call volume over the last 24 hours.

Set the knowledge base embedding model

Runbook embedding model embeds and searches the knowledge base. Fill in the Model ID; leave the endpoint and key empty to reuse the chat model's. Select Test connection, and Vector dimensions fills in. Turn on Enable and save. You cannot save before the test passes.

Field masking

Before sending data to the model, the gateway masks it according to the masking mode. Content pushed to external channels such as Feishu or email is masked the same way. The mode is set under Field masking in Settings; all three modes are available in every edition.

ModeWhat happens before sendingUse when
CloudIPs keep their first three octets, such as 10.30.12.x. Account names keep their first and last characters; the local part of email addresses is hidden. Secrets, tokens, cookies, ID numbers and similar become [REDACTED]. IPs and phone numbers in free text are masked tooThe model runs in a public cloud
PrivateIPs keep their first two octets, such as 10.30.x.x. Account names and email addresses are partly hidden. Secrets and tokens are still replacedThe model runs in your own data center
Air-gappedNothing is replacedA local model; data never leaves the network

If the page has no setting, the mode comes from RST_MASKING_MODE in .env, and defaults to Cloud.

Review failed cases

The Feedback tab lists the queries users marked as wrong in the chat, with the question, the generated API call, the API call the user suggested and any comment. Filter or delete them, or select Export JSON to use them as a regression set when changing prompts.

On this page