Synced from monorepo Changes: - Workspace server: report `/ready` as failed with dwell on hub connect failure - Refresh OIDC token for the Grok agent in the shell - ACP terminal output recorder - Cross-platform provider auth commands in the shell - Default `/resume` to Grok sessions with a hint for hidden external sessions - Resume sessions by title with `--resume` - Limit app-builder archive size - Data-driven tag labels for slash commands - Doctor fixes for tmux - Custom provider gateways and subprocess environment policy in the shell - `/tutorial` — opt-in onboarding tour of Grok Build - Soft and required CLI version checks in the shell - Privacy banner env overrides survive live settings updates - Add remote flag to override the image-edit model - Return profile fields from auth info even when the access token is expired - Add edit control on queued prompt rows - Keep fail-closed policy when clearing orphans with no team - Setting to disable the Ctrl+Space/F8 voice shortcut - Pass `--raw` to pw-record so Linux dictation works on older PipeWire - Validate git URLs when adding marketplace entries - Stop shipping stale tool-doc parameter and tool names - Re-point dashboard attach after `/fork` only when the parent was attached - Surface Grok Computer media-generation results as file-path chunks - Clear web background-task tray on kill and keep the task description - Show privacy upsell banner in agent view until acted on - Add tools-server client callback surface - Protect persistent global hook sources Source-Revision: 95d84f443eddcbed6cbfd6eed22e2eafe6b3939d
13 KiB
Custom Models
Grok connects to custom model endpoints for alternative providers, self-hosted models, and overriding built-in settings. This guide explains how to select models, configure endpoints, and integrate third-party providers.
Default Models
By default, Grok uses models hosted by SpaceXAI, and new sessions start with grok-build. Default models require no configuration. Authenticate with grok login or an API key, then start a session.
List all available models:
grok models
Selecting a Model
CLI Flag
grok -p "Hello" -m grok-build
Slash Command
In the TUI, switch models during a session:
/model grok-build
Or use the alias:
/m grok-build
Model Picker (Ctrl+M)
Press Ctrl+M from the scrollback pane to open the model picker. It lists all available models, both built-in and custom, and lets you switch with a single keystroke. With the prompt focused, Ctrl+M toggles multiline input instead -- use /model to switch without leaving the prompt.
Config Default
Set a persistent default in ~/.grok/config.toml:
[models]
default = "grok-build"
Supported API Backends
Grok supports three API backends. Set api_backend in your [model.*] config to choose which protocol the model uses:
| Value | API | Default |
|---|---|---|
"chat_completions" |
OpenAI Chat Completions (/v1/chat/completions) |
Yes |
"responses" |
OpenAI Responses (/v1/responses) |
|
"messages" |
Anthropic Messages (/v1/messages) |
When you omit api_backend, Grok uses chat_completions.
To send provider-specific authentication or version headers -- for example, Anthropic's x-api-key -- use the extra_headers field described below. Grok sends those headers verbatim with every request to the endpoint.
Configuring Custom Models
Add custom model endpoints in ~/.grok/config.toml under [model.<name>] sections:
[model.my-model]
model = "model-id" # Model identifier sent to the API
base_url = "https://api.example.com/v1" # OpenAI-compatible endpoint
name = "Display Name" # Shown in the model picker
description = "Model description" # Optional description
api_key = "sk-..." # API key for this provider (optional)
env_key = "XAI_API_KEY" # Env var holding the API key (optional; string or array)
api_backend = "chat_completions" # "chat_completions", "responses", or "messages"
temperature = 0.7 # Sampling temperature
top_p = 0.95 # Nucleus sampling parameter
max_completion_tokens = 8192 # Maximum tokens per response
context_window = 128000 # Total context window in tokens
extra_headers = { "x-api-key" = "sk-..." } # Extra request headers, sent verbatim (optional)
query_params = { api-version = "2026-07-22" } # Query params appended to every request URL (optional)
env_http_headers = { "X-Tenant" = "TENANT_TOKEN" } # Headers from env vars, resolved at client build (optional)
Credential Resolution
Grok resolves the API key in this order:
- The
api_keyfield in the model config - The environment variable(s) named by
env_key— a single string or an array of names. The first set, non-empty value wins (for exampleenv_key = ["ANTHROPIC_AUTH_TOKEN", "LC_ANTHROPIC_AUTH_TOKEN"]for SSHLC_*forwarding) - Your signed-in session token (from
grok login), for a model with noapi_key/env_keyof its own - The
XAI_API_KEYenvironment variable (global fallback; Grok also acceptsGROK_CODE_XAI_API_KEYfor backward compatibility)
Context Window
The context_window value tells Grok when to trigger auto-compaction. When you override a known model, Grok inherits that model's context window. When you define a new model and omit context_window, Grok defaults to 200,000 tokens, so set it explicitly to match your provider.
Global Default Headers
To apply the same headers to every model in the catalog -- built-in, prefetched from /v1/models, or custom -- set them once under the global [models] section instead of repeating them per model:
[models]
extra_headers = { "X-Request-Tags" = "team=example,env=prod" }
These act as a base for each model's inference requests. A per-model [model.<id>].extra_headers entry overrides the global default per key (matched case-insensitively): a key set on the model wins, while any global-only keys are still inherited by that model. Like the per-model field, they ride on that model's inference calls -- not on separate services such as image generation or video generation -- which makes them handy for attribution tags (for example, cost tracking) without re-declaring them whenever a new model appears.
Global Default Values
A few common per-model settings can also be set once under [models] as a default for every model. A per-model [model.<id>] value always wins; the global only fills in where a model (or the server's model list) left the field unset:
[models]
temperature = 0.7
top_p = 0.95
max_completion_tokens = 8192
max_retries = 8
inference_idle_timeout_secs = 600
stream_tool_calls = true
This is a small, fixed set of environment-wide knobs. Settings that identify a specific model (model, base_url, api_key, context_window, ...) cannot be defaulted this way, and a few settings with their own dedicated configuration -- auto-compaction ([session]), the system-prompt label ([agent]), and reasoning effort ([models].default_reasoning_effort) -- keep their existing homes.
Note on
stream_tool_calls: this one affects request shape, not just sampling. A few endpoints (some BYOK providers) expect it left unset; if a globalstream_tool_calls = truecauses problems for such a model, opt that model out withstream_tool_calls = falsein its[model.<id>]block.
Request Query Parameters
Some gateways route or version on the query string. query_params appends percent-encoded query parameters to every request Grok makes for a model. For example, a gateway that selects an API version this way:
[model.my-gateway]
model = "my-model"
base_url = "https://gateway.example/v1"
api_backend = "responses"
env_key = "GATEWAY_API_KEY"
query_params = { api-version = "2026-07-22" }
A key that also appears in the base_url query string is overridden (last value wins) rather than duplicated. Query parameters are saved in the session, so do not put secrets in them: use env_http_headers for a secret.
Environment-Variable Headers
env_http_headers maps a request header to the name of an environment variable that supplies its value, so a per-request secret never has to be written into config.toml:
[model.gateway]
model = "my-model"
base_url = "https://gateway.example/v1"
env_http_headers = { "X-Tenant-Token" = "GATEWAY_TENANT_TOKEN" }
Grok reads each variable when it builds the client for a session and places the value in the request headers only, never on disk. A header is skipped when its variable is unset or blank, and a resolved value overrides an extra_headers entry of the same name. Use extra_headers for a static value and env_http_headers for one that comes from the environment.
Both fields also work on a shared [model_providers.<id>] block. A model that points at a provider with model_provider = "<id>" inherits the provider's query_params and env_http_headers when it sets none of its own, matching how extra_headers is inherited.
Overriding Built-in Models
You can override specific fields of built-in models without redefining everything. Only specify the fields you want to change:
# Override only the API key for a default model
[model.grok-build]
api_key = "my-api-key"
# Override temperature and add a custom API key
[model.grok-build]
temperature = 0.5
api_key = "sk-custom"
When you override a built-in model, Grok starts with the default configuration (including the correct base_url), then applies only the fields you specify. Unspecified fields inherit from the default.
Priority Order
- Your config (
[model.*]) -- highest priority - Prefetched models from remote
/v1/models - Hardcoded defaults -- lowest priority
Provider Examples
Anthropic (Claude)
Use Claude models directly via the Anthropic Messages API:
[model.claude-opus]
model = "claude-opus-4-6"
base_url = "https://api.anthropic.com/v1"
name = "Claude Opus 4.6"
api_backend = "messages"
context_window = 200000
extra_headers = { "x-api-key" = "sk-ant-...", "anthropic-version" = "2023-06-01" }
The messages backend uses the Anthropic Messages protocol. Anthropic authenticates with an x-api-key header rather than Authorization: Bearer, so pass your key through extra_headers, which Grok sends verbatim.
OpenAI (Chat Completions)
[model.gpt-4o]
model = "gpt-4o"
base_url = "https://api.openai.com/v1"
name = "GPT-4o"
env_key = "OPENAI_API_KEY"
api_backend defaults to "chat_completions", so you don't need to set it explicitly for OpenAI.
OpenAI (Responses API)
If your provider supports the newer Responses API:
[model.gpt-4o-responses]
model = "gpt-4o"
base_url = "https://api.openai.com/v1"
name = "GPT-4o (Responses)"
api_backend = "responses"
env_key = "OPENAI_API_KEY"
Ollama (Local Models)
Run models locally with Ollama:
[model.ollama-codellama]
model = "codellama"
base_url = "http://localhost:11434/v1"
name = "CodeLlama (Ollama)"
Make sure Ollama is running (ollama serve) and the model is pulled (ollama pull codellama).
Together AI
[model.together-mixtral]
model = "mistralai/Mixtral-8x7B-Instruct-v0.1"
base_url = "https://api.together.xyz/v1"
name = "Mixtral 8x7B"
env_key = "TOGETHER_API_KEY"
Local OpenAI-Compatible Server
Any server that implements the OpenAI Chat Completions or Responses API:
[model.local-llama]
model = "llama-3.1-70b"
base_url = "http://localhost:8080/v1"
name = "Local Llama"
temperature = 0.8
Custom Models Endpoint
Point Grok at a custom OpenAI-compatible /v1/models endpoint instead of the default. Use this when your models sit behind a corporate gateway or a self-hosted inference service.
Environment Variables
| Variable | Required | Description |
|---|---|---|
GROK_MODELS_BASE_URL |
Yes | Base URL for inference. Grok fetches the model list from {base_url}/models. |
XAI_API_KEY |
Yes | API key sent as Authorization: Bearer. Grok also accepts GROK_CODE_XAI_API_KEY. |
GROK_MODELS_LIST_URL |
No | Override the model-list URL when it differs from {base_url}/models. |
Setup
export GROK_MODELS_BASE_URL="https://api.acme.com/v1"
export XAI_API_KEY="xai-..."
grok
Config File Alternative
[endpoints]
models_base_url = "https://api.acme.com/v1"
# Override only the API key for a specific model
[model.grok-build]
api_key = "my-api-key"
When you use [endpoints] with partial model overrides, Grok inherits the base_url from the endpoints config, so you do not need to specify it in each [model.*] section.
Auth Behavior
When you set models_base_url, Grok uses API key auth (Authorization: Bearer) instead of session auth. You do not need grok login -- the API key is enough.
Web Search Model
The web_search tool uses a separate model. Configure it with:
[models]
web_search = "grok-4.20-multi-agent"
Or via environment variable:
export GROK_WEB_SEARCH_MODEL="grok-4.20-multi-agent"
If you point web search at a custom model, you also need a [model.*] entry so Grok can reach it. Server-side ("backend") web search runs only when the model sets supports_backend_search = true (and the build enables backend search); it does not depend on api_backend:
[models]
web_search = "my-custom-model"
[model.my-custom-model]
model = "my-custom-model"
supports_backend_search = true
Using Custom Models
# List available models (including custom)
grok models
# Use in the TUI via slash command
/model my-model
# Use in headless mode
grok -p "Hello" -m my-model
# Set as default in config.toml:
[models]
default = "my-model"
Enterprise Deployment
A complete config for an enterprise deployment with custom models:
[cli]
auto_update = false
[auth]
auth_provider_command = "/usr/local/bin/my-company-auth-provider"
auth_provider_label = "Acme Corp"
auth_token_ttl = 3600
[models]
default = "company-grok"
[model.company-grok]
model = "grok-build"
base_url = "https://grok-proxy.acme.com/"
name = "Grok Build Latest (Proxy)"
context_window = 128000
[features]
telemetry = false
Troubleshooting
Model Not Found
# List available models
grok models
# Check config.toml for typos in [model.*] sections
Connection Errors
Verify the endpoint is reachable:
curl -s https://api.example.com/v1/models \
-H "Authorization: Bearer $XAI_API_KEY"
Debug Logging
RUST_LOG=debug GROK_LOG_FILE=/tmp/grok.log grok
tail -f /tmp/grok.log
Look for log entries containing model or sampling to trace model selection and API calls.