> For the complete documentation index, see [llms.txt](https://docs.warp.dev/llms.txt).
> Markdown versions of each page are available by appending .md to any URL.

# Custom inference endpoint

Connect agents to any OpenAI-compatible inference endpoint — OpenRouter, LiteLLM, z.ai, or an internal gateway exposed at a public URL.

Warp supports **custom inference endpoints** for users who want to power agents with any OpenAI-compatible inference endpoint — a model router, hosted gateway, or internal infrastructure they already run.

**Your endpoint must be reachable at a public URL.** Requests route through Warp’s servers (see [How it works](#how-it-works)), so Warp must be able to reach your endpoint over the public internet. `localhost`, private or internal network addresses, and internal-only services — such as a LiteLLM proxy that’s only reachable inside your network — are rejected. To use an internal or local endpoint, first expose it at a public HTTPS URL. See [Network requirements](#network-requirements) for details.

Note

Custom inference endpoints are available on Free and all eligible paid plans for individual users and organizations with 10 or fewer employees, subject to Warp’s [Terms of Service](https://www.warp.dev/legal/terms-of-service). Larger organizations need a Business or Enterprise plan. See [Warp pricing](https://www.warp.dev/pricing) for current availability.

## Key features

-   **OpenAI-compatible** - Works with any endpoint that implements the OpenAI Chat Completions API.
-   **Provider flexibility** - Use a model router (OpenRouter, LiteLLM), a model provider with an OpenAI-compatible surface (z.ai), or your own internal gateway exposed at a public URL.
-   **Provider-billed inference** - Your endpoint provider bills inference directly. On Business and Enterprise, local runs that use an endpoint still incur [platform charges](https://docs.warp.dev/support-and-community/plans-and-billing/platform-credits/).
-   **Personal configuration** - Your endpoint configuration is stored on your device for local Warp Agent runs. Business and Enterprise admins can instead configure [team-managed endpoints](https://docs.warp.dev/enterprise/enterprise-features/team-managed-keys-and-endpoints/) for local and cloud runs.

## How it works

Your endpoint must implement the **OpenAI Chat Completions API** (`POST /v1/chat/completions`). Compatible services include:

-   **OpenRouter** - Aggregates many model providers behind a single OpenAI-compatible API and consolidated billing.
-   **LiteLLM** - A self-hosted proxy that exposes a unified, OpenAI-compatible API across providers.
-   **z.ai** - A model provider with an OpenAI-compatible API surface for its models.
-   **Internal gateways (exposed at a public URL)** - An in-house service that fronts model providers behind an OpenAI-compatible endpoint (for example, a corporate AI gateway with logging, redaction, or access control). The gateway must be reachable from the public internet — an internal-only service, such as a LiteLLM proxy that only resolves inside your network or VPN, won’t work until it’s exposed at a public URL (see [Network requirements](#network-requirements)).

Personal endpoint configurations are stored on your device, with API keys in secure storage. When you select an endpoint model, your prompt, context, endpoint URL, and key pass through Warp’s servers to your endpoint. Warp does not retain your personal endpoint API key.

Caution

Personal endpoint configurations are not available to [cloud runs](https://docs.warp.dev/platform/). Business and Enterprise teams can configure [team-managed endpoints](https://docs.warp.dev/enterprise/enterprise-features/team-managed-keys-and-endpoints/) for cloud inference; applicable compute and platform charges remain.

## Enabling a custom inference endpoint

To configure a personal endpoint in the Warp app:

1.  In Warp, open **Settings** and search for `inference endpoint` to jump to the configuration.
2.  Add your endpoint URL (the base URL that exposes `/v1/chat/completions`) and any required credentials (typically an API key).
3.  Specify the model identifier(s) you want to route through this endpoint.
4.  Save the configuration. Once added, you’ll see your custom models appear in the model picker.

When you select an endpoint-routed model, Warp routes inference through your endpoint instead of consuming Warp-provided inference usage.

## Network requirements

Warp routes inference requests through its servers, so **your endpoint must be reachable from the public internet**. `localhost`, `127.0.0.1`, and other private or local network URLs are rejected when configuring a custom inference endpoint.

This requirement applies to any endpoint that isn’t already publicly accessible:

-   **Internal gateways and proxies** - An internal LiteLLM proxy, corporate AI gateway, or other service that only resolves inside your private network or VPN can’t be reached by Warp. Expose it at a public HTTPS URL — for example, through a load balancer, an API gateway, or a tunneling service — before configuring it in Warp.
-   **Local models** - To route through a model running on your own machine (for example, Ollama, LM Studio, vLLM, or llama.cpp), expose it through a tunneling service like [ngrok](https://ngrok.com/) and use the public tunnel URL as the base URL in your endpoint configuration.

For example, with a default Ollama install listening on port `11434`, run `ngrok http 11434` and use the resulting `https://*.ngrok-free.app/v1` URL as your endpoint. Other tunneling services that produce a publicly reachable HTTPS URL (Cloudflare Tunnel, Tailscale Funnel, and similar) work the same way.

## Billing behavior

### Inference usage

Your endpoint provider bills endpoint-routed inference according to their pricing, rather than drawing from your Warp usage balance.

Local endpoint-routed runs on Business and Enterprise still incur [platform charges](https://docs.warp.dev/support-and-community/plans-and-billing/platform-credits/). External inference bills are not fully represented in Warp’s usage totals.

### Auto routing uses Warp-provided inference

Warp’s **Auto** models use Warp-provided inference and consume your available usage, even if you’ve configured an endpoint.

To use your endpoint, select the specific endpoint-routed model from the model picker rather than an Auto option.

[Custom routers](https://docs.warp.dev/agents/inference/custom-routers/) can’t use your endpoint either: routing targets must be Warp-supported models, so a router never resolves to an endpoint-routed model. Custom routers do apply [BYOK](https://docs.warp.dev/agents/inference/bring-your-own-api-key/) provider keys after resolving a model.

### Other AI features in Warp

Other features use their own inference configuration and are unaffected by a personal endpoint. See the [feature breakdown on the BYOK page](https://docs.warp.dev/agents/inference/bring-your-own-api-key/#byok-usage-and-billing-behavior).

## Zero Data Retention (ZDR)

Warp is **SOC 2 compliant** and has **Zero Data Retention (ZDR)** agreements with all of its contracted LLM providers.

Custom inference endpoint prompts and responses transit Warp’s backend (see [How it works](#how-it-works)). Warp does not use this content for training; retention and analytics handling follow the same account-level privacy and telemetry settings that apply to Warp-billed traffic.

When you use a custom inference endpoint:

-   Data retention on the **provider side** is determined by your endpoint provider and any upstream model providers they route to.
-   Warp **cannot enforce ZDR** for requests sent through a custom inference endpoint.
-   If your endpoint provider does not have ZDR with the underlying model provider, your requests may be retained according to their terms.

Review your endpoint provider’s data handling and retention policies before routing sensitive prompts through a custom inference endpoint.

## Centrally managed configuration

Personal endpoints are configured on each user’s device and work for local Warp Agent runs, not [cloud agents](https://docs.warp.dev/platform/).

Business and Enterprise admins can configure [team-managed endpoints](https://docs.warp.dev/enterprise/enterprise-features/team-managed-keys-and-endpoints/) in the [Admin Panel](https://docs.warp.dev/enterprise/team-management/admin-panel/) for local and cloud runs. For Enterprise AWS Bedrock or Gemini Enterprise (Vertex AI) routing, see [Bring Your Own LLM](https://docs.warp.dev/enterprise/enterprise-features/bring-your-own-llm/).

## How custom inference endpoints differ from BYOK and BYOLLM

Choose a configuration based on who manages it and where your agents run:

| Name | Meaning | Plans |
| --- | --- | --- |
| **[Personal BYOK](https://docs.warp.dev/agents/inference/bring-your-own-api-key/)** | Use your own API key for OpenAI, Anthropic, or Google models in local Warp Agent runs. Keys are stored on your device. | Free and all eligible paid plans |
| **Personal custom inference endpoint** | Connect local Warp Agent runs to an OpenAI-compatible endpoint such as OpenRouter, LiteLLM, z.ai, or an internal gateway. | Free and all eligible paid plans |
| **[Team-managed API keys and endpoints](https://docs.warp.dev/enterprise/enterprise-features/team-managed-keys-and-endpoints/)** | An admin configures shared provider keys or endpoints for local and cloud Warp Agent runs. | Business and Enterprise |
| **[Bring Your Own LLM](https://docs.warp.dev/enterprise/enterprise-features/bring-your-own-llm/)** (BYOLLM) | Route inference through your organization’s AWS Bedrock or Gemini Enterprise (Vertex AI) account. See the provider guides for cloud support. | Enterprise only |

Local runs on Business and Enterprise that use customer-supplied inference incur [platform charges](https://docs.warp.dev/support-and-community/plans-and-billing/platform-credits/).

## Related pages

-   **[Bring Your Own API Key](https://docs.warp.dev/agents/inference/bring-your-own-api-key/)** - Use your own OpenAI, Anthropic, or Google API keys.
-   **[Bring Your Own LLM](https://docs.warp.dev/enterprise/enterprise-features/bring-your-own-llm/)** - Enterprise-managed inference through your cloud provider or approved infrastructure.
-   **[Model choice](https://docs.warp.dev/agents/inference/model-choice/)** - Full list of supported models and `model_id` values.
-   **[Usage and billing](https://docs.warp.dev/support-and-community/plans-and-billing/credits/)** - Inference, compute, and platform charges.
