Skip to content

Agents > Inference & providers

Custom inference endpoint

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

Connect agents to any OpenAI-compatible inference endpoint — OpenRouter, LiteLLM, z.ai, or an internal gateway exposed at a public URL.

Warp supports custom inference endpoints for users who want to power agents with any OpenAI-compatible inference endpoint — a model router, hosted gateway, or internal infrastructure they already run.

Your endpoint must be reachable at a public URL. Requests route through Warp’s servers (see How it works), so Warp must be able to reach your endpoint over the public internet. localhost, private or internal network addresses, and internal-only services — such as a LiteLLM proxy that’s only reachable inside your network — are rejected. To use an internal or local endpoint, first expose it at a public HTTPS URL. See Network requirements for details.

  • OpenAI-compatible - Works with any endpoint that implements the OpenAI Chat Completions API.
  • Provider flexibility - Use a model router (OpenRouter, LiteLLM), a model provider with an OpenAI-compatible surface (z.ai), or your own internal gateway exposed at a public URL.
  • Provider-billed inference - Your endpoint provider bills inference directly. On Business and Enterprise, local runs that use an endpoint still incur platform charges.
  • Personal configuration - Your endpoint configuration is stored on your device for local Warp Agent runs. Business and Enterprise admins can instead configure team-managed endpoints for local and cloud runs.

Your endpoint must implement the OpenAI Chat Completions API (POST /v1/chat/completions). Compatible services include:

  • OpenRouter - Aggregates many model providers behind a single OpenAI-compatible API and consolidated billing.
  • LiteLLM - A self-hosted proxy that exposes a unified, OpenAI-compatible API across providers.
  • z.ai - A model provider with an OpenAI-compatible API surface for its models.
  • Internal gateways (exposed at a public URL) - An in-house service that fronts model providers behind an OpenAI-compatible endpoint (for example, a corporate AI gateway with logging, redaction, or access control). The gateway must be reachable from the public internet — an internal-only service, such as a LiteLLM proxy that only resolves inside your network or VPN, won’t work until it’s exposed at a public URL (see Network requirements).

Personal endpoint configurations are stored on your device, with API keys in secure storage. When you select an endpoint model, your prompt, context, endpoint URL, and key pass through Warp’s servers to your endpoint. Warp does not retain your personal endpoint API key.

To configure a personal endpoint in the Warp app:

  1. In Warp, open Settings and search for inference endpoint to jump to the configuration.
  2. Add your endpoint URL (the base URL that exposes /v1/chat/completions) and any required credentials (typically an API key).
  3. Specify the model identifier(s) you want to route through this endpoint.
  4. Save the configuration. Once added, you’ll see your custom models appear in the model picker.

When you select an endpoint-routed model, Warp routes inference through your endpoint instead of consuming Warp-provided inference usage.

Warp routes inference requests through its servers, so your endpoint must be reachable from the public internet. localhost, 127.0.0.1, and other private or local network URLs are rejected when configuring a custom inference endpoint.

This requirement applies to any endpoint that isn’t already publicly accessible:

  • Internal gateways and proxies - An internal LiteLLM proxy, corporate AI gateway, or other service that only resolves inside your private network or VPN can’t be reached by Warp. Expose it at a public HTTPS URL — for example, through a load balancer, an API gateway, or a tunneling service — before configuring it in Warp.
  • Local models - To route through a model running on your own machine (for example, Ollama, LM Studio, vLLM, or llama.cpp), expose it through a tunneling service like ngrok and use the public tunnel URL as the base URL in your endpoint configuration.

For example, with a default Ollama install listening on port 11434, run ngrok http 11434 and use the resulting https://*.ngrok-free.app/v1 URL as your endpoint. Other tunneling services that produce a publicly reachable HTTPS URL (Cloudflare Tunnel, Tailscale Funnel, and similar) work the same way.

Your endpoint provider bills endpoint-routed inference according to their pricing, rather than drawing from your Warp usage balance.

Local endpoint-routed runs on Business and Enterprise still incur platform charges. External inference bills are not fully represented in Warp’s usage totals.

Warp’s Auto models use Warp-provided inference and consume your available usage, even if you’ve configured an endpoint.

To use your endpoint, select the specific endpoint-routed model from the model picker rather than an Auto option.

Custom routers can’t use your endpoint either: routing targets must be Warp-supported models, so a router never resolves to an endpoint-routed model. Custom routers do apply BYOK provider keys after resolving a model.

Other features use their own inference configuration and are unaffected by a personal endpoint. See the feature breakdown on the BYOK page.

Warp is SOC 2 compliant and has Zero Data Retention (ZDR) agreements with all of its contracted LLM providers.

Custom inference endpoint prompts and responses transit Warp’s backend (see How it works). Warp does not use this content for training; retention and analytics handling follow the same account-level privacy and telemetry settings that apply to Warp-billed traffic.

When you use a custom inference endpoint:

  • Data retention on the provider side is determined by your endpoint provider and any upstream model providers they route to.
  • Warp cannot enforce ZDR for requests sent through a custom inference endpoint.
  • If your endpoint provider does not have ZDR with the underlying model provider, your requests may be retained according to their terms.

Review your endpoint provider’s data handling and retention policies before routing sensitive prompts through a custom inference endpoint.

Personal endpoints are configured on each user’s device and work for local Warp Agent runs, not cloud agents.

Business and Enterprise admins can configure team-managed endpoints in the Admin Panel for local and cloud runs. For Enterprise AWS Bedrock or Gemini Enterprise (Vertex AI) routing, see Bring Your Own LLM.

How custom inference endpoints differ from BYOK and BYOLLM

Section titled “How custom inference endpoints differ from BYOK and BYOLLM”

Choose a configuration based on who manages it and where your agents run:

NameMeaningPlans
Personal BYOKUse your own API key for OpenAI, Anthropic, or Google models in local Warp Agent runs. Keys are stored on your device.Free and all eligible paid plans
Personal custom inference endpointConnect local Warp Agent runs to an OpenAI-compatible endpoint such as OpenRouter, LiteLLM, z.ai, or an internal gateway.Free and all eligible paid plans
Team-managed API keys and endpointsAn admin configures shared provider keys or endpoints for local and cloud Warp Agent runs.Business and Enterprise
Bring Your Own LLM (BYOLLM)Route inference through your organization’s AWS Bedrock or Gemini Enterprise (Vertex AI) account. See the provider guides for cloud support.Enterprise only

Local runs on Business and Enterprise that use customer-supplied inference incur platform charges.