Self-Hosted Model Integration in Agent Command Center

Route requests to Ollama, vLLM, LM Studio, or any OpenAI-compatible inference server alongside cloud providers. All gateway features apply to local models.

About

Agent Command Center can route requests to models running on your own hardware alongside cloud providers. Self-hosted models are configured as providers with a base_url pointing to your local inference server. All gateway features (routing, caching, failover, guardrails) work the same way.

You can add a model server in two places: in the gateway’s config.yaml, as in the examples on this page, or as a provider your organization adds in the dashboard under Gateway > Providers. They serve different API keys:

  • providers: in config.yaml serve only the gateway’s internal keys: the platform’s own calls in a self-hosted Future AGI install, and keys in config.yaml with key_type: "internal" on a gateway you run on its own.
  • Providers added in the dashboard serve the API keys of that organization, including the keys created in the dashboard.

So in a self-hosted Future AGI install, add the model server under Gateway > Providers to call it from your code: see Reach a model server. The two places also treat private and local addresses differently: see Private and local provider URLs.


Supported inference servers

Serverapi_formatNotes
OllamaopenaiServes an OpenAI-compatible API under /v1
vLLMopenaiOpenAI-compatible server for production inference
LM StudioopenaiDesktop app with local server mode
Any OpenAI-compatible serveropenaiAny server that implements /v1/chat/completions

The gateway has no preset for these servers, so a type value such as ollama, vllm or lm_studio has no effect. Every provider in config.yaml needs base_url and api_format, or the gateway refuses to start with provider "<name>": base_url is required or provider "<name>": api_format is required.


Configuration

These examples go in the providers: section of the gateway’s config.yaml, for a gateway you run on its own (Self-hosted deployment). Only keys with key_type: "internal" can use them: see Authentication. In a self-hosted Future AGI install, config.yaml providers serve only the platform’s own calls (LLM Gateway), and API keys created in the dashboard cannot reach them.

The examples reach the model server at host.docker.internal, the Docker host as seen from inside a container. Docker Desktop, Colima and OrbStack resolve it; on Linux, see Reach a model server. Use localhost only when the gateway runs directly on the machine that runs the model server, not in a container: inside a container, localhost is the container itself.

Give base_url without /v1: the gateway adds /v1 to each call, such as /v1/chat/completions. A base_url that already ends in /v1 works too, since the gateway does not repeat it.

Ollama

providers:
  ollama:
    base_url: "http://host.docker.internal:11434"
    api_format: "openai"
    # No models list: the gateway discovers them from /v1/models when it starts

Start Ollama with OLLAMA_HOST=0.0.0.0 so it listens on an address containers can reach, and read the caution in Reach a model server first. The gateway lists Ollama’s models when it starts, so after you pull a new model (ollama pull llama3.1), restart the gateway to pick it up.

vLLM

providers:
  vllm:
    base_url: "http://gpu-server:8000"
    api_format: "openai"
    models:
      - "meta-llama/Llama-3.1-70B-Instruct"

LM Studio

providers:
  lm-studio:
    base_url: "http://host.docker.internal:1234"
    api_format: "openai"

As with Ollama, LM Studio’s server must listen on an address the gateway’s container can reach.

Generic OpenAI-compatible server

Any server that implements the /v1/chat/completions endpoint:

providers:
  my-server:
    base_url: "http://inference.internal:8080"
    api_format: "openai"
    models:
      - "my-custom-model"

Private and local provider URLs

The two ways to add a model server follow different URL rules:

Where you add itWho adds itPrivate and local addresses
providers: in the gateway’s config.yamlThe operator of the gatewayAllowed. These entries are trusted and have no address check.
Gateway > Providers in the dashboard (preset Custom / Self-hosted), or the org config APIMembers of an organizationFollow the policy below

The policy for providers that organizations add:

AddressExamplesDefaultWith AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS=true
PublicAny address not listed belowAllowedAllowed
Private10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10 (CGNAT, such as Tailscale), fc00::/7, and names that resolve there: Docker service names, host.docker.internal, LAN host namesRefusedAllowed
Loopback127.0.0.0/8, ::1, localhostRefusedRefused
Link-local, cloud metadata, multicast and unspecified169.254.0.0/16, fe80::/10, 224.0.0.0/4, 0.0.0.0/8, and metadata host names such as metadata.google.internalRefusedRefused

The gateway checks the address when it sets up the provider and again on every connection, after DNS resolution, so a host name that later resolves to a refused address is refused too. A host name with several addresses must pass on all of them. When you save a provider, the Future AGI API runs a similar check, and it also refuses the ports 6379, 5432, 3306, 27017, 9200, 11211 and 2379.

Releases from before this setting let 100.64.0.0/10 addresses through. After you upgrade, a provider on a Tailscale address needs the setting below.

Allow private addresses

Turn the setting on for both the Future AGI API, which checks the URL when you save a provider and when it fetches the provider’s model list, and the gateway, which checks every request.

In Standalone and Distributed, one line in .env reaches both the API and the gateway:

echo "AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS=true" >> .env
docker compose up -d

Set agentccGateway.allowPrivateProviderURLs to true. The chart passes it to the gateway and the backend.

agentccGateway:
  allowPrivateProviderURLs: true

In the gateway’s config.yaml:

org_providers:
  allow_private_urls: true

Or set AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS=true in the gateway’s environment. Set the same variable on the Future AGI API the gateway syncs from.

A gateway or API from before this setting ignores it and refuses every private URL. While the setting is on, the gateway logs a warning at startup.

Warning

The setting is off by default because on a shared gateway it lets any organization reach internal hosts. Once it is on, anyone who can add a provider can point the gateway and the API at any address on your private networks, databases included. Turn it on only where everyone who can add a provider is trusted.

Installs made before Standalone became the default may run ClickHouse with no password (CH_PASSWORD empty), and neither check refuses ClickHouse’s ports (8123, 9000). Give ClickHouse a password before you turn the setting on: see Secrets the installer generates.

Reach a model server

The addresses below are all private, so a provider that uses one needs the setting above. In Gateway > Providers, click Add Provider and select Custom / Self-hosted. Enter the server’s OpenAI-compatible base URL, such as http://host.docker.internal:11434/v1. With the default API Path Prefix of /v1, the address without /v1 works too.

The dialog requires an API key. For a server that checks no key, such as Ollama or a vLLM server started without one, enter any placeholder, such as ollama. The dialog then fetches the server’s models: select the ones to serve, and click Add Provider.

On the Docker host. Use http://host.docker.internal:11434/v1 for Ollama, started with OLLAMA_HOST=0.0.0.0, or http://host.docker.internal:8000/v1 for vLLM. OLLAMA_HOST=0.0.0.0 makes Ollama listen on every network interface of the host, and Ollama does not authenticate requests, so block port 11434 from other machines with a firewall unless the host is only on a network you trust. Docker Desktop, Colima and OrbStack resolve host.docker.internal. On Linux, add it in docker-compose.override.yml, to the app service in Standalone:

services:
  app:
    extra_hosts:
      - "host.docker.internal:host-gateway"

In Distributed, add the same extra_hosts entry to agentcc-gateway and backend instead of app.

As a Compose service. Add the model server to docker-compose.override.yml and use its service name, such as http://ollama:11434/v1. Single-label host names such as ollama are accepted.

On Distributed, Compose reads docker-compose.override.yml only when COMPOSE_FILE in .env lists it: see Move to managed data stores. Otherwise the entries above have no effect.

Elsewhere on your network. A LAN host, a Tailscale address or a Kubernetes Service name works the same way once the setting is on.

Do not use localhost or 127.0.0.1. From the gateway, that address is the gateway itself, so it is never allowed.

Errors

  • When you save the provider, the API answers 400 with Base URL points to a private network address, which is refused by default. A self-hosted deployment can allow private/LAN provider URLs by setting AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS=true on the backend and the gateway. For a loopback, link-local or cloud metadata address, it answers Invalid base URL: it must be an http(s) URL, and loopback, link-local and cloud metadata addresses are never allowed.
  • On a request, the gateway answers 403 with code provider_base_url_blocked when the address is refused, and 502 with code provider_base_url_unusable when the base URL is not a valid http(s) URL or its host does not resolve from the gateway. The message names the model, the provider and the reason, for example: model "llama3.1" is served by your organization's provider "ollama", but the gateway will not call it: its base_url points to a private network address, which this gateway refuses by default. A self-hosted deployment can allow private/LAN provider URLs by setting AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS=true on the gateway and the backend.

Note

Custom models in Settings > AI Providers are a different feature, and AGENTCC_ALLOW_PRIVATE_PROVIDER_URLS does not apply to them. Future AGI calls a custom model’s API Base URL from its own containers, so use host.docker.internal for a server on the Docker host (on Linux, with the extra_hosts entry above) or a Compose service name, not localhost. See AI Providers.


Hybrid routing

The main value of self-hosted models through Agent Command Center is hybrid routing: use cheap local models for simple requests and fall back to cloud providers for complex ones.

Cost-based routing

The cost-optimized strategy chooses among the providers that serve the requested model. It does not read prices: it picks the target with the lowest priority in routing.targets. When your own server and a cloud provider serve the same model, give your server the lower value. It then takes every request while it is up, and failover moves requests to the cloud provider when it fails:

routing:
  default_strategy: "cost-optimized"
  failover:
    enabled: true
  targets:
    "meta-llama/Llama-3.1-70B-Instruct":
      - provider: vllm
        priority: 0
      - provider: cloud-llama
        priority: 1

providers:
  vllm:
    base_url: "http://gpu-server:8000"
    api_format: "openai"
    models: ["meta-llama/Llama-3.1-70B-Instruct"]

  cloud-llama:
    base_url: "https://llm.example.com"
    api_key: "${CLOUD_LLAMA_API_KEY}"
    api_format: "openai"
    models: ["meta-llama/Llama-3.1-70B-Instruct"]

Failover from local to cloud

Failover only moves a request between providers of the same model. To send a request for a local model to a different cloud model when the local server fails, list the cloud model in routing.model_fallbacks:

routing:
  default_strategy: "round-robin"
  failover:
    enabled: true
    on_status_codes: [429, 500, 502, 503, 504]
    on_timeout: true
  model_fallbacks:
    llama3.1:
      - gpt-4o

providers:
  ollama:
    base_url: "http://host.docker.internal:11434"
    api_format: "openai"

  openai:
    base_url: "https://api.openai.com"
    api_key: "${OPENAI_API_KEY}"
    api_format: "openai"
    models: ["gpt-4o"]

When a llama3.1 request to Ollama fails with one of on_status_codes or times out, the gateway sends it again as gpt-4o to OpenAI. An Ollama that is down counts as a 502. The fallback needs failover.enabled and a default_strategy: without them, the gateway returns Ollama’s error.

Complexity-based routing

Route simple queries to a local model and complex queries to a cloud model:

routing:
  complexity:
    enabled: true
    tiers:
      simple:
        max_score: 30
        model: "llama3.1"
        provider: "ollama"
      complex:
        max_score: 100
        model: "gpt-4o"
        provider: "openai"

See Routing > Complexity-based routing for the full scoring system.


Using self-hosted models from code

Once configured, self-hosted models are used the same way as cloud models.

In a self-hosted Future AGI install, first add the model server under Gateway > Providers (Reach a model server), then call it with an API key created in the dashboard. The gateway listens on port 8090 (AGENTCC_GATEWAY_PORT):

from agentcc import AgentCC

client = AgentCC(
    api_key="sk-agentcc-your-key",  # an API key created in the dashboard
    base_url="http://localhost:8090",  # the gateway in a self-hosted Future AGI install
)

# llama3.1 must be one of the models selected for the provider
response = client.chat.completions.create(
    model="llama3.1",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

A gateway container you run on its own, as in Self-hosted deployment, listens on 8080 and serves the config.yaml providers above to keys with key_type: "internal".


Limitations

  • Self-hosted models don’t support the Assistants API (threads are stored on OpenAI’s servers)
  • Embedding endpoints require the inference server to implement /v1/embeddings
  • With cost_tracking.enabled: true, the gateway prices each request from its model database, which has no prices for self-hosted models, so it reports no cost for them. A provider has no pricing field. To price a model, add it under model_database.overrides with pricing.input_per_token and pricing.output_per_token (USD per token). The gateway then also checks requests against the entry’s capabilities, so set function_calling, vision or response_schema to true for what the model supports, or requests that use them are refused.

Next Steps

Was this page helpful?

Questions & Discussion