Self-Hosted Command Center

Deploy Agent Command Center on your own infrastructure via Docker or Go binary: full control over data residency, routing, failover, caching, and rate limiting.

About

Agent Command Center is distributed as a Go binary and Docker image. Self-hosting gives you full control over data residency, network topology, and configuration. All requests stay within your infrastructure.

Whether you’re running a single instance for development or scaling to production, Agent Command Center handles routing, failover, caching, and rate limiting across multiple LLM providers.

Note

A self-hosted Future AGI install already runs this gateway: inside the app container in Standalone and as the agentcc-gateway container in Distributed, both on port 8090, and as the <release>-agentcc-gateway Service in Helm. See Self-hosting. This page is for running the gateway on its own.

Requirements

  • Docker (for container deployment) or Go 1.25+ (to build from source)
  • Provider API keys for any cloud LLM providers you want to use

Quick start with Docker

Create a configuration file

Save this as config.yaml:

server:
  port: 8080

providers:
  openai:
    base_url: "https://api.openai.com"
    api_key: "${OPENAI_API_KEY}"
    api_format: "openai"
    models:
      - gpt-4o
      - gpt-4o-mini

auth:
  enabled: true
  keys:
    - name: "my-key"
      key: "sk-agentcc-my-key-here"
      key_type: "internal"

logging:
  level: info

Set your API key

export OPENAI_API_KEY="sk-..."

Run the container

docker run -d \
  -p 8080:8080 \
  -v $(pwd)/config.yaml:/app/config.yaml \
  -e OPENAI_API_KEY="$OPENAI_API_KEY" \
  --name agentcc-gateway \
  futureagi/agentcc-gateway:latest

Verify it's running

curl http://localhost:8080/healthz

Expected response: {"status":"ok"}

Note

Replace config.yaml with your actual configuration file. Environment variables referenced in the config (like ${OPENAI_API_KEY}) are resolved at runtime. The container runs as uid 65532, so config.yaml must be readable by that user (mode 0644).

key_type: "internal" lets the key use the providers in config.yaml. Without it, requests fail with 403: see Authentication.

Configuration file

Basic configuration

Here’s a minimal config for getting started with OpenAI:

server:
  port: 8080
  host: "0.0.0.0"

providers:
  openai:
    base_url: "https://api.openai.com"
    api_key: "${OPENAI_API_KEY}"
    api_format: "openai"
    models:
      - gpt-4o
      - gpt-4o-mini

auth:
  enabled: true
  keys:
    - name: "my-key"
      key: "sk-agentcc-my-key-here"
      key_type: "internal"

logging:
  level: info

Adding multiple providers

Combine OpenAI, Anthropic, and a self-hosted Ollama instance:

server:
  port: 8080

providers:
  openai:
    base_url: "https://api.openai.com"
    api_key: "${OPENAI_API_KEY}"
    api_format: "openai"
    models:
      - gpt-4o
      - gpt-4o-mini

  anthropic:
    base_url: "https://api.anthropic.com"
    api_key: "${ANTHROPIC_API_KEY}"
    api_format: "anthropic"
    models:
      - claude-sonnet-4-6

  ollama:
    base_url: "http://host.docker.internal:11434"
    api_format: "openai"

auth:
  enabled: true
  keys:
    - name: "my-key"
      key: "sk-agentcc-my-key-here"
      key_type: "internal"

logging:
  level: info

Tip

For Ollama, models are auto-discovered from the /v1/models endpoint when the gateway starts. You don’t need to list them, but restart the gateway after you pull a new model.

host.docker.internal is the Docker host as seen from the gateway’s container. Docker Desktop, Colima and OrbStack resolve it; on Linux, add --add-host=host.docker.internal:host-gateway to docker run. If you run the Go binary on the machine that runs Ollama, use http://localhost:11434. See Self-hosted models for more servers.

Enabling routing and failover

Add intelligent routing across multiple providers:

routing:
  default_strategy: "round-robin"
  failover:
    enabled: true
    max_attempts: 3
    on_status_codes: [429, 500, 502, 503, 504]
    on_timeout: true
  circuit_breaker:
    enabled: true
    failure_threshold: 5
    success_threshold: 2
    cooldown: 30s
  retry:
    enabled: true
    max_retries: 2
    initial_delay: 500ms
    max_delay: 10s
    multiplier: 2.0

This configuration:

  • Routes requests round-robin across providers
  • Fails over to the next provider on 429, 5xx errors, or timeouts
  • Opens circuit breaker after 5 consecutive failures
  • Automatically retries with exponential backoff

Enabling caching

Cache responses to reduce latency and API costs:

cache:
  enabled: true
  default_ttl: 5m
  max_entries: 10000

Warning

Caching is based on request content. Ensure your use case is compatible with cached responses (e.g., deterministic queries, not real-time data).

Rate limiting

Control request volume:

rate_limiting:
  enabled: true
  global_rpm: 1000

Set global_rpm: 0 for unlimited requests.

Authentication

Restrict access with API keys:

auth:
  enabled: true
  keys:
    - name: "dev-key"
      key: "sk-agentcc-dev-key-for-testing"
      key_type: "internal"
      owner: "dev-team"
      models:
        - gpt-4o
        - gpt-4o-mini

    - name: "prod-key"
      key: "sk-agentcc-prod-key-here"
      key_type: "internal"
      owner: "production"

The models field is optional. If omitted, the key can access all models.

Only keys with key_type: "internal" can use the providers in config.yaml. A key without key_type is a byok key, which can use only the providers an organization adds in a connected Future AGI app, so on a gateway you run on its own its requests fail with 403 model "<model>" is not available for this API key. Requests made with an internal key get no request.trace log line and are not sent to a connected Future AGI app, and the gateway does not enforce the key’s providers list. It still enforces models.

Server configuration reference

SettingDefaultDescription
server.port8080Port to listen on
server.host0.0.0.0Host to bind to
server.read_timeout5sRequest read timeout
server.write_timeout300sResponse write timeout
server.idle_timeout120sIdle connection timeout
server.shutdown_timeout30sHow long shutdown waits for in-flight requests. See Graceful shutdown.
server.max_request_body_size52428800Max request body (50MB). config.example.yaml and the Future AGI Helm chart set it to 10485760 (10MB).
server.default_request_timeout60sDefault timeout for provider requests

Provider configuration reference

Each provider in the providers: section supports:

SettingRequiredDescription
api_keyCloud onlyAPI key (can use ${ENV_VAR} syntax). Not needed for self-hosted providers like Ollama.
api_formatYesFormat: openai, anthropic, gemini, bedrock, cohere, azure
base_urlYesProvider endpoint, such as https://api.openai.com. The gateway refuses to start without it.
typeNoHas no effect in config.yaml, since base_url and api_format are required anyway. Leave it out.
modelsNoList of available models. If you leave it out, the gateway asks the server’s /v1/models endpoint when it starts, without credentials. That works for a server that checks no key, such as Ollama. List the models for cloud providers.
default_timeoutNoRequest timeout for this provider
max_concurrentNoMax concurrent requests
conn_pool_sizeNoConnection pool size

Connecting to a Future AGI app

A gateway connected to a Future AGI app (the control plane) loads API keys and organization settings from it, sends it request logs, and serves the providers that organizations add in the dashboard. API keys created in the dashboard can use only their organization’s providers, not providers: in config.yaml. The settings below apply only when the gateway is connected. Each has a config.yaml key and an environment variable; the variable wins when both are set.

SettingEnvironment variableDefaultDescription
control_plane.urlAGENTCC_CONTROL_PLANE_URLnoneURL of the Future AGI app
control_plane.admin_tokenAGENTCC_CONTROL_PLANE_TOKENThe gateway admin token (admin.token, AGENTCC_ADMIN_TOKEN)Token for the gateway’s calls to the app
control_plane.sync_on_startupAGENTCC_SYNC_ON_STARTUPfalseLoad keys and organization settings at startup. The gateway retries with backoff until the app answers, and serves only the keys in its config.yaml until then. Without it, keys created in the dashboard stop working when the gateway restarts.
control_plane.sync_intervalAGENTCC_SYNC_INTERVALUnset (off)Load them again this often, as a Go duration such as 60s or 5m. Unset or 0 turns periodic sync off. An AGENTCC_SYNC_INTERVAL that does not parse is ignored; in config.yaml, one stops the gateway from starting. Set it when you run more than one replica, so a key created on one reaches the others.
org_providers.allow_private_urlsAGENTCC_ALLOW_PRIVATE_PROVIDER_URLSfalseLet providers that organizations add use private and LAN addresses. It does not apply to providers: in config.yaml, which are trusted. See Private and local provider URLs.

The gateway logs control plane startup sync succeeded once the startup sync has loaded everything. While it keeps failing, it logs control plane not ready yet, retrying startup sync, and after two minutes a warning that starts with control plane sync still failing.

Health checks

Verify the gateway is running and ready:

curl http://localhost:8080/healthz
curl http://localhost:8080/readyz

/healthz returns {"status":"ok"} while the process is up. /readyz returns {"status":"ready"} while the gateway is serving, and 503 with {"status":"not_ready"} once it starts shutting down.

Connecting your application

Once running, point your application to the self-hosted gateway:

from agentcc import AgentCC

client = AgentCC(
    api_key="sk-agentcc-my-key-here",
    base_url="http://localhost:8080",
)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
import { AgentCC } from "@futureagi/agentcc";

const client = new AgentCC({
  apiKey: "sk-agentcc-my-key-here",
  baseUrl: "http://localhost:8080",
});

const response = await client.chat.completions.create({
  model: "gpt-4o",
  messages: [{ role: "user", content: "Hello!" }],
});

console.log(response.choices[0].message.content);
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer sk-agentcc-my-key-here" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Tip

For production, use a public endpoint (e.g., behind a reverse proxy with TLS). Replace http://localhost:8080 with your actual gateway URL.

Building from source

The gateway’s source is in the public future-agi repository, under agentcc-gateway/. Build it with Go 1.25 or later:

git clone https://github.com/future-agi/future-agi.git
cd future-agi/agentcc-gateway
go build -o agentcc-gateway ./cmd/agentcc
./agentcc-gateway --config config.yaml

agentcc-gateway/config.example.yaml documents every configuration field.

Environment variables

All values in config.yaml that use ${VAR_NAME} syntax are resolved from environment variables at startup. For example:

providers:
  openai:
    api_key: "${OPENAI_API_KEY}"

Set the variable before running:

export OPENAI_API_KEY="sk-..."
docker run -e OPENAI_API_KEY="$OPENAI_API_KEY" ...

Logging

Control verbosity with the logging.level setting:

logging:
  level: debug  # debug, info, warn, error

View logs from the container:

docker logs -f agentcc-gateway

Graceful shutdown

When the gateway gets SIGTERM or SIGINT, it:

  1. Stops taking new requests and lets in-flight requests finish, for up to server.shutdown_timeout (default 30s).
  2. Sends the request logs it still holds to the Future AGI app, when it is connected to one. This takes up to about 5 seconds.
  3. Exits.

If requests are still running when shutdown_timeout ends, the gateway logs shutdown error, still sends the buffered request logs, and exits with status 1. A second signal stops it at once.

Give the container at least shutdown_timeout plus 5 seconds to stop before it is killed: 35 seconds with the default. With OTLP export (otel.exporter: otlp), add up to 15 seconds for the exporter’s last flush. In Docker Compose, set stop_grace_period:

services:
  agentcc-gateway:
    stop_grace_period: 35s

Two log lines report what happened to the buffered request logs:

  • log flusher: delivered buffered request logs before shutdown, with count: all of them reached the app.
  • log flusher: request logs not delivered before shutdown, with undelivered: some did not. It is a warning when the app was already gone, for example when the connection was refused, and an error otherwise.

In a self-hosted Future AGI install:

  • Helm gives the gateway enough time by default: see Configure.
  • Distributed sets no stop_grace_period on agentcc-gateway, so Docker’s default stop timeout applies. To give it the time above, set stop_grace_period: 35s on agentcc-gateway in docker-compose.override.yml, and list the override in COMPOSE_FILE: see Move to managed data stores.
  • Standalone runs the gateway inside the app container, where it stops after the API. Request logs it still holds at that point cannot reach the app, and it logs log flusher: request logs not delivered before shutdown.

If you alert on the gateway’s log lines, see LLM gateway changes for the request-log lines that changed.

Next Steps

Was this page helpful?

Questions & Discussion