live atapi.inference.club

A distributed inference network powered by consumer GPUs and Tailscale

inference.club merges local AI with the cloud — a tailnet that puts your computer on the world's largest network of distributed consumer AI hardware. After all, the cloud is just someone else's computer.

How it fits together

One tailnet, every modality.

inference.club is a Tailscale tailnet that joins consumer hardware — RTX PCs, the DGX Spark, Apple silicon — so members can safely expose their inference through one unified API, across the whole range of AI modalities: chat, images, video, speech, music, 3D.

For consumers

Drop-in for the OpenAI SDK.

Sign up, mint a token, point your client at api.inference.club/v1. One simple API — a superset of OpenAI, NVIDIA NIM, and other popular standards — serving the best and most popular open models for every major modality.

export OPENAI_API_KEY=ic_xxxxxxxxxxxxxxxxxxxx
export OPENAI_BASE_URL=https://api.inference.club/v1

curl $OPENAI_BASE_URL/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.6-27b",
    "messages": [
      {"role": "user", "content": "explain MoE in one sentence"}
    ]
  }'
For providers

Wrap any OpenAI-compatible server.

Exposing local inference usually means ngrok, Cloudflare tunnels, or a custom tangle of ports and firewall rules. Here it's one simple container: run the agent next to vllm, llama.cpp, or ollama and it routes requests straight to your local services.

# Already running vLLM, llama.cpp, or Ollama on your GPU?
# Point the agent at it and join the network.

export INFERENCE_CLUB_API_KEY=ic_xxxxxxxxxxxxxxxxxxxx
export OPENAI_BASE_URL=http://localhost:8000/v1
export OPENAI_API_KEY=local-key  # whatever your local server expects

docker run --rm -d --name club-agent --network host \
  -e INFERENCE_CLUB_API_KEY \
  -e OPENAI_BASE_URL \
  -e OPENAI_API_KEY \
  ghcr.io/inference-club/inference-club-agent:latest

Architecture

Three pieces. Nothing magic.

Follow a single request from your code to a GPU and back. The cloud control plane authenticates, applies your privacy rules, and routes — but the model itself runs on hardware you own.

step01

Operators run agents

Members run inference-club-agent next to their local LLM server. The agent advertises whatever models the server is hosting.

step02

Agents join the tailnet

Each agent receives a short-lived Tailscale key and joins our private mesh. No ngrok, no Cloudflare configs, no ports or firewall holes. Just WireGuard.

step03

Consumers send requests

Calls to api.inference.club route to an online agent serving the requested model. Streaming works. Latency is direct.

Your application

curl · OpenAI SDK · the Playground · your agents

base_url = "https://api.inference.club/v1"
api_key = "ic_xxxxxxxxxxxxxxxxxxxx"
HTTPS · Authorization: Bearer ic_…

api.inference.club — the control plane

one small cloud VPS (Hetzner). It routes; it never runs the model.

Caddy

TLS · reverse proxy

Django + DRF

OpenAI-compatible /v1 router · auth · routing

Access control

visibility · per-service ACLs · kill switch

Celery workers

async jobs · batches · workflow DAG

Postgres + Redis

state · queue · throttling

GCS

images · video · voice · music

The inference.club tailnet

a private Tailscale mesh — pure WireGuard

SOCKS5 sidecarMagicDNS · club-host-17:443short-lived auth keysno ports · no firewall holes

Your rig — where inference actually happens

a GPU you own, at home, on hardware you trust

inference-club-agentcontainer · --network host

Joins the tailnet with its minted key, advertises models from agent.yaml, and forwards each request to whatever you already run locally:

vLLMllama.cppOllamaLM StudioDiaLTX-2

→ http://localhost:1234/v1

running onbrian's 4090M3 Ultra · 192GBDGX Spark2× 3090 rigk3s home cluster

Follow one request

  1. 1Your code calls api.inference.club/v1 with your ic_ key — the same request you’d send OpenAI.
  2. 2Caddy terminates TLS; Django authenticates the key and applies your privacy + access rules.
  3. 3The router picks a healthy, online node that actually serves the requested model.
  4. 4Django (via a Tailscale SOCKS5 sidecar) dials the node by MagicDNS over WireGuard — no ports, no tunnels.
  5. 5The agent container hands the request to your local LLM server on localhost.
  6. 6Tokens (or images / video / audio) stream back along the exact same path.

In one breath

inference.club is what happens when you point an OpenAI-compatible API at a pile of consumer GPUs you actually own and trust — a private Tailscale tailnet quietly stitching a 4090 here, an M3 Ultra there, a DGX Spark and a couple of 3090s into one WireGuard mesh with no ports forwarded and no firewall holes, where a little inference-club-agent container sits next to whatever you’re already running — vLLM, llama.cpp, Ollama, LM Studio — and advertises it through a manifest, while back in the cloud a Django + Celery server behind Caddy authenticates your ic_ key, enforces your privacy and per-service access controls, and routes the call over the tailnet by MagicDNS to a healthy online node, with Redis and Postgres driving async jobs, batches and a whole workflow DAG engine, GCS holding the images, video, voice and music that come back, a Nuxt playground and dashboard to poke at all of it, the home fleet itself migrating from Docker to k3s, and the entire thing — chat, images, LTX-2 video, Dia voice cloning, speech, the works — sitting behind one base URL you can curl, so go ahead and build something, and, as the prompt says: Make no mistakes.

Why inference.club

The cloud is just someone else's computer — own yours.

OpenAI-compatible

A superset of the APIs you already use — OpenAI, NVIDIA NIM, and friends. Swap the base URL and key; your existing SDKs and prompts just work.

Real GPUs, real models

Members serve the best open-weight models on hardware they own — Qwen, Llama, DeepSeek, LTX — for chat, images, video, speech, and more.

Private by default

Requests reach providers over Tailscale, end-to-end encrypted. No public endpoints to scrape.

A club, not a vendor

Connect with passionate local-AI enthusiasts and evangelists. Pool compute with people you trust, and showcase the best results consumer hardware can produce.

Ready to plug in?

Sign in, mint a key, and you're live in under a minute. Bring a node whenever you have spare cycles — your hardware is the cloud now.