Any GPU. Any cloud. One inference service you run yourself.
Opod installs into your own Kubernetes — cloud, VPC, on-prem, or air-gapped — and turns your GPU machines into governed, autoscaling, OpenAI-compatible endpoints. Serving runs on NVIDIA today; AMD is measured and packed by the same ledger, with its worker image next. Prompts never leave your cluster.
Free tier, no key, no sign-up · open-source core (Apache-2.0) · or a two-week read-only assessment of your bill.
- Fleet time
- 14:00
- GPUs powered
- —
- Endpoints
- 4
- Utilization
- —
Each lit square is a worker on a GPU. Overnight, workers scale to zero and GPUs are released; the four leaders stay up, and the first morning request wakes everything within its SLO.
- ≤ 2 msadded by the gateway at p99
- 127 spreflight to first token on a real GPU node
- 0inbound vendor connections
The problem
Inference burns money three ways.
GPUs sit at 5–30 % utilization while frontier-API token bills triple as agents replace chat. The engines that fix this are excellent and free — the management layer above them is the gap.
1 · GPUs
5–30 %
typical enterprise utilization. Most of an always-on GPU bill is paid-for idle — $3–8K per GPU per month.
2 · Tokens
50–500×
more tokens per task from agents than chat. Frontier APIs charge $12–30 per million output tokens; open models served well cost $0.04–1.
3 · DIY
$1M+/yr
to build this layer yourself, over 6–18 months — and it will never be your product.
What Opod is
One control plane. Any cloud, GPU, engine, model.
Opod composes the best open engines instead of forking them, runs on your Kubernetes (or brings its own k3s), and gives every team self-serve endpoints inside quota — metered per request.
Everything stays in your cloud.
Zero inbound vendor connections. Prompts and completions never leave your cluster. Air-gapped installs are first-class.
How it works
The endpoint is the object. The money path waits for no one.
An endpoint is a stable URL with keys and policy, served by one leader and the workers that join it — each worker one engine instance on a GPU share. The control plane places everything and stores everything of record, but it is never on the request path.
- Static stability. Lose the control plane and requests keep flowing; usage spools locally and streams back later.
- One product API. Dashboard, CLI, CRDs and Terraform all drive the same versioned plan per endpoint.
- Everything is an event. One append-only log; every store is a projection you can replay.
- Trust is proven. A node-agent gates every machine before any workload lands.
Requests served
0
Usage spooled on leaders
0
Status
Healthy. Control plane pushes plans; leaders route; usage streams back.
Worker
one engine instance inside its VRAM budget; many per GPU; disposable, rebuilt from heartbeats.Leader
the endpoint's front door — auth, queues, streaming, one usage record per request. Failover is a ~10 s restart.Control plane
ledger, placement, rollouts, tenancy, usage, console. One per org, many cells. Never on the request path.Capacity ledger · placement · autoscale
Utilization is the product.
A deterministic placement library with a per-GPU VRAM ledger — fleet state and workloads in, assignments out. Replay it against a fleet before you buy it.
Every deployment is a reservation.
Weights + KV budget + overhead + headroom, per GPU. The dashboard shows reserved, free, and what still fits.
Honest bin-packing.
Dedicated means a whole GPU and a hard SLA; shared means packed. We never claim isolation we can't enforce.
Priority and preemption.
critical › standard › preemptible — nightly batch soaks idle GPUs and is evicted in seconds when morning traffic arrives.
Three autoscale tiers.
Sleep, pod-zero with cached weights, node-zero — each endpoint picks its wake SLO. Sharded models scale as whole gangs.
Target: ≥ 2× your utilization baseline within 30 days, measured by the GPU-hours-saved report per model and team.
Where the biggest number comes from
Pack the GPU.
The field's rule is one model per card. Opod's ledger knows each model's footprint — weights plus a KV-cache reserve — and each worker's budget, so it places several models on one card and leaves the rest free or asleep.
Drag the slider to change how many models you want resident. Footprints are illustrative sizes for common open models, listed in the table under the chart.
One model per card versus first-fit packing against a per-worker budget, with 8 GB of headroom kept on every card.
- Cards, one per model
- —
- Cards, packed by the ledger
- —
- VRAM idle, one per model
- —
- VRAM idle, packed
- —
- Cards freed for sleep or batch
- —
Model footprints and the current placement
Footprint = weights at the listed precision + a KV-cache reserve sized for typical agent context. Packing is first-fit decreasing on a single node class. Real placement also respects engine compatibility, vendor and gang rules, so the real number is a little lower than this model. The point is the ratio, not the decimals.
Verified rollouts
Every deploy step produces a proof.
Not "the pod is Running" — a machine-checkable proof at every step, with host-level GPU evidence. On failure: diagnostics captured before rollback, plus a suggested remedy. The same saga upgrades the platform itself.
- 1
Preflight
placement validated against the ledger · capacity present · node gate passed on every target
proof: pending
- 2
Images
digest-pinned, cosign-verified, SBOM published, present in the cell mirror — built in CI, never on nodes
proof: pending
- 3
Launch
plan ConfigMap written · leader Deployment and worker pods scheduled on the reserved GPU shares
proof: pending
- 4
Place
weights resident from node NVMe cache · every worker heartbeats loaded with measured VRAM
proof: pending
- 5
Health
leader readiness · a real completion round-trips through the gateway
proof: pending
- 6
GPU proof
host-level utilization evidence from the node-agent, not the container
proof: pending
- 7
Resilience
kill one worker pod · the leader heals within its plan · zero dropped requests
proof: pending
For developers
Ship on it in an afternoon. Keep your SDK.
Create an endpoint from the console or one API call and point the OpenAI SDK you already use at it. Every change rolls out with proof; roll back with one call.
- ✓OpenAI-compatible /v1/chat/completions and /v1/models, streaming; embeddings endpoints in the next chart.
- ✓Models from the catalog or any Hugging Face GGUF / safetensors repo — images pinned by digest.
- ✓Per-key RPM/TPM limits enforced by the leader, guardrail webhooks, a fallback target.
- ✓One REST API behind the console — POST /api/v1/endpoints/preview dry-runs a create.
Opod core, Apache-2.0. No Kubernetes needed: opod up on one box, opod join from the rest, one OpenAI endpoint. Point Cursor, Aider or Codex CLI at it. Quickstart →
# pip install openai — nothing else changes from openai import OpenAI client = OpenAI( base_url="http://opod-ep-chat.opod.svc:8080/v1", # your endpoint, in your cluster api_key="sk-orc-…", # minted per person, with limits ) stream = client.chat.completions.create( model="qwen2.5-0.5b-gguf", messages=[{"role": "user", "content": "Summarise this contract."}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="")
# create an endpoint — leave the node out and the placer picks a GPU with room curl -X POST $OPOD/api/v1/endpoints \ -H "Authorization: Bearer $ADMIN_TOKEN" -H "Content-Type: application/json" \ -d '{"name":"chat","model":"qwen2.5-0.5b-gguf", "plan":{"model":{"id":"qwen2.5-0.5b-gguf","engine":"llamacpp"}, "colocation":"shared", "workers":[{"vramBudgetGb":6,"engine":"llamacpp"}], "autoscale":{"floor":0,"max":3,"idleSeconds":600}}}' # "serving" means the leader produced a token, not that a pod is Ready curl -H "Authorization: Bearer $ADMIN_TOKEN" $OPOD/api/v1/endpoints/$ID/status → {"phase":"serving","proof":"first token in 40 ms", …}
# open-source core, no Kubernetes: one box + friends $ opod up # leader + OpenAI gateway on :8080 $ opod token create --node # join token $ opod join "http://leader:8080?token=…" # from every other machine $ opod shard create <model> 2 --nodes a,b # split a model too big for one box $ opod connect cursor # copy-paste config for your tool
# Path A — you already have Kubernetes (EKS / GKE / AKS / k3s ≥ 1.28) $ helm install opod oci://ghcr.io/opod-io/charts/opod --version 0.1.0 \ -n opod --create-namespace --set admin.token=$ADMIN_TOKEN ✔ preflight Job: k8s version, nodes Ready, StorageClass, CoreDNS ✔ control plane + console up · no databases installed # Path B — bare Linux machines, no Kubernetes yet: we bring k3s over SSH $ opodcp infra plan -f machines.yaml # host facts + gate, changes nothing $ opodcp infra up -f machines.yaml # k3s server + agents, GPU vendor labels
Quickstart · free, no sign-up
From nothing to a first token.
Pick the path that matches what you have. The control plane is free with no key up to 4 GPUs in one cell, and every image pulls anonymously from ghcr.io/opod-io for amd64 and arm64. The laptop path was re-run word for word against the public 0.1.0 chart on 2026-09-12. Installing the control plane accepts its EULA; core is Apache-2.0 and needs no agreement.
Where the control plane runs
- Inside your cluster, inside your network.
opodcpis one Deployment in theopodnamespace, next to your GPU nodes — in your VPC, on your bare metal, or air-gapped. There is no hosted part and nothing of ours runs outside it. - Network it needs: outbound pulls from
ghcr.io(or your mirror) and the model source; no inbound connection from us, ever. The console and API are reached the way you reach anything in that cluster — port-forward, your ingress, or a NodePort on the LAN. - Serving does not depend on it. Every endpoint keeps answering from its own plan if the control plane is down, upgraded, or deleted.
EKS, GKE, AKS, k3s, RKE2 — any Kubernetes ≥ 1.28 with a default StorageClass. About two minutes with warm images.
- 1
Label your GPU nodes
The chart's device plugins and VRAM probes schedule by this label.
opodcp infra updoes it for you on bare machines.kubectl label node <gpu-node> opod.io/gpu-vendor=nvidia
- 2
Check the cluster (optional — the chart runs the same preflight as a pre-install Job)
docker run --rm --network host -v ~/.kube:/kube:ro \ ghcr.io/opod-io/opodcp:0.1.0 preflight --kubeconfig /kube/config
- 3
Install the control plane
One Deployment with the console inside, a preflight Job, RBAC and per-vendor probes. No databases.
export ADMIN_TOKEN=$(openssl rand -hex 16) helm install opod oci://ghcr.io/opod-io/charts/opod --version 0.1.0 \ -n opod --create-namespace --set admin.token=$ADMIN_TOKEN kubectl -n opod port-forward svc/opod-opodcp 8080:8080 # console: http://127.0.0.1:8080 — sign in with $ADMIN_TOKEN
- 4
Create an endpoint
In the console, Endpoints → New, or from the API. Leave the node out and the placer picks a GPU with room.
export OPOD=http://127.0.0.1:8080 curl -s -X POST $OPOD/api/v1/endpoints \ -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \ -d '{"name":"chat","model":"qwen2.5-0.5b-gguf", "plan":{"model":{"id":"qwen2.5-0.5b-gguf","engine":"llamacpp"}, "workers":[{"replicas":1,"engine":"llamacpp"}]}}' - 5
Wait for a real token, then chat
pending → converging → serving, andservingonly after the leader produced a token. In the cluster the endpoint ishttp://opod-ep-chat.opod.svc:8080/v1.curl -s -H "Authorization: Bearer $ADMIN_TOKEN" $OPOD/api/v1/endpoints/<id>/status | jq .phase,.proof curl -s -X POST $OPOD/api/v1/endpoints/<id>/gateway/v1/chat/completions \ -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \ -d '{"model":"qwen2.5-0.5b-gguf","messages":[{"role":"user","content":"hello"}]}'
The same product on a laptop with a llama.cpp CPU worker — Apple silicon works. This is also what our CI runs.
- 1
Make a cluster and install
kind create cluster --name opod export ADMIN_TOKEN=$(openssl rand -hex 16) helm install opod oci://ghcr.io/opod-io/charts/opod --version 0.1.0 \ -n opod --create-namespace --set admin.token=$ADMIN_TOKEN --wait kubectl -n opod port-forward svc/opod-opodcp 8080:8080 &
- 2
Create a CPU endpoint
With no GPU, name the node — the placer never picks a CPU-only node by itself.
export OPOD=http://127.0.0.1:8080 curl -s -X POST $OPOD/api/v1/endpoints \ -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \ -d '{"name":"chat","model":"qwen2.5-0.5b-gguf", "plan":{"model":{"id":"qwen2.5-0.5b-gguf","engine":"llamacpp"}, "workers":[{"node":"opod-control-plane","replicas":1,"engine":"llamacpp"}]}}' - 3
Chat
curl -s -H "Authorization: Bearer $ADMIN_TOKEN" $OPOD/api/v1/endpoints/<id>/status | jq .phase,.proof curl -s -X POST $OPOD/api/v1/endpoints/<id>/gateway/v1/chat/completions \ -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \ -d '{"model":"qwen2.5-0.5b-gguf","messages":[{"role":"user","content":"hello"}]}'
Measured cold on an Apple-silicon laptop, nothing cached (2026-09-12): helm install --wait 52 s; serving · first token in 144 ms 253 s after the create — mostly the first image pulls and the model download.
Linux machines with GPUs and SSH, no Kubernetes. opodcp infra builds k3s over SSH, labels each node by GPU vendor and writes the kubeconfig.
- 1
Get the
opodcpCLISigned binaries come with the first tagged release; until then the CLI ships inside the image. Run on a Linux box that can reach the others over SSH.
id=$(docker create ghcr.io/opod-io/opodcp:0.1.0) docker cp "$id":/opodcp ./opodcp && docker rm "$id"
- 2
Describe the machines
# machines.yaml cell: lab ssh: { user: ubuntu, keyFile: ~/.ssh/id_ed25519, sudo: true } server: { name: cp-1, host: 192.0.2.10 } agents: - { name: gpu-1, host: 192.0.2.11 } # GPU vendor detected on the host - { name: gpu-2, host: 192.0.2.12 } - 3
Plan, build, check, install
Then continue from step 4 of the Kubernetes path.
--localwhen the server is the machine you're on../opodcp infra plan -f machines.yaml # host facts + gate, changes nothing ./opodcp infra up -f machines.yaml # k3s server + agents over SSH ./opodcp preflight --kubeconfig ~/.opodcp/lab/kubeconfig export KUBECONFIG=~/.opodcp/lab/kubeconfig helm install opod oci://ghcr.io/opod-io/charts/opod --version 0.1.0 \ -n opod --create-namespace --set admin.token=$(openssl rand -hex 16)
Measured on a design-partner cell: infra up built a 9-machine cluster in 80 s; preflight to first token on an RTX A6000 took 127 s.
Only the open-source runtime, on a Mac or Linux box — no Kubernetes, no control plane. One OpenAI URL in front of the engine you already run.
- 1
Build it (Go 1.25+)
Prebuilt binaries arrive with the first tagged release; building takes seconds.
git clone https://github.com/opod-io/opod-core cd opod-core && go build -o opod ./cmd/opod
- 2
Run an engine, start a leader
Ollama is simplest; vLLM, llama.cpp and MLX-LM work too —
./opod doctorsays which fits.# macOS: brew install --cask ollama && open -a Ollama # Linux: curl -fsSL https://ollama.com/install.sh | sh OPOD_DEFAULT_MODEL=llama-3.2-1b ./opod up # prints the API URL and an admin key, once
- 3
Point your tools at it, add machines
curl http://127.0.0.1:8080/v1/chat/completions \ -H "Authorization: Bearer sk-orc-…" -H 'Content-Type: application/json' \ -d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"hello"}]}' ./opod connect --list # Cursor, Aider, Continue, Codex CLI, Open WebUI… ./opod token create --node # on the leader ./opod join "http://<leader>:8080?token=<node-token>" # on each other machine
Early usages
What you can run on it today.
Labelled honestly: ready in chart 0.1.0 and proven on real hardware · early works, proven narrowly · next chart merged, not yet published.
A private OpenAI endpoint for your team
readyPoint Codex CLI, Aider, Continue, Open WebUI or the OpenAI SDK at your own GPUs. Mint a key per person with RPM/TPM limits — the leader enforces them even with the control plane down.
POST /api/v1/endpoints/<id>/keys
{"name":"alice","rpmLimit":120}
Several models on one GPU
readyGive each worker a VRAM budget; the placer packs them best-fit against the card's measured free memory and refuses an overcommit with that number. Proven: three endpoints on one 48 GB card.
"colocation": "shared",
"workers": [{"vramBudgetGb": 6, "engine": "llamacpp"}]
Idle models scale to zero
readyA floor of 0 releases the GPU when nobody is asking; the first request wakes a worker while the leader answers 503 + Retry-After. Measured wake: 54 s with llama.cpp.
"autoscale": {"floor": 0, "max": 3,
"targetInflight": 4, "idleSeconds": 600}
Change models and images safely
readyA rollout validates, applies, converges, probes and smoke-chats, recording a proof per step; any failure restores the previous plan. Rollback is one call.
POST /api/v1/endpoints/<id>/rollouts # body = plan POST /api/v1/rollouts/<id>/rollback
Govern it like production
readyGuardrail webhooks before each request, a fallback target when capacity is gone, an audit log of every change, alert rules to Slack / PagerDuty / Alertmanager, usage facts exported to CSV, Parquet, ClickHouse or S3, OIDC SSO.
A hosted model, same controls
readyA remote endpoint is a leader with no GPU in front of any OpenAI-compatible vendor, with a write-only key — same keys, limits, guardrails, audit and usage as your own models.
"model": {"source": "remote",
"remoteUrl": "https://api.vendor.example/v1"}
Find idle GPUs first
readyA read-only report on any cluster, before you install anything: probe pods in a throwaway namespace, removed afterwards. No licence.
opodcp audit --kubeconfig ~/.kube/config \ --duration 5m --price 2.50
A model bigger than one machine
earlyA llama.cpp pipeline-parallel gang across GPU nodes, placed and healed as a unit — proven with a two-node gang on consumer GPUs. A tensor group may cross nodes too; the placer says what that costs on your fabric instead of refusing.
Mixed GPU vendors, one console
earlyServing on NVIDIA. AMD is measured — VRAM, busy and temperature from the driver, placeable by the ledger. Intel Arc and Tenstorrent are inventoried.
Name the model, not the shards
next chartPlacement auto picks the engine, the cards and the split from measured VRAM — one GPU, then tensor parallel inside a node, then pipeline across nodes — and pins the answer with the reason for every choice. It moves only on a rollout.
Embeddings endpoints
next chartFor RAG — proven on a GPU: 768 dimensions, first vector in 175 ms.
Training and batch jobs
next chartJobs claim cards from the same ledger; preemptible jobs yield to serving and resume from a checkpoint.
HA, quotas, SGLang
next chartControl-plane replicas on Postgres, per-team GPU quotas, SGLang as a third engine, Pod Security baseline, and the 4-GPU free tier (0.1.0 still enforces 8).
For the people who run the fleet
Self-serve for dev teams. Governed by platform teams. Trusted by security.
Platform engineer
Zero-to-serving in under an hour.
- Helm on existing Kubernetes, or discover bare machines and enroll them — gate, drivers, k3s join, node-agent. Retries through reboots.
- Live fleet view: reserved / free / what still fits per machine; health, latency and SLOs per endpoint.
- No database pods to run: SQLite, an in-process event bus and /metrics are built in. Connect your own Postgres, NATS, ClickHouse, Vault, OIDC or KEDA with one helm upgrade --set.
Security · compliance
Zero-access, by construction.
- No inbound vendor connection — every link is outbound and resumable. Air-gapped installs from an offline bundle.
- OIDC SSO + RBAC, scoped API keys, TLS everywhere, signed leader↔worker commands.
- Hash-chained audit exportable to your SIEM. Digest-pinned, cosign-verified images; residency pinning per cell; prompt logging off / redacted / full per endpoint.
Engineering leader · FinOps
Every token attributed. Rated in your tools.
- One usage record per request — tenant, team, endpoint, model — in tokens, audio-seconds, images, and always GPU-seconds.
- Hourly Parquet to your object store; showback runs in the tools you already own.
- Per-team quotas and budgets enforced at the leader, plus a GPU-hours-saved report per model and team.
The landscape
Everyone rents you their cloud. We run AI on yours.
Inference clouds make enterprises rent their GPUs. NVIDIA bought the scheduler layer and open-sourced it, NVIDIA-first.
The quadrant enterprises with data gravity must buy — a complete serving control plane, on your own cloud, across vendors — is open.
We compose, we don't fight. vLLM, SGLang, llama.cpp, Kubernetes, KEDA, llm-d, Dynamo — the moat is placement, ledger, management and the console, not router code.
It's genuinely hard
Model-aware autoscaling, gang scheduling, verified rollouts, and a VRAM ledger that packs several models onto one card from measured memory rather than a guess. Four silicon vendors run in our own mixed fleet; NVIDIA serves today and AMD is measured by the same ledger.Clouds won't build it
Managed AI services exist to keep you on that cloud's hardware at that cloud's prices. A plane that cuts your GPU-hours and moves freely across clouds has to be independent — and yours.Inference clouds can't
Their model needs your workloads to move to them. Regulated teams won't move — BYOC is our entire architecture.For the VP of Engineering
Your team could build most of this. Here's the honest math.
The open stack is excellent and we compose it, not compete with it: vLLM and SGLang are solved, llm-d and Dynamo route and disaggregate, KEDA scales, Kueue queues. What it ships is components. What you run in production is a platform.
- Assembly is a project: 2–4 platform engineers, 6–12 months, to wire gateway, tenancy, metering, autoscale, rollouts and images into one system.
- Ownership is a payroll: the same 2+ engineers forever, tracking a stack where every layer ships monthly — $1M+/yr loaded, and it will never be your product.
- The comparison isn't license vs. free. It's license vs. headcount — plus the utilization you don't recover while you build.
1 · No amount of headcount assembles this
Mixed silicon as one service.
One endpoint on NVIDIA + AMD weighted by measured throughput; memory-proportional splits across mixed VRAM; prefill/decode paired across vendors. llm-d is vLLM-centric, Dynamo is NVIDIA-governed — cross-vendor is the piece the ecosystem admits is missing.
2 · No amount of headcount assembles this
Fractional GPUs with an accountable ledger.
Kubernetes hands out whole GPUs; MIG is NVIDIA-only and rigid. The per-GPU VRAM ledger packs N models per GPU with derived KV budgets and corrects itself from measured loads — the mechanism behind the utilization number.
3 · No amount of headcount assembles this
Sovereign operations, end to end.
Bare NAT'd machines with no Kubernetes and no public IP, checkbox-enrolled. Air-gapped bundles, zero inbound connections, offline licences, kill-the-brain survivability. The CNCF stack assumes you already have K8s and connectivity.
And when you want out
The exit is part of the product.
North side is pure OpenAI API — redeploying two endpoints on the open stack is a weekend test we invite you to run. Usage lands as Parquet in your object store. Serving survives us indefinitely: no phone-home, no remote kill. Source escrow available on enterprise terms.
If you're homogeneous NVIDIA on managed Kubernetes with a strong platform team, DIY is a defensible call — we'll tell you so in the assessment. The moment your fleet mixes vendors, touches bare metal, or answers to a regulator, the build option stops existing.
The numbers
The sales deck is your own cloud bill.
A platform fee plus a per-GPU meter, against tens to hundreds of thousands removed per month. Every lever is independently documented — run your own numbers.
- open model vs frontier API, per token50–92 % less
- routing simple traffic to small models10–40× cheaper
- autoscale-to-zero on idle GPU-hoursup to ~67 %
- spot / preemptible, orchestrated safely70–91 % less
- right-sizing provider / instance40–60 %
The calculator uses your inputs and the platform fee alone — an estimate, not a quote.
Formula: today = GPUs × rate × 730 h. With Opod = GPUs × rate × 730 h × (busy hours ÷ 24) + Opod's platform fee. Ignores spot, right-sizing and token routing — the levers above stack on top.
- Today / mo
- —
- With Opod / mo
- —
- You keep / yr
- —
API-heavy product
5B tokens/mo on a frontier API ≈ $75K. Routed 70/30 to open models on 4–8 autoscaled GPUs in their cloud: ~$20–30K all-in.
saves 60–70 %
Idle reserved cluster
16 reserved H100s ≈ $70K/mo at 5–15 % utilization. Scale-to-zero + consolidation at ~35 % busy: ~$35K all-in.
saves ~50 %
Agents at scale
50B tokens/mo ≈ $750K on frontier APIs. Routing + self-hosted open models on autoscaled capacity: ~$80–120K.
saves 83–87 %
Pricing
Priced per managed GPU. Never per token on your own hardware.
Compute runs in your cloud at your cost; our meter is the GPUs we manage — auditable, and it never shrinks because our optimizations worked. Licences are offline-verified: no phone-home, no remote kill, two months' grace.
Community
Free
open-source core · the full control plane up to 4 GPUs in one cell
- core: leader + workers, OpenAI-compatible gateway, API keys — Apache-2.0, no usage limit
- llama.cpp-RPC sharding across machines
- control plane: ledger, autoscale-to-zero, verified rollouts, console — no key, no card
- free tier: 4 GPUs in one cell
Enterprise · BYOC
Per managed GPU
priced by volume · platform floor · annual
- the full control plane in your Kubernetes: ledger, autoscale-to-zero, verified rollouts, tenancy, usage export, console
- managed by us: upgrades, engine tuning, incident response, 24/7 SLA
- self-serve capacity: pick a GPU count, upgrade instantly
- frontier passthrough at provider cost plus a small margin
Enterprise · flat
Custom
regulated, sovereign and multi-cell estates
- compliance pack: audit exports, residency pinning, SOC 2 groundwork
- air-gapped bundle install; contractual audit rights, no phone-home
- multi-cell console, dedicated SRE, custom engine tuning
- marketplace private offers that burn committed spend
Open source
Read it, run it, fork it.
The runtime that holds your requests, keys and weights is Apache-2.0 with no usage limit — anyone may run, modify, redistribute or compete with it. It never needs the control plane to serve, and it has no telemetry.
- ✓Contributions with a DCO sign-off (
git commit -s) — no CLA. - ✓
make checkruns vet, tests and build locally. - ✓Security reports privately, per SECURITY.md.
opod-core
Apache-2.0Leader and worker runtime, OpenAI gateway, API keys, engine drivers (vLLM, llama.cpp, Ollama, MLX-LM), model catalog, llama.cpp RPC sharding, the /admin/v1 contract.
github.com/opod-io/opod-core →
opod-sdk
Apache-2.0Wire types and an admin client for the leader's /admin/v1 surface — build your own manager or wire opod into a platform you already run.
github.com/opod-io/opod-sdk →
Images & chart
public pullLeader, workers (llama.cpp CPU / NVIDIA, vLLM NVIDIA), the control plane and its Helm chart — anonymous pull from ghcr.io/opod-io, mirrorable for air-gapped installs.
ghcr.io/opod-io →
Docs
- Core README — CLI, config, API
- Core quickstart — one and many machines
- Architecture — leader/worker, routing, sharding
- Model catalog
- Chart values:
helm show values oci://ghcr.io/opod-io/charts/opod
Straight talk · early software
What is not there yet.
No signed release yet. The first tagged release brings signed binaries, SBOMs and a signed chart; until then images are pinned by digest when a plan is written.
Serving is proven on NVIDIA. AMD is measured and placeable, its worker image unpublished; Intel and Tenstorrent are inventoried.
Two engines in the control plane: vLLM and llama.cpp. SGLang is merged for the next chart; Ollama and MLX-LM are core-only.
Shared packing is not a security boundary. VRAM budgets keep models apart, not tenants — use dedicated (the default) for untrusted neighbours.
One control-plane replica in 0.1.0. HA on Postgres lands in the next chart, with the 4-GPU free tier (0.1.0 enforces 8).
OpenAI API only. No Anthropic Messages, audio or rerank shapes; usage facts, not invoices.
Questions we get
Straight answers.
What exactly runs in my cluster?
helm upgrade --set. Everything lives inside your perimeter.What happens when the control plane goes down?
Do I need Kubernetes?
Which GPUs and engines?
Do you see my prompts or my usage?
GET /api/v1/licence shows exactly which GPUs are counted.Is Opod open source?
How does billing behave if we stop paying?
Why wouldn't we build this on llm-d, Dynamo and Kubernetes ourselves?
What happens if Opod shuts down or gets acquired?
How do we validate the claims before committing budget?
What isn't in scope?
Early-access program
Send us your cloud bill.
We'll send back the savings.
A two-week, read-only assessment of your GPU and API spend, then a paid pilot on one workload — in your cluster, in days. Early-access customers get preferred pricing for an audited case study.
Make our proofs your acceptance criteria — pilot payment can follow them: