Opod Control Plane Enterprise early access

Any GPU. Any cloud. One inference service you run yourself.

Opod installs into your own Kubernetes — cloud, VPC, on-prem, or air-gapped — and turns your GPU machines into governed, autoscaling, OpenAI-compatible endpoints. Serving runs on NVIDIA today; AMD is measured and packed by the same ledger, with its worker image next. Prompts never leave your cluster.

Free tier, no key, no sign-up · open-source core (Apache-2.0) · or a two-week read-only assessment of your bill.

Fleet time
14:00
GPUs powered
—
Endpoints
4
Utilization
—
live scene · one cell · mixed-vendor fleet NVIDIA AMD Intel Tenstorrent
day · demand high

Each lit square is a worker on a GPU. Overnight, workers scale to zero and GPUs are released; the four leaders stay up, and the first morning request wakes everything within its SLO.

  • ≤ 2 ms
    added by the gateway at p99
  • 127 s
    preflight to first token on a real GPU node
  • 0
    inbound vendor connections

The problem

Inference burns money three ways.

GPUs sit at 5–30 % utilization while frontier-API token bills triple as agents replace chat. The engines that fix this are excellent and free — the management layer above them is the gap.

1 · GPUs

5–30 %

typical enterprise utilization. Most of an always-on GPU bill is paid-for idle — $3–8K per GPU per month.

2 · Tokens

50–500×

more tokens per task from agents than chat. Frontier APIs charge $12–30 per million output tokens; open models served well cost $0.04–1.

3 · DIY

$1M+/yr

to build this layer yourself, over 6–18 months — and it will never be your product.

Paid GPUs vs. traffic over one day always-on reservation with Opod autoscale (three tiers) traffic · GPUs the load needs saving · paid for, never used GPUs · 16-GPU reservation
Sleep tier 0.1–0.8 s wake, GPU retained Pod-zero 10–60 s, weights cached on node NVMe Node-zero 2–10 min, node released to the cloud Scale-to-zero on an 8 h/day pattern saves up to ~67 % of the GPU line.

What Opod is

One control plane. Any cloud, GPU, engine, model.

Opod composes the best open engines instead of forking them, runs on your Kubernetes (or brings its own k3s), and gives every team self-serve endpoints inside quota — metered per request.

Everything stays in your cloud.

Zero inbound vendor connections. Prompts and completions never leave your cluster. Air-gapped installs are first-class.

Any cloud
AWS · EKSGoogle Cloud · GKEAzure · AKSNeocloudsOn-prem · k3sAir-gapped
Any GPU
NVIDIA servingAMD measuredIntel inventoriedTenstorrent inventoriedone ledger, one placer, one console across all of them
Any engine
vLLMSGLangllama.cppOllamaTEIWhisperTTSdiffusionadapters over official upstreams, never forks
Any model
LlamaDeepSeekQwenGLMMistralKimiGPT-OSSyour fine-tunesLLM · embeddings · speech · images · custom
One API
/v1/chat/completions/v1/embeddings/v1/modelsstreaming everywhereunchanged OpenAI SDKs

How it works

The endpoint is the object. The money path waits for no one.

An endpoint is a stable URL with keys and policy, served by one leader and the workers that join it — each worker one engine instance on a GPU share. The control plane places everything and stores everything of record, but it is never on the request path.

  • Static stability. Lose the control plane and requests keep flowing; usage spools locally and streams back later.
  • One product API. Dashboard, CLI, CRDs and Terraform all drive the same versioned plan per endpoint.
  • Everything is an event. One append-only log; every store is a projection you can replay.
  • Trust is proven. A node-agent gates every machine before any workload lands.
One cell, two endpoints — break things

Requests served

0

Usage spooled on leaders

0

Status

Healthy. Control plane pushes plans; leaders route; usage streams back.

Worker

one engine instance inside its VRAM budget; many per GPU; disposable, rebuilt from heartbeats.

Leader

the endpoint's front door — auth, queues, streaming, one usage record per request. Failover is a ~10 s restart.

Control plane

ledger, placement, rollouts, tenancy, usage, console. One per org, many cells. Never on the request path.
Per-GPU VRAM ledger · one 4×80 GB machine
weights KV budget (derived) activation · overhead headroom free · what still fits

Capacity ledger · placement · autoscale

Utilization is the product.

A deterministic placement library with a per-GPU VRAM ledger — fleet state and workloads in, assignments out. Replay it against a fleet before you buy it.

Every deployment is a reservation.

Weights + KV budget + overhead + headroom, per GPU. The dashboard shows reserved, free, and what still fits.

Honest bin-packing.

Dedicated means a whole GPU and a hard SLA; shared means packed. We never claim isolation we can't enforce.

Priority and preemption.

critical › standard › preemptible — nightly batch soaks idle GPUs and is evicted in seconds when morning traffic arrives.

Three autoscale tiers.

Sleep, pod-zero with cached weights, node-zero — each endpoint picks its wake SLO. Sharded models scale as whole gangs.

Target: ≥ 2× your utilization baseline within 30 days, measured by the GPU-hours-saved report per model and team.

Where the biggest number comes from

Pack the GPU.

The field's rule is one model per card. Opod's ledger knows each model's footprint — weights plus a KV-cache reserve — and each worker's budget, so it places several models on one card and leaves the rest free or asleep.

Drag the slider to change how many models you want resident. Footprints are illustrative sizes for common open models, listed in the table under the chart.

Cards needed to keep N models resident · 80 GB cards model, not a measurement

One model per card versus first-fit packing against a per-worker budget, with 8 GB of headroom kept on every card.

Cards, one per model
—
Cards, packed by the ledger
—
VRAM idle, one per model
—
VRAM idle, packed
—
Cards freed for sleep or batch
—
weights KV-cache reserve headroom unused VRAM
Model footprints and the current placement

Footprint = weights at the listed precision + a KV-cache reserve sized for typical agent context. Packing is first-fit decreasing on a single node class. Real placement also respects engine compatibility, vendor and gang rules, so the real number is a little lower than this model. The point is the ratio, not the decimals.

Verified rollouts

Every deploy step produces a proof.

Not "the pod is Running" — a machine-checkable proof at every step, with host-level GPU evidence. On failure: diagnostics captured before rollback, plus a suggested remedy. The same saga upgrades the platform itself.

llama-3.1-70b · 2 gangs · NVIDIA + AMD
  1. 1

    Preflight

    placement validated against the ledger · capacity present · node gate passed on every target

    proof: pending

  2. 2

    Images

    digest-pinned, cosign-verified, SBOM published, present in the cell mirror — built in CI, never on nodes

    proof: pending

  3. 3

    Launch

    plan ConfigMap written · leader Deployment and worker pods scheduled on the reserved GPU shares

    proof: pending

  4. 4

    Place

    weights resident from node NVMe cache · every worker heartbeats loaded with measured VRAM

    proof: pending

  5. 5

    Health

    leader readiness · a real completion round-trips through the gateway

    proof: pending

  6. 6

    GPU proof

    host-level utilization evidence from the node-agent, not the container

    proof: pending

  7. 7

    Resilience

    kill one worker pod · the leader heals within its plan · zero dropped requests

    proof: pending

For developers

Ship on it in an afternoon. Keep your SDK.

Create an endpoint from the console or one API call and point the OpenAI SDK you already use at it. Every change rolls out with proof; roll back with one call.

  • ✓OpenAI-compatible /v1/chat/completions and /v1/models, streaming; embeddings endpoints in the next chart.
  • ✓Models from the catalog or any Hugging Face GGUF / safetensors repo — images pinned by digest.
  • ✓Per-key RPM/TPM limits enforced by the leader, guardrail webhooks, a fallback target.
  • ✓One REST API behind the console — POST /api/v1/endpoints/preview dry-runs a create.
Free

Opod core, Apache-2.0. No Kubernetes needed: opod up on one box, opod join from the rest, one OpenAI endpoint. Point Cursor, Aider or Codex CLI at it. Quickstart →

# pip install openai — nothing else changes
from openai import OpenAI

client = OpenAI(
    base_url="http://opod-ep-chat.opod.svc:8080/v1",  # your endpoint, in your cluster
    api_key="sk-orc-…",                               # minted per person, with limits
)

stream = client.chat.completions.create(
    model="qwen2.5-0.5b-gguf",
    messages=[{"role": "user", "content": "Summarise this contract."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Quickstart · free, no sign-up

From nothing to a first token.

Pick the path that matches what you have. The control plane is free with no key up to 4 GPUs in one cell, and every image pulls anonymously from ghcr.io/opod-io for amd64 and arm64. The laptop path was re-run word for word against the public 0.1.0 chart on 2026-09-12. Installing the control plane accepts its EULA; core is Apache-2.0 and needs no agreement.

Where the control plane runs

  • Inside your cluster, inside your network. opodcp is one Deployment in the opod namespace, next to your GPU nodes — in your VPC, on your bare metal, or air-gapped. There is no hosted part and nothing of ours runs outside it.
  • Network it needs: outbound pulls from ghcr.io (or your mirror) and the model source; no inbound connection from us, ever. The console and API are reached the way you reach anything in that cluster — port-forward, your ingress, or a NodePort on the LAN.
  • Serving does not depend on it. Every endpoint keeps answering from its own plan if the control plane is down, upgraded, or deleted.

EKS, GKE, AKS, k3s, RKE2 — any Kubernetes ≥ 1.28 with a default StorageClass. About two minutes with warm images.

  1. 1

    Label your GPU nodes

    The chart's device plugins and VRAM probes schedule by this label. opodcp infra up does it for you on bare machines.

    kubectl label node <gpu-node> opod.io/gpu-vendor=nvidia
  2. 2

    Check the cluster (optional — the chart runs the same preflight as a pre-install Job)

    docker run --rm --network host -v ~/.kube:/kube:ro \
      ghcr.io/opod-io/opodcp:0.1.0 preflight --kubeconfig /kube/config
  3. 3

    Install the control plane

    One Deployment with the console inside, a preflight Job, RBAC and per-vendor probes. No databases.

    export ADMIN_TOKEN=$(openssl rand -hex 16)
    helm install opod oci://ghcr.io/opod-io/charts/opod --version 0.1.0 \
      -n opod --create-namespace --set admin.token=$ADMIN_TOKEN
    
    kubectl -n opod port-forward svc/opod-opodcp 8080:8080
    # console: http://127.0.0.1:8080 — sign in with $ADMIN_TOKEN
  4. 4

    Create an endpoint

    In the console, Endpoints → New, or from the API. Leave the node out and the placer picks a GPU with room.

    export OPOD=http://127.0.0.1:8080
    curl -s -X POST $OPOD/api/v1/endpoints \
      -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \
      -d '{"name":"chat","model":"qwen2.5-0.5b-gguf",
           "plan":{"model":{"id":"qwen2.5-0.5b-gguf","engine":"llamacpp"},
                   "workers":[{"replicas":1,"engine":"llamacpp"}]}}'
  5. 5

    Wait for a real token, then chat

    pending → converging → serving, and serving only after the leader produced a token. In the cluster the endpoint is http://opod-ep-chat.opod.svc:8080/v1.

    curl -s -H "Authorization: Bearer $ADMIN_TOKEN" $OPOD/api/v1/endpoints/<id>/status | jq .phase,.proof
    
    curl -s -X POST $OPOD/api/v1/endpoints/<id>/gateway/v1/chat/completions \
      -H "Authorization: Bearer $ADMIN_TOKEN" -H 'Content-Type: application/json' \
      -d '{"model":"qwen2.5-0.5b-gguf","messages":[{"role":"user","content":"hello"}]}'

Early usages

What you can run on it today.

Labelled honestly: ready in chart 0.1.0 and proven on real hardware · early works, proven narrowly · next chart merged, not yet published.

A private OpenAI endpoint for your team

ready

Point Codex CLI, Aider, Continue, Open WebUI or the OpenAI SDK at your own GPUs. Mint a key per person with RPM/TPM limits — the leader enforces them even with the control plane down.

POST /api/v1/endpoints/<id>/keys
{"name":"alice","rpmLimit":120}

Several models on one GPU

ready

Give each worker a VRAM budget; the placer packs them best-fit against the card's measured free memory and refuses an overcommit with that number. Proven: three endpoints on one 48 GB card.

"colocation": "shared",
"workers": [{"vramBudgetGb": 6, "engine": "llamacpp"}]

Idle models scale to zero

ready

A floor of 0 releases the GPU when nobody is asking; the first request wakes a worker while the leader answers 503 + Retry-After. Measured wake: 54 s with llama.cpp.

"autoscale": {"floor": 0, "max": 3,
  "targetInflight": 4, "idleSeconds": 600}

Change models and images safely

ready

A rollout validates, applies, converges, probes and smoke-chats, recording a proof per step; any failure restores the previous plan. Rollback is one call.

POST /api/v1/endpoints/<id>/rollouts  # body = plan
POST /api/v1/rollouts/<id>/rollback

Govern it like production

ready

Guardrail webhooks before each request, a fallback target when capacity is gone, an audit log of every change, alert rules to Slack / PagerDuty / Alertmanager, usage facts exported to CSV, Parquet, ClickHouse or S3, OIDC SSO.

A hosted model, same controls

ready

A remote endpoint is a leader with no GPU in front of any OpenAI-compatible vendor, with a write-only key — same keys, limits, guardrails, audit and usage as your own models.

"model": {"source": "remote",
  "remoteUrl": "https://api.vendor.example/v1"}

Find idle GPUs first

ready

A read-only report on any cluster, before you install anything: probe pods in a throwaway namespace, removed afterwards. No licence.

opodcp audit --kubeconfig ~/.kube/config \
  --duration 5m --price 2.50

A model bigger than one machine

early

A llama.cpp pipeline-parallel gang across GPU nodes, placed and healed as a unit — proven with a two-node gang on consumer GPUs. A tensor group may cross nodes too; the placer says what that costs on your fabric instead of refusing.

Mixed GPU vendors, one console

early

Serving on NVIDIA. AMD is measured — VRAM, busy and temperature from the driver, placeable by the ledger. Intel Arc and Tenstorrent are inventoried.

Name the model, not the shards

next chart

Placement auto picks the engine, the cards and the split from measured VRAM — one GPU, then tensor parallel inside a node, then pipeline across nodes — and pins the answer with the reason for every choice. It moves only on a rollout.

Embeddings endpoints

next chart

For RAG — proven on a GPU: 768 dimensions, first vector in 175 ms.

Training and batch jobs

next chart

Jobs claim cards from the same ledger; preemptible jobs yield to serving and resume from a checkpoint.

HA, quotas, SGLang

next chart

Control-plane replicas on Postgres, per-team GPU quotas, SGLang as a third engine, Pod Security baseline, and the 4-GPU free tier (0.1.0 still enforces 8).

For the people who run the fleet

Self-serve for dev teams. Governed by platform teams. Trusted by security.

Platform engineer

Zero-to-serving in under an hour.

  • Helm on existing Kubernetes, or discover bare machines and enroll them — gate, drivers, k3s join, node-agent. Retries through reboots.
  • Live fleet view: reserved / free / what still fits per machine; health, latency and SLOs per endpoint.
  • No database pods to run: SQLite, an in-process event bus and /metrics are built in. Connect your own Postgres, NATS, ClickHouse, Vault, OIDC or KEDA with one helm upgrade --set.

Security · compliance

Zero-access, by construction.

  • No inbound vendor connection — every link is outbound and resumable. Air-gapped installs from an offline bundle.
  • OIDC SSO + RBAC, scoped API keys, TLS everywhere, signed leader↔worker commands.
  • Hash-chained audit exportable to your SIEM. Digest-pinned, cosign-verified images; residency pinning per cell; prompt logging off / redacted / full per endpoint.

Engineering leader · FinOps

Every token attributed. Rated in your tools.

  • One usage record per request — tenant, team, endpoint, model — in tokens, audio-seconds, images, and always GPU-seconds.
  • Hourly Parquet to your object store; showback runs in the tools you already own.
  • Per-team quotas and budgets enforced at the leader, plus a GPU-hours-saved report per model and team.

The landscape

Everyone rents you their cloud. We run AI on yours.

Inference clouds make enterprises rent their GPUs. NVIDIA bought the scheduler layer and open-sourced it, NVIDIA-first.

The quadrant enterprises with data gravity must buy — a complete serving control plane, on your own cloud, across vendors — is open.

We compose, we don't fight. vLLM, SGLang, llama.cpp, Kubernetes, KEDA, llm-d, Dynamo — the moat is placement, ledger, management and the console, not router code.

Where inference platforms sithover a point

It's genuinely hard

Model-aware autoscaling, gang scheduling, verified rollouts, and a VRAM ledger that packs several models onto one card from measured memory rather than a guess. Four silicon vendors run in our own mixed fleet; NVIDIA serves today and AMD is measured by the same ledger.

Clouds won't build it

Managed AI services exist to keep you on that cloud's hardware at that cloud's prices. A plane that cuts your GPU-hours and moves freely across clouds has to be independent — and yours.

Inference clouds can't

Their model needs your workloads to move to them. Regulated teams won't move — BYOC is our entire architecture.

For the VP of Engineering

Your team could build most of this. Here's the honest math.

The open stack is excellent and we compose it, not compete with it: vLLM and SGLang are solved, llm-d and Dynamo route and disaggregate, KEDA scales, Kueue queues. What it ships is components. What you run in production is a platform.

  • Assembly is a project: 2–4 platform engineers, 6–12 months, to wire gateway, tenancy, metering, autoscale, rollouts and images into one system.
  • Ownership is a payroll: the same 2+ engineers forever, tracking a stack where every layer ships monthly — $1M+/yr loaded, and it will never be your product.
  • The comparison isn't license vs. free. It's license vs. headcount — plus the utilization you don't recover while you build.

1 · No amount of headcount assembles this

Mixed silicon as one service.

One endpoint on NVIDIA + AMD weighted by measured throughput; memory-proportional splits across mixed VRAM; prefill/decode paired across vendors. llm-d is vLLM-centric, Dynamo is NVIDIA-governed — cross-vendor is the piece the ecosystem admits is missing.

2 · No amount of headcount assembles this

Fractional GPUs with an accountable ledger.

Kubernetes hands out whole GPUs; MIG is NVIDIA-only and rigid. The per-GPU VRAM ledger packs N models per GPU with derived KV budgets and corrects itself from measured loads — the mechanism behind the utilization number.

3 · No amount of headcount assembles this

Sovereign operations, end to end.

Bare NAT'd machines with no Kubernetes and no public IP, checkbox-enrolled. Air-gapped bundles, zero inbound connections, offline licences, kill-the-brain survivability. The CNCF stack assumes you already have K8s and connectivity.

And when you want out

The exit is part of the product.

North side is pure OpenAI API — redeploying two endpoints on the open stack is a weekend test we invite you to run. Usage lands as Parquet in your object store. Serving survives us indefinitely: no phone-home, no remote kill. Source escrow available on enterprise terms.

If you're homogeneous NVIDIA on managed Kubernetes with a strong platform team, DIY is a defensible call — we'll tell you so in the assessment. The moment your fleet mixes vendors, touches bare metal, or answers to a regulator, the build option stops existing.

The numbers

The sales deck is your own cloud bill.

A platform fee plus a per-GPU meter, against tens to hundreds of thousands removed per month. Every lever is independently documented — run your own numbers.

  • open model vs frontier API, per token50–92 % less
  • routing simple traffic to small models10–40× cheaper
  • autoscale-to-zero on idle GPU-hoursup to ~67 %
  • spot / preemptible, orchestrated safely70–91 % less
  • right-sizing provider / instance40–60 %

The calculator uses your inputs and the platform fee alone — an estimate, not a quote.

Idle-GPU calculatorreserved GPUs · always-on vs autoscaled

Formula: today = GPUs × rate × 730 h. With Opod = GPUs × rate × 730 h × (busy hours ÷ 24) + Opod's platform fee. Ignores spot, right-sizing and token routing — the levers above stack on top.

Today / mo
—
With Opod / mo
—
You keep / yr
—

API-heavy product

5B tokens/mo on a frontier API ≈ $75K. Routed 70/30 to open models on 4–8 autoscaled GPUs in their cloud: ~$20–30K all-in.

saves 60–70 %

Idle reserved cluster

16 reserved H100s ≈ $70K/mo at 5–15 % utilization. Scale-to-zero + consolidation at ~35 % busy: ~$35K all-in.

saves ~50 %

Agents at scale

50B tokens/mo ≈ $750K on frontier APIs. Routing + self-hosted open models on autoscaled capacity: ~$80–120K.

saves 83–87 %

Pricing

Priced per managed GPU. Never per token on your own hardware.

Compute runs in your cloud at your cost; our meter is the GPUs we manage — auditable, and it never shrinks because our optimizations worked. Licences are offline-verified: no phone-home, no remote kill, two months' grace.

Community

Free

open-source core · the full control plane up to 4 GPUs in one cell

  • core: leader + workers, OpenAI-compatible gateway, API keys — Apache-2.0, no usage limit
  • llama.cpp-RPC sharding across machines
  • control plane: ledger, autoscale-to-zero, verified rollouts, console — no key, no card
  • free tier: 4 GPUs in one cell
Start the quickstart

Enterprise · BYOC

Per managed GPU

priced by volume · platform floor · annual

  • the full control plane in your Kubernetes: ledger, autoscale-to-zero, verified rollouts, tenancy, usage export, console
  • managed by us: upgrades, engine tuning, incident response, 24/7 SLA
  • self-serve capacity: pick a GPU count, upgrade instantly
  • frontier passthrough at provider cost plus a small margin
Talk to us

Enterprise · flat

Custom

regulated, sovereign and multi-cell estates

  • compliance pack: audit exports, residency pinning, SOC 2 groundwork
  • air-gapped bundle install; contractual audit rights, no phone-home
  • multi-cell console, dedicated SRE, custom engine tuning
  • marketplace private offers that burn committed spend
Request a proposal

Open source

Read it, run it, fork it.

The runtime that holds your requests, keys and weights is Apache-2.0 with no usage limit — anyone may run, modify, redistribute or compete with it. It never needs the control plane to serve, and it has no telemetry.

  • ✓Contributions with a DCO sign-off (git commit -s) — no CLA.
  • ✓make check runs vet, tests and build locally.
  • ✓Security reports privately, per SECURITY.md.

Straight talk · early software

What is not there yet.

No signed release yet. The first tagged release brings signed binaries, SBOMs and a signed chart; until then images are pinned by digest when a plan is written.

Serving is proven on NVIDIA. AMD is measured and placeable, its worker image unpublished; Intel and Tenstorrent are inventoried.

Two engines in the control plane: vLLM and llama.cpp. SGLang is merged for the next chart; Ollama and MLX-LM are core-only.

Shared packing is not a security boundary. VRAM budgets keep models apart, not tenants — use dedicated (the default) for untrusted neighbours.

One control-plane replica in 0.1.0. HA on Postgres lands in the next chart, with the 4-GPU free tier (0.1.0 enforces 8).

OpenAI API only. No Anthropic Messages, audio or rerank shapes; usage facts, not invoices.

Questions we get

Straight answers.

What exactly runs in my cluster?
A small leader pod per endpoint, worker pods on whole GPUs or GPU shares, a per-vendor probe, and the control plane — one Helm release with its state in SQLite on a volume, no database pods. Postgres, NATS, ClickHouse, Vault and OIDC connect later with helm upgrade --set. Everything lives inside your perimeter.
What happens when the control plane goes down?
Nothing, from your users' point of view. Leaders keep serving from a mounted plan and cached auth, spool usage locally, and reconcile when the link returns — you just can't create new endpoints until it's back. We drill it: kill the control plane, serving is unaffected.
Do I need Kubernetes?
Kubernetes is the executor — EKS, GKE, AKS or k3s ≥ 1.28. For bare machines we bring k3s: discover over SSH or cloud credentials, checkbox, and enrollment converges them. NAT'd machines work; the community CLI needs no Kubernetes at all.
Which GPUs and engines?
Straight answer, because it is easy to check: serving is proven on NVIDIA. AMD is inventoried and its VRAM measured, so the ledger packs it like any other card; the ROCm worker image is the next step. Intel and Tenstorrent are inventoried — present, driver state known, not yet serving. All four run in our own mixed fleet. Engines are adapters over official upstream images (vLLM, SGLang, llama.cpp, Ollama, …); we never compile or fork an engine.
Do you see my prompts or my usage?
No. Prompts and completions never leave your cluster, and content logging is a per-endpoint policy in your own log backend. Licensing is an offline capacity token — no phone-home; GET /api/v1/licence shows exactly which GPUs are counted.
Is Opod open source?
The leader/worker core and CLI are open source under Apache-2.0 with no usage limit (opod-io/opod-core), and so is the SDK. The control plane is commercial: publicly pullable, free up to 4 GPUs in one cell with no key. Remove the control plane and the core still serves on its own.
How does billing behave if we stop paying?
Payment failure never touches the data plane. Two months of full service past expiry, then vendor-side degrade only: no new enrollment, dashboard read-only. Never a remote kill; running nodes are never touched.
Why wouldn't we build this on llm-d, Dynamo and Kubernetes ourselves?
If you're all-NVIDIA on managed Kubernetes with a platform team to spare, you genuinely can — budget 2–4 engineers for 6–12 months of assembly, then 2+ forever for day-2 ownership. What you can't assemble at any headcount: mixed-vendor serving (including cross-vendor prefill/decode), fractional-GPU packing with an accountable VRAM ledger, and sovereign onboarding of bare NAT'd machines. We compose the same open stack underneath — buying us is buying the assembled, owned system plus the three pieces that don't exist in the open. See Build vs buy.
What happens if Opod shuts down or gets acquired?
Your service doesn't notice. Serving runs entirely in your cluster from cached plans, licences verify offline, and there is no remote kill — running nodes serve indefinitely. Your data was always in your stores; usage exports as Parquet. Because the north side is the plain OpenAI API, redeploying endpoints on the open stack is a weekend migration, not a rewrite — we encourage you to rehearse it during the pilot. Source escrow and change-of-control terms are available on enterprise contracts.
How do we validate the claims before committing budget?
Don't take the demo's word — the product is built to prove itself on your fleet. Five drills, all machine-checkable: kill the control plane with zero dropped requests; the ledger packing your real model mix multiple-per-GPU; one plan serving on NVIDIA + AMD; 24 hours of autoscale inside your wake SLOs; exported usage reconciling to the token against your own gateway logs. Make them the pilot's acceptance criteria — payment milestones can follow them.
What isn't in scope?
We don't build engines, we don't compute invoices (we export usage facts; your FinOps tools rate them), we don't claim hard GPU isolation outside MIG, and training orchestration comes after Acts 1–2.

Early-access program

Send us your cloud bill.
We'll send back the savings.

A two-week, read-only assessment of your GPU and API spend, then a paid pilot on one workload — in your cluster, in days. Early-access customers get preferred pricing for an audited case study.

Make our proofs your acceptance criteria — pilot payment can follow them:

kill the control plane · zero dropped requests your model mix packed multi-per-GPU by the ledger the same plan serving on NVIDIA + AMD 24 h of autoscale inside your wake SLOs exported usage reconciles to the token
Start the assessment Watch a live cluster demo

Replies within one business day.