Dedicated compute, provisioned on request.
Beyond the shared API, Sable can stand up dedicated compute for a specific workload (open-model GPU inference, or metered code-execution sandboxes) on Sable-operated infrastructure. It runs when there’s committed demand, so there’s no idle box billed to you. Everything Sable already does still applies: metered at the gateway, budget-capped per key, paid in prepaid USDT, and returned with a signed receipt.
Steady volume, dedicated throughput, or proof of what ran.
Teams with steady inference volume that want compute sized to it. Agent platforms that want dedicated throughput instead of sharing a pool. Workloads that run code and need a receipt proving exactly what executed. If that’s you, the shared API is where you start; dedicated compute is where you go when the volume is real.
Apply, and Sable sets it up.
Apply with your workload.
Tell us what you want to run (open-model inference or code-execution sandboxes) and your rough volume. There’s no self-serve spin-up here; a person reads it and replies.
Sable provisions it.
We stand up dedicated compute for that workload on Sable-operated infrastructure. Nothing runs, and nothing costs, until it’s set up for you, so there’s no idle box billed to your account.
Call it through the same API.
You reach it over the OpenAI/Anthropic-compatible surface you already use. Every unit is metered at the gateway, capped by your key’s budget, and returned with a signed receipt.
The whole control plane comes with it.
Dedicated compute isn’t a different product with different rules. It’s the same gateway, the same metering, the same proof, pointed at capacity provisioned for you.
Metered per second.
Compute is measured at the gateway: vCPU-seconds and memory for sandboxes, tokens for inference. You pay for what runs, not for a reservation.
Budget-capped per key.
Every API key carries a spend limit and can mint bounded sub-keys. A dedicated workload doesn’t loosen any of that: the same caps apply to the same keys.
Paid in prepaid USDT.
Top up a balance in USDT; compute draws it down. No invoices, no monthly minimums, no credit line to negotiate.
Signed, verifiable receipts.
Each request returns a receipt anyone can verify: metadata and a content fingerprint, never your prompt or code. Proof of what ran, independent of us.
What it is.
- Sable-operated dedicated compute, provisioned per customer for a specific workload.
- Provisioned on committed demand: no idle capacity, so no idle cost passed to you.
- Reached through the same metered, receipted OpenAI/Anthropic-compatible API.
What it isn’t.
- Not a marketplace. No third-party operators serve traffic today, and we won’t call it one until they do.
- Not instant self-serve. It’s an application a person answers, not a button that spins up a box.
- Not a rented VM or SSH box. You get metered API access to compute, not a machine to log into.
Tell us what you want to run.
There’s no signup wizard. Email us and a person answers. The more you can say up front, the faster we can size it.
- Whether it’s open-model inference, code-execution sandboxes, or both.
- Rough volume: requests or vCPU-hours per day, and how bursty it is.
- Which models you need served, or what your sandbox jobs do.
- Anything that shapes it: latency needs, region, whether you need the receipt for compliance.