Open-weight inference on Kubox

Open-weight model inference

With drop in replacement for OpenAI or Anthropic SDK. Pick a model, a hosting location and set your budget.

Every model shows where it runs and what it costs.

No guessing and no fine print. You’ll see the hosting location and the rate for each model before you connect anything, so you can choose the right fit for the job and for your compliance team.

The Kubox model catalogue, listing open-weight models with their capabilities, hosting location and per-token rate

From sign-up to your first response in a few minutes.

  1. 1

    Pick a model

    Compare what each one can do, where it runs and what it costs.

  2. 2

    Create a key

    Give it the name of the app that will use it.

  3. 3

    Send a request

    Copy our example, run it, and watch the usage appear straight away.

Stay in control of spending without spreadsheets.

One key per app

You’ll always know which app made which call.

See where the money went

Check what each key has spent, then open the individual calls behind it.

Track spending against a budget

Set a budget per key and see how much remains.

In partnership with SCX

Your prompts can stay in Australia.

We partnered with SCX so selected models run on infrastructure in Australia. SCX runs isolated inference on purpose-built hardware, with IRAP-aligned controls. SCX states that prompts are not cached or used for training, and are never stored, indexed or replayed.

You can read SCX’s security details at scx.ai. Each model in the catalogue shows exactly where it runs, so you always know what you’re choosing.

Already using the OpenAI SDK? Change two lines.

Point your existing client at Kubox, swap in your key, and you’re done.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.kubox.cloud/v1",
    api_key="YOUR_KUBOX_API_KEY",
)

response = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "Summarise this changelog."}],
)

print(response.choices[0].message.content)

FAQs

Do I need to run any infrastructure?
No. We host the models, so you just call the endpoint.
Are all models hosted in Australia?
Not all of them, so we show the location for every model up front. If everything needs to stay onshore, choose a model carrying the Sovereign badge.
How do I choose a model?
Start with what you’re building and what your hosting rules allow, then try a handful of real requests. If you’d like a second opinion, book a quick chat and we’ll suggest a starting point.
What will it cost?
Each model shows its rate per million tokens before you send anything.

Need it fully air-gapped?

Some workloads can’t leave your network at all. We can run open-weight models entirely inside your environment. Tell us what you need and we’ll work it through with you.