ServerNet AI Gateway

One API for the AI in your product

Choose a model, create a project and connect using the OpenAI request format. Toman credit and a record of every request help you control spending.

from openai import OpenAI

import os

client = OpenAI(
    api_key=os.environ["SERVERNET_API_KEY"],
    base_url="https://servernet.cloud/v1",
    max_retries=0,
)

models = client.models.list()
if not models.data:
    raise RuntimeError("No models available")
available_model_id = models.data[0].id
response = client.chat.completions.create(
    model=available_model_id,
    messages=[{"role":"user","content":"Hello"}],
    max_tokens=128,
)
Only available models are exposed by the API. Availability requires a configured provider, valid prices and a supported execution path.
01

Clear request costs

Input, output and tax are calculated separately. The final charge and request ID appear in your usage report.

02

Budgets per project

Set monthly project budgets and daily key limits. Issue and revoke keys independently.

03

Connect with OpenAI SDK

Change the base URL and API key. The compatible endpoint supports text chat, tool calls and streaming.

Compare models

XiaomiMiMo

MiMo-V2.6-Pro

PreparingText chat

Context window: 1,048,576 tokens

Input99,705
Output201,729
Toman per million tokens
tencent

Hy4-preview

PreparingText chat

Context window: 1,048,576 tokens

Input193,381
Output579,911
Toman per million tokens
deepseek-ai

DeepSeek-V4-Flash

PreparingText chat

Context window: 1,048,576 tokens

Input20,869
Output41,737
Toman per million tokens
zai-org

GLM-5.3

PreparingText chat

Context window: 1,048,576 tokens

Input208,685
Output927,486
Toman per million tokens
deepseek-ai

DeepSeek-V4-Flash-Vision-Exp

PreparingText chat

Context window: 1,048,576 tokens

Input102,024
Output306,071
Toman per million tokens
zai-org

GLM-5.2

PreparingText chat

Context window: 1,048,576 tokens

Input173,904
Output556,492
Toman per million tokens

From model selection to your first response

  1. Check the model and its availability.
  2. Sign in, create a project and a key with ai:chat and ai:models:read.
  3. Fund your wallet and send a test request with a small output limit.
  4. Review the request usage and charge in your dashboard.
Integration guide
10% tax is added to usage charges. Prices are frozen at request admission and settlement uses reported consumption.

Frequently asked questions

How is usage billed?

Billing uses reported model consumption and the prices frozen for your request. Usage and tax are recorded separately. When a response becomes uncertain after sending, consumption is reconciled according to the request state.

Can I use every listed model?

Each model shows its availability. Only available models can be purchased through the API. Preparing and discontinued models remain visible for catalogue reference.

How can I limit team spending?

Set a monthly project budget and a daily API key limit. Holds for concurrent requests count towards these limits.