DocumentationDokumentasi

Everything you need to use Zenicat AI Gateway: how tokens are counted, the specifications of each model, endpoints, code examples, and what every error code means. Semua yang perlu kamu tahu untuk memakai Zenicat AI Gateway: cara menghitung token, spesifikasi tiap model, endpoint, contoh kode, dan arti setiap kode error.

Quick startMulai cepat

This API follows the OpenAI standard, so the libraries and apps you already use will work — you only change the base URL. API ini mengikuti standar OpenAI, jadi pustaka dan aplikasi yang sudah kamu pakai bisa langsung dipakai — cukup ganti alamat dasarnya.

Base URLAlamat dasar https://zenicat.net/ai/v1
HeaderHeader Authorization: Bearer sk-ksr-…
ModelsModel loading…
Keep your key safe. It is shown only once when created. If it leaks, ask us to disable it and issue a new one — unused credit is not lost. Simpan kuncimu sebaik mungkin. Kunci hanya ditampilkan sekali saat dibuat. Kalau bocor, minta kami matikan lalu buat yang baru — saldo yang belum terpakai tidak hilang.

What is a tokenApa itu token

A token is the smallest unit of text the model recognises — it can be a word, a number, or even a punctuation mark. Tokens are what you are billed for, not the number of messages or characters. Token adalah satuan terkecil teks yang dikenali model — bisa satu kata, satu angka, bahkan satu tanda baca. Token inilah yang dihitung sebagai biaya, bukan jumlah pesan atau jumlah huruf.

In other words, one long message does not cost the same as one short message. What counts is input tokens (your message) and output tokens (the model's reply). Artinya: mengirim satu pesan panjang tidak sama biayanya dengan satu pesan pendek. Yang dihitung adalah token masuk (pesanmu) dan token keluar (balasan model).

Rough character-to-token ratiosPerkiraan kasar karakter ke token

Official figures from the model provider, for Latin-script text: Angka resmi dari penyedia model, untuk teks berhuruf Latin:

Indonesian has no official ratio. Because of its many affixes, Indonesian text tends to use slightly more tokens per word than English. Use 1 token ≈ 3 characters as an estimate, then confirm with the real number (shown below). Bahasa Indonesia belum punya angka resmi. Karena banyak imbuhan, teks Indonesia cenderung sedikit lebih banyak token per kata daripada Inggris. Pakai patokan 1 token ≈ 3 karakter untuk perkiraan, lalu pastikan dengan angka sebenarnya (caranya di bawah).

The big pictureGambaran besarnya

ContentIsi Approx. tokensPerkiraan token
One short chat message (± 50 words)Satu pesan chat pendek (± 50 kata)± 70
One page of a document (± 500 words)Satu halaman dokumen (± 500 kata)± 700–900
A long article (± 2,000 words)Artikel panjang (± 2.000 kata)± 2,800–3,600
A whole thin novel (± 40,000 words)Seluruh isi novel tipis (± 40.000 kata)± 55,000
1,000,000 tokens1.000.000 token ± 750,000 words± 750.000 kata

The figures above are for planning only. The amount billed is always the real number. Perkiraan di atas hanya untuk mengira-ngira. Angka yang ditagih selalu angka sebenarnya.

How to know exactly: read usageCara tahu pasti: baca usage

Every response includes the actual token counts. This is the authoritative source, not an estimate: Setiap balasan menyertakan jumlah token sebenarnya. Ini sumber yang benar, bukan perkiraan:

{
  "choices": [ ... ],
  "usage": {
    "prompt_tokens": 52,        // your message
    "completion_tokens": 118,   // the reply
    "total_tokens": 170,
    "prompt_cache_hit_tokens": 32,
    "prompt_cache_miss_tokens": 20
  },
  "x_ksr": {                    // added by us
    "dipotong_idr": 2.14,       // what was actually deducted
    "saldo_idr": 19997.86
  }
}

The x_ksr field answers the most common question: “how much did this call cost?” — without you having to work it out. Field x_ksr menjawab pertanyaan yang paling sering muncul: “berapakah biaya panggilan ini?” — tanpa perlu menghitung sendiri.

Cached input is far cheaper. If the beginning of your prompt matches an earlier request (for example a fixed system prompt), the model recognises it and counts it 50× cheaper. That is why prompt_cache_hit_tokens is reported separately from prompt_cache_miss_tokens. Putting fixed instructions in system saves a lot for routine usage. Input yang di-cache jauh lebih murah. Kalau bagian awal pesanmu sama dengan permintaan sebelumnya (misalnya system prompt yang tetap), model mengenali dan menghitungnya 50× lebih murah. Inilah sebabnya prompt_cache_hit_tokens terpisah dari prompt_cache_miss_tokens. Menaruh instruksi tetap di system menghemat banyak untuk pemakaian rutin.

Thinking mode is on by default — and you pay for itMode berpikir menyala secara bawaan — dan kamu membayarnya

This is the single most expensive surprise on this API. The model reasons before answering by default, and every reasoning token is billed as an output token — taken from the same max_tokens budget as your answer. Ini kejutan paling mahal di API ini. Model berpikir dulu sebelum menjawab secara bawaan, dan setiap token penalaran ditagih sebagai token keluaran — diambil dari jatah max_tokens yang sama dengan jawabanmu.

Measured result: with max_tokens: 60 and thinking left on, a simple request returns empty content with finish_reason: "length" — all 60 tokens went to reasoning. The same request with thinking off answers in 3 tokens. Hasil pengukuran: dengan max_tokens: 60 dan berpikir dibiarkan menyala, permintaan sederhana mengembalikan content kosong dengan finish_reason: "length" — 60 token habis dipakai penalaran. Permintaan yang sama dengan berpikir dimatikan menjawab dalam 3 token.

To turn it off, add one field to your request: Untuk mematikannya, tambahkan satu field ke permintaanmu:

{
  "model": "deepseek-flash",
  "messages": [{"role": "user", "content": "Hello"}],
  "max_tokens": 60,
  "thinking": {"type": "disabled"}
}
Why bother? A one-sentence answer cost 94 output tokens with thinking on and 3 with it off — about 30× more. Leave thinking on only when you genuinely need step-by-step reasoning, and raise max_tokens accordingly, because reasoning alone can take hundreds of tokens. Kenapa penting? Jawaban satu kalimat memakan 94 token keluaran saat berpikir menyala dan 3 token saat dimatikan — sekitar 30× lebih mahal. Biarkan berpikir menyala hanya kalau kamu memang butuh penalaran bertahap, dan naikkan max_tokens-nya, karena penalaran saja bisa memakan ratusan token.

How to spot it: content empty or short while completion_tokens_details.reasoning_tokens is large. Reasoning tokens are counted inside completion_tokens, so the amount billed is correct — it is the result and your budget that suffer. Cara mengenalinya: content kosong atau pendek padahal completion_tokens_details.reasoning_tokens besar. Token penalaran termasuk di dalam completion_tokens, jadi jumlah yang ditagih sudah benar — yang dirugikan adalah hasil dan anggaranmu.

Models & specificationsModel & spesifikasi

This list is read live from our system, so it always reflects the models currently available. Daftar ini dibaca langsung dari sistem kami, jadi selalu mengikuti model yang tersedia.

ModelModel VersionVersi ContextKonteks Max outputKeluaran maks VisionGambar ToolsTool JSON Best forPaling cocok untuk
loading…memuat…

Context is the total conversation length the model can still remember (your messages plus its replies). Max output is the longest single reply it can produce. Konteks = panjang total percakapan yang masih bisa diingat model (pesanmu + balasannya). Keluaran maks = panjang balasan terpanjang yang bisa dihasilkan sekali jalan.

The name to send in requests is the first column. There is also a deepseek-chat alias for apps that hard-code that name — it is served by the same model as deepseek-flash. Nama model yang dipakai di permintaan adalah kolom pertama. Ada juga alias deepseek-chat untuk aplikasi yang menuliskan nama itu secara tetap — alias itu dilayani model yang sama dengan deepseek-flash.

PricingHarga

Rupiah per 1 million tokens. Two things make prices differ: Satuan rupiah per 1 juta token. Ada dua hal yang membuat harga berbeda-beda:

ModelModel WhenSaat InputMasukan Input (cached)Masukan (cache) OutputKeluaran
loading…memuat…

Peak hours are 08:00–11:00 and 13:00–17:00 WIB (UTC+7), on weekdays. Outside those windows (including Saturday, Sunday, and Chinese public holidays) the rate is half. If you want to spend less, move heavy work outside those windows. Jam sibuk = 08:00–11:00 dan 13:00–17:00 WIB, hari kerja. Di luar itu (termasuk Sabtu, Minggu, dan hari libur Tiongkok) tarifnya setengah harga. Kalau ingin lebih hemat, geser pekerjaan berat ke luar jam tersebut.

A breakdown of the cost of a typical reply is on the main page. Rincian biaya satu balasan biasa juga bisa dilihat di halaman utama.

EndpointsEndpoint

MethodMetode AddressAlamat Key?Kunci? PurposeKegunaan
POST/ai/v1/chat/completions yesya Main endpoint — send a message, get a replyEndpoint utama — kirim pesan, terima balasan
GET/ai/v1/models yesya Models available to youDaftar model yang tersedia untukmu
GET/ai/v1/usage yesya Usage over the last 30 daysPemakaian 30 hari terakhir
GET/ai/v1/balance yesya Remaining credit and daily limitSisa saldo dan batas harian
GET/ai/v1/pricing notidak Price list (this page reads from it)Daftar harga (data di halaman ini berasal dari sini)
GET/ai/health notidak Service statusStatus layanan

Code examplesContoh kode

curl
curl https://zenicat.net/ai/v1/chat/completions -H "Authorization: Bearer sk-ksr-YOURKEY" -H "Content-Type: application/json" -d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Hello!"}]}'
Python (openai library)Python (pustaka openai)
from openai import OpenAI

client = OpenAI(
    base_url="https://zenicat.net/ai/v1",
    api_key="sk-ksr-YOURKEY",
)

r = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
Node.js
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://zenicat.net/ai/v1",
  apiKey: "sk-ksr-YOURKEY",
});

const r = await client.chat.completions.create({
  model: "deepseek-flash",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(r.choices[0].message.content);
PHP
$data = json_encode([
  "model" => "deepseek-flash",
  "messages" => [["role" => "user", "content" => "Hello!"]],
]);

$ch = curl_init("https://zenicat.net/ai/v1/chat/completions");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
  "Authorization: Bearer sk-ksr-YOURKEY",
  "Content-Type: application/json",
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, $data);
$out = json_decode(curl_exec($ch), true);
echo $out["choices"][0]["message"]["content"];
Keeping a conversationMenyimpan percakapan
history = [{"role": "system", "content": "You are a friendly assistant."}]

def ask(message):
    history.append({"role": "user", "content": message})
    r = client.chat.completions.create(
        model="deepseek-flash",
        messages=history,
    )
    answer = r.choices[0].message.content
    history.append({"role": "assistant", "content": answer})
    return answer

Note that the whole history is sent every time, so input tokens grow as the conversation gets longer. To save money, trim the history or put fixed instructions in system so they hit the cache. Perhatikan: seluruh riwayat ikut dikirim setiap kali, jadi token masuknya bertambah seiring percakapan memanjang. Untuk menghemat, pangkas riwayat atau letakkan instruksi tetap di system agar kena cache.

Streaming

For replies that appear word by word, as in a chat app, send "stream": true. The reply arrives as data: chunks and ends with data: [DONE]. Untuk balasan yang muncul kata demi kata seperti di aplikasi chat, kirim "stream": true. Balasan datang sebagai potongan data: dan diakhiri data: [DONE].

stream = client.chat.completions.create(
    model="deepseek-flash",
    messages=[{"role": "user", "content": "Tell me a short story about coffee."}],
    stream=True,
)

for chunk in stream:
    piece = chunk.choices[0].delta.content
    if piece:
        print(piece, end="", flush=True)
Our system always requests the token counts from the provider, including in streaming mode. So usage in this mode is still recorded and billed correctly — nothing is lost. Sistem kami selalu meminta jumlah token dari penyedia, termasuk pada mode streaming. Jadi pemakaian mode ini tetap tercatat dan ditagih dengan benar — tidak ada yang hilang.

Connecting to appsMenyambung ke aplikasi

Almost every AI app has a Base URL or API Base field. Enter https://zenicat.net/ai/v1 and paste your key. Hampir semua aplikasi AI punya kolom Base URL atau API Base. Isi dengan https://zenicat.net/ai/v1 lalu tempel kuncimu.

Chatbox / LobeChat / AnythingLLM

Choose the OpenAI Compatible (or “Custom”) provider, fill in the base URL and API key, then type the model name. Pilih penyedia OpenAI Compatible (atau “Custom”), isi Base URL dan API Key, lalu tulis nama modelnya.

n8n / Make

Use the OpenAI node and change the base URL to ours. Credentials stay a normal API key. Pakai node OpenAI, lalu ubah Base URL ke alamat kami. Kredensialnya tetap API Key biasa.

LangChain / LlamaIndex

Pass the base URL parameter to the OpenAI provider class. Lewatkan parameter Base URL ke kelas penyedia OpenAI-nya.

WhatsApp botBot WhatsApp

Use the WhatsApp library of your choice for messages in and out, with our API as the brain. We recommend using your own WhatsApp number. Pakai pustaka WhatsApp pilihanmu untuk masuk-keluar pesan, dan API kami sebagai otaknya. Sarannya pakai nomor WhatsApp milikmu sendiri.

Limits & quotasBatas & kuota

LimitBatas ValueNilai MeaningArtinya
Requests per minutePermintaan per menit 30 Prevents an accidental loop from draining your creditMencegah pemakaian tak sengaja yang menguras saldo
Requests per 24 hoursPermintaan per 24 jam 500 Reset automatically every dayDireset otomatis setiap hari
Reply length per callPanjang balasan sekali jalan 4,096 Larger requests are trimmed automaticallyPermintaan lebih besar dipangkas otomatis
Conversation contextKonteks percakapan 1,000,000 A limit of the model, not of oursBatas dari model, bukan dari kami

If you need different limits, contact us — they are set per key. Kalau butuh batas berbeda untuk kebutuhanmu, hubungi kami — batas ini diatur per kunci.

Error codesArti kode error

All errors are returned as the same JSON shape, with a readable description: Semua error dikirim dalam bentuk JSON yang sama, dengan keterangan yang bisa dibaca:

{"error": {"message": "saldo habis. Silakan isi ulang untuk melanjutkan.",
          "type": "invalid_request_error", "code": 402}}
CodeKode MeaningArtinya What to doYang harus dilakukan
400 Model name is unavailable, or the request body is not valid JSON Nama model tidak tersedia, atau isi permintaan bukan JSON yang sah Check the model field (see models) and your JSON Periksa model (lihat daftar model) dan bentuk JSON-mu
401 The key was not sent, or is not recognised Kunci tidak dikirim, atau tidak dikenal Make sure the Authorization: Bearer sk-ksr-… header is sent intact Pastikan header Authorization: Bearer sk-ksr-… terkirim utuh
402 Credit exhaustedSaldo habis Top up; requests are accepted again automatically once funded Isi ulang saldo; permintaan diterima lagi otomatis setelah terisi
403 The key or account has been disabled Kunci dimatikan, atau akun dinonaktifkan Contact usHubungi kami
429 You hit the per-minute or daily limit Kena batas per menit atau batas harian Wait a moment, or slow down your requests Tunggu sebentar, atau kurangi kecepatan permintaan
502 A problem on the model provider's side (not your key) Masalah di sisi penyedia model (bukan kuncimu) Retry shortly; failed usage is not billed Coba lagi beberapa saat; pemakaian yang gagal tidak ditagih
Failed usage is not billed. If a request is rejected before the model ever answers, your credit is not reduced. Pemakaian yang gagal tidak ditagih. Kalau permintaan ditolak sebelum sempat dijawab model, saldomu tidak berkurang.

Balance & usageSaldo & pemakaian

Check any time: Cek kapan saja:

curl https://zenicat.net/ai/v1/balance -H "Authorization: Bearer sk-ksr-YOURKEY"

# {"saldo_idr": 19997.86, "batas_harian": 500}

curl https://zenicat.net/ai/v1/usage -H "Authorization: Bearer sk-ksr-YOURKEY"

The /v1/usage response contains the number of requests, input tokens, output tokens, total tokens, and the total billed over the last 30 days. Balasan /v1/usage berisi jumlah permintaan, token masuk, token keluar, total token, dan total yang sudah ditagih dalam 30 hari terakhir.

Credit never expires. There is no subscription and no expiry date — credit is used until it runs out, and you only pay for what you actually use. Saldo tidak hangus. Tidak ada langganan dan tidak ada masa aktif — saldo terpakai sampai habis, dan kamu hanya membayar apa yang benar-benar dipakai.

FAQTanya jawab

Is it exactly the same as the OpenAI API?Apakah sama persis dengan API OpenAI?

The request and response shapes are the same, so the OpenAI libraries work as-is. The differences are the model list and the extra x_ksr field in responses. Bentuk permintaan dan balasannya sama, jadi pustaka OpenAI bisa langsung dipakai. Yang berbeda hanya daftar model dan tambahan x_ksr pada balasan.

Why are the model names different from what I know?Kenapa nama modelnya berbeda dari yang saya kenal?

We follow the official names from the model provider. Older aliases such as deepseek-chat are still accepted so existing apps do not need changes. Kami mengikuti nama resmi dari penyedia model. Alias lama seperti deepseek-chat tetap kami terima supaya aplikasi yang sudah ada tidak perlu diubah.

Are my conversations stored?Apakah percakapan saya disimpan?

We store only the token counts and the cost, for billing. Not the content of your messages. The conversation content is processed by the model provider under their own policy. Kami hanya menyimpan jumlah token dan biaya untuk keperluan penagihan — bukan isi pesanmu. Isi percakapan diproses oleh penyedia model sesuai kebijakan mereka.

Can one key be used by several apps?Bisa dipakai untuk banyak aplikasi dengan satu kunci?

Yes. Usage is pooled into one balance, and you can request separate keys if you prefer separate reporting. Bisa. Pemakaiannya digabung ke satu saldo, dan kamu bisa minta kunci terpisah kalau ingin memisahkan pencatatannya.

What if I want to stop?Bagaimana kalau saya ingin berhenti?

There is no contract. Just stop using it; any remaining credit is still yours. Tidak ada ikatan. Tinggal berhenti memakai; saldo yang tersisa tetap milikmu.

‹ Zenicat