DocumentationDokumentasi
Everything you need to use Zenicat AI Gateway: how tokens are counted, the specifications of each model, endpoints, code examples, and what every error code means. Semua yang perlu kamu tahu untuk memakai Zenicat AI Gateway: cara menghitung token, spesifikasi tiap model, endpoint, contoh kode, dan arti setiap kode error.
Quick startMulai cepat
This API follows the OpenAI standard, so the libraries and apps you already use will work — you only change the base URL. API ini mengikuti standar OpenAI, jadi pustaka dan aplikasi yang sudah kamu pakai bisa langsung dipakai — cukup ganti alamat dasarnya.
| Base URLAlamat dasar | https://zenicat.net/ai/v1 |
|---|---|
| HeaderHeader | Authorization: Bearer sk-ksr-… |
| ModelsModel | loading… |
What is a tokenApa itu token
A token is the smallest unit of text the model recognises — it can be a word, a number, or even a punctuation mark. Tokens are what you are billed for, not the number of messages or characters. Token adalah satuan terkecil teks yang dikenali model — bisa satu kata, satu angka, bahkan satu tanda baca. Token inilah yang dihitung sebagai biaya, bukan jumlah pesan atau jumlah huruf.
In other words, one long message does not cost the same as one short message. What counts is input tokens (your message) and output tokens (the model's reply). Artinya: mengirim satu pesan panjang tidak sama biayanya dengan satu pesan pendek. Yang dihitung adalah token masuk (pesanmu) dan token keluar (balasan model).
Rough character-to-token ratiosPerkiraan kasar karakter ke token
Official figures from the model provider, for Latin-script text: Angka resmi dari penyedia model, untuk teks berhuruf Latin:
- 1 Latin character ≈ 0.3 token — so 1 token ≈ 3–4 characters 1 karakter Latin ≈ 0,3 token — jadi 1 token ≈ 3–4 karakter
- One English word averages ≈ 1.3 tokens Satu kata bahasa Inggris rata-rata ≈ 1,3 token
- Code is denser: ≈ 8–12 tokens per realistic line Kode program lebih padat: ≈ 8–12 token per baris yang realistis
The big pictureGambaran besarnya
| ContentIsi | Approx. tokensPerkiraan token |
|---|---|
| One short chat message (± 50 words)Satu pesan chat pendek (± 50 kata) | ± 70 |
| One page of a document (± 500 words)Satu halaman dokumen (± 500 kata) | ± 700–900 |
| A long article (± 2,000 words)Artikel panjang (± 2.000 kata) | ± 2,800–3,600 |
| A whole thin novel (± 40,000 words)Seluruh isi novel tipis (± 40.000 kata) | ± 55,000 |
| 1,000,000 tokens1.000.000 token | ± 750,000 words± 750.000 kata |
The figures above are for planning only. The amount billed is always the real number. Perkiraan di atas hanya untuk mengira-ngira. Angka yang ditagih selalu angka sebenarnya.
How to know exactly: read usageCara tahu pasti: baca usage
Every response includes the actual token counts. This is the authoritative source, not an estimate: Setiap balasan menyertakan jumlah token sebenarnya. Ini sumber yang benar, bukan perkiraan:
{
"choices": [ ... ],
"usage": {
"prompt_tokens": 52, // your message
"completion_tokens": 118, // the reply
"total_tokens": 170,
"prompt_cache_hit_tokens": 32,
"prompt_cache_miss_tokens": 20
},
"x_ksr": { // added by us
"dipotong_idr": 2.14, // what was actually deducted
"saldo_idr": 19997.86
}
}
The x_ksr field answers the most common question:
“how much did this call cost?” — without you having to work it out.
Field x_ksr menjawab pertanyaan yang paling sering muncul:
“berapakah biaya panggilan ini?” — tanpa perlu menghitung sendiri.
prompt_cache_hit_tokens is reported separately from
prompt_cache_miss_tokens. Putting fixed instructions in
system saves a lot for routine usage.
Input yang di-cache jauh lebih murah. Kalau bagian awal pesanmu sama
dengan permintaan sebelumnya (misalnya system prompt yang tetap), model mengenali dan
menghitungnya 50× lebih murah. Inilah sebabnya
prompt_cache_hit_tokens terpisah dari prompt_cache_miss_tokens.
Menaruh instruksi tetap di system menghemat banyak untuk pemakaian rutin.
Thinking mode is on by default — and you pay for itMode berpikir menyala secara bawaan — dan kamu membayarnya
This is the single most expensive surprise on this API. The model
reasons before answering by default, and every reasoning token is billed as an
output token — taken from the same max_tokens budget as your
answer.
Ini kejutan paling mahal di API ini. Model berpikir dulu sebelum
menjawab secara bawaan, dan setiap token penalaran ditagih sebagai token keluaran
— diambil dari jatah max_tokens yang sama dengan jawabanmu.
Measured result: with max_tokens: 60 and thinking left on, a
simple request returns empty content with
finish_reason: "length" — all 60 tokens went to reasoning. The same
request with thinking off answers in 3 tokens.
Hasil pengukuran: dengan max_tokens: 60 dan berpikir
dibiarkan menyala, permintaan sederhana mengembalikan content kosong
dengan finish_reason: "length" — 60 token habis dipakai penalaran.
Permintaan yang sama dengan berpikir dimatikan menjawab dalam 3 token.
To turn it off, add one field to your request: Untuk mematikannya, tambahkan satu field ke permintaanmu:
{
"model": "deepseek-flash",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 60,
"thinking": {"type": "disabled"}
}
max_tokens
accordingly, because reasoning alone can take hundreds of tokens.
Kenapa penting? Jawaban satu kalimat memakan 94 token keluaran
saat berpikir menyala dan 3 token saat dimatikan — sekitar 30× lebih
mahal. Biarkan berpikir menyala hanya kalau kamu memang butuh penalaran bertahap, dan
naikkan max_tokens-nya, karena penalaran saja bisa memakan ratusan
token.
How to spot it: content empty or short while
completion_tokens_details.reasoning_tokens is large. Reasoning tokens
are counted inside completion_tokens, so the amount billed is correct
— it is the result and your budget that suffer.
Cara mengenalinya: content kosong atau pendek padahal
completion_tokens_details.reasoning_tokens besar. Token penalaran
termasuk di dalam completion_tokens, jadi jumlah yang ditagih sudah
benar — yang dirugikan adalah hasil dan anggaranmu.
Models & specificationsModel & spesifikasi
This list is read live from our system, so it always reflects the models currently available. Daftar ini dibaca langsung dari sistem kami, jadi selalu mengikuti model yang tersedia.
| ModelModel | VersionVersi | ContextKonteks | Max outputKeluaran maks | VisionGambar | ToolsTool | JSON | Best forPaling cocok untuk |
|---|---|---|---|---|---|---|---|
| loading…memuat… | |||||||
Context is the total conversation length the model can still remember (your messages plus its replies). Max output is the longest single reply it can produce. Konteks = panjang total percakapan yang masih bisa diingat model (pesanmu + balasannya). Keluaran maks = panjang balasan terpanjang yang bisa dihasilkan sekali jalan.
deepseek-chat alias for apps that hard-code that name — it is served by the
same model as deepseek-flash.
Nama model yang dipakai di permintaan adalah kolom pertama. Ada juga alias
deepseek-chat untuk aplikasi yang menuliskan nama itu secara tetap — alias itu
dilayani model yang sama dengan deepseek-flash.
PricingHarga
Rupiah per 1 million tokens. Two things make prices differ: Satuan rupiah per 1 juta token. Ada dua hal yang membuat harga berbeda-beda:
- Input or output — output tokens are always more expensive than input. Masuk atau keluar — token keluaran selalu lebih mahal daripada masukan.
- Peak hours — at certain hours the rate is double. Jam sibuk — di jam tertentu tarifnya dua kali lipat.
| ModelModel | WhenSaat | InputMasukan | Input (cached)Masukan (cache) | OutputKeluaran |
|---|---|---|---|---|
| loading…memuat… | ||||
A breakdown of the cost of a typical reply is on the main page. Rincian biaya satu balasan biasa juga bisa dilihat di halaman utama.
EndpointsEndpoint
| MethodMetode | AddressAlamat | Key?Kunci? | PurposeKegunaan |
|---|---|---|---|
| POST | /ai/v1/chat/completions |
yesya | Main endpoint — send a message, get a replyEndpoint utama — kirim pesan, terima balasan |
| GET | /ai/v1/models |
yesya | Models available to youDaftar model yang tersedia untukmu |
| GET | /ai/v1/usage |
yesya | Usage over the last 30 daysPemakaian 30 hari terakhir |
| GET | /ai/v1/balance |
yesya | Remaining credit and daily limitSisa saldo dan batas harian |
| GET | /ai/v1/pricing |
notidak | Price list (this page reads from it)Daftar harga (data di halaman ini berasal dari sini) |
| GET | /ai/health |
notidak | Service statusStatus layanan |
Code examplesContoh kode
curl https://zenicat.net/ai/v1/chat/completions -H "Authorization: Bearer sk-ksr-YOURKEY" -H "Content-Type: application/json" -d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Hello!"}]}'
from openai import OpenAI
client = OpenAI(
base_url="https://zenicat.net/ai/v1",
api_key="sk-ksr-YOURKEY",
)
r = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://zenicat.net/ai/v1",
apiKey: "sk-ksr-YOURKEY",
});
const r = await client.chat.completions.create({
model: "deepseek-flash",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(r.choices[0].message.content);
$data = json_encode([
"model" => "deepseek-flash",
"messages" => [["role" => "user", "content" => "Hello!"]],
]);
$ch = curl_init("https://zenicat.net/ai/v1/chat/completions");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
"Authorization: Bearer sk-ksr-YOURKEY",
"Content-Type: application/json",
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, $data);
$out = json_decode(curl_exec($ch), true);
echo $out["choices"][0]["message"]["content"];
history = [{"role": "system", "content": "You are a friendly assistant."}]
def ask(message):
history.append({"role": "user", "content": message})
r = client.chat.completions.create(
model="deepseek-flash",
messages=history,
)
answer = r.choices[0].message.content
history.append({"role": "assistant", "content": answer})
return answer
Note that the whole history is sent every time, so input tokens grow as the
conversation gets longer. To save money, trim the history or put fixed instructions in
system so they hit the cache.
Perhatikan: seluruh riwayat ikut dikirim setiap kali, jadi token masuknya
bertambah seiring percakapan memanjang. Untuk menghemat, pangkas riwayat atau letakkan instruksi
tetap di system agar kena cache.
Streaming
For replies that appear word by word, as in a chat app, send
"stream": true. The reply arrives as data: chunks and ends with
data: [DONE].
Untuk balasan yang muncul kata demi kata seperti di aplikasi chat, kirim
"stream": true. Balasan datang sebagai potongan data: dan diakhiri
data: [DONE].
stream = client.chat.completions.create(
model="deepseek-flash",
messages=[{"role": "user", "content": "Tell me a short story about coffee."}],
stream=True,
)
for chunk in stream:
piece = chunk.choices[0].delta.content
if piece:
print(piece, end="", flush=True)
Connecting to appsMenyambung ke aplikasi
Almost every AI app has a Base URL or API Base field. Enter
https://zenicat.net/ai/v1 and paste your key.
Hampir semua aplikasi AI punya kolom Base URL atau API Base. Isi
dengan https://zenicat.net/ai/v1 lalu tempel kuncimu.
Chatbox / LobeChat / AnythingLLM
Choose the OpenAI Compatible (or “Custom”) provider, fill in the base URL and API key, then type the model name. Pilih penyedia OpenAI Compatible (atau “Custom”), isi Base URL dan API Key, lalu tulis nama modelnya.
n8n / Make
Use the OpenAI node and change the base URL to ours. Credentials stay a normal API key. Pakai node OpenAI, lalu ubah Base URL ke alamat kami. Kredensialnya tetap API Key biasa.
LangChain / LlamaIndex
Pass the base URL parameter to the OpenAI provider class. Lewatkan parameter Base URL ke kelas penyedia OpenAI-nya.
WhatsApp botBot WhatsApp
Use the WhatsApp library of your choice for messages in and out, with our API as the brain. We recommend using your own WhatsApp number. Pakai pustaka WhatsApp pilihanmu untuk masuk-keluar pesan, dan API kami sebagai otaknya. Sarannya pakai nomor WhatsApp milikmu sendiri.
Limits & quotasBatas & kuota
| LimitBatas | ValueNilai | MeaningArtinya |
|---|---|---|
| Requests per minutePermintaan per menit | 30 | Prevents an accidental loop from draining your creditMencegah pemakaian tak sengaja yang menguras saldo |
| Requests per 24 hoursPermintaan per 24 jam | 500 | Reset automatically every dayDireset otomatis setiap hari |
| Reply length per callPanjang balasan sekali jalan | 4,096 | Larger requests are trimmed automaticallyPermintaan lebih besar dipangkas otomatis |
| Conversation contextKonteks percakapan | 1,000,000 | A limit of the model, not of oursBatas dari model, bukan dari kami |
If you need different limits, contact us — they are set per key. Kalau butuh batas berbeda untuk kebutuhanmu, hubungi kami — batas ini diatur per kunci.
Error codesArti kode error
All errors are returned as the same JSON shape, with a readable description: Semua error dikirim dalam bentuk JSON yang sama, dengan keterangan yang bisa dibaca:
{"error": {"message": "saldo habis. Silakan isi ulang untuk melanjutkan.",
"type": "invalid_request_error", "code": 402}}
| CodeKode | MeaningArtinya | What to doYang harus dilakukan |
|---|---|---|
| 400 | Model name is unavailable, or the request body is not valid JSON Nama model tidak tersedia, atau isi permintaan bukan JSON yang sah | Check the model field (see models) and your JSON
Periksa model (lihat daftar model) dan bentuk JSON-mu |
| 401 | The key was not sent, or is not recognised Kunci tidak dikirim, atau tidak dikenal | Make sure the Authorization: Bearer sk-ksr-… header is sent intact
Pastikan header Authorization: Bearer sk-ksr-… terkirim utuh |
| 402 | Credit exhaustedSaldo habis | Top up; requests are accepted again automatically once funded Isi ulang saldo; permintaan diterima lagi otomatis setelah terisi |
| 403 | The key or account has been disabled Kunci dimatikan, atau akun dinonaktifkan | Contact usHubungi kami |
| 429 | You hit the per-minute or daily limit Kena batas per menit atau batas harian | Wait a moment, or slow down your requests Tunggu sebentar, atau kurangi kecepatan permintaan |
| 502 | A problem on the model provider's side (not your key) Masalah di sisi penyedia model (bukan kuncimu) | Retry shortly; failed usage is not billed Coba lagi beberapa saat; pemakaian yang gagal tidak ditagih |
Balance & usageSaldo & pemakaian
Check any time: Cek kapan saja:
curl https://zenicat.net/ai/v1/balance -H "Authorization: Bearer sk-ksr-YOURKEY"
# {"saldo_idr": 19997.86, "batas_harian": 500}
curl https://zenicat.net/ai/v1/usage -H "Authorization: Bearer sk-ksr-YOURKEY"
The /v1/usage response contains the number of requests, input
tokens, output tokens, total tokens, and the total billed over the last 30 days.
Balasan /v1/usage berisi jumlah permintaan, token masuk, token
keluar, total token, dan total yang sudah ditagih dalam 30 hari terakhir.
Credit never expires. There is no subscription and no expiry date — credit is used until it runs out, and you only pay for what you actually use. Saldo tidak hangus. Tidak ada langganan dan tidak ada masa aktif — saldo terpakai sampai habis, dan kamu hanya membayar apa yang benar-benar dipakai.
FAQTanya jawab
Is it exactly the same as the OpenAI API?Apakah sama persis dengan API OpenAI?
The request and response shapes are the same, so the OpenAI libraries work as-is.
The differences are the model list and the extra x_ksr field in responses.
Bentuk permintaan dan balasannya sama, jadi pustaka OpenAI bisa langsung dipakai.
Yang berbeda hanya daftar model dan tambahan x_ksr pada balasan.
Why are the model names different from what I know?Kenapa nama modelnya berbeda dari yang saya kenal?
We follow the official names from the model provider. Older aliases such as
deepseek-chat are still accepted so existing apps do not need changes.
Kami mengikuti nama resmi dari penyedia model. Alias lama seperti
deepseek-chat tetap kami terima supaya aplikasi yang sudah ada tidak perlu
diubah.
Are my conversations stored?Apakah percakapan saya disimpan?
We store only the token counts and the cost, for billing. Not the content of your messages. The conversation content is processed by the model provider under their own policy. Kami hanya menyimpan jumlah token dan biaya untuk keperluan penagihan — bukan isi pesanmu. Isi percakapan diproses oleh penyedia model sesuai kebijakan mereka.
Can one key be used by several apps?Bisa dipakai untuk banyak aplikasi dengan satu kunci?
Yes. Usage is pooled into one balance, and you can request separate keys if you prefer separate reporting. Bisa. Pemakaiannya digabung ke satu saldo, dan kamu bisa minta kunci terpisah kalau ingin memisahkan pencatatannya.
What if I want to stop?Bagaimana kalau saya ingin berhenti?
There is no contract. Just stop using it; any remaining credit is still yours. Tidak ada ikatan. Tinggal berhenti memakai; saldo yang tersisa tetap milikmu.