Growth
Up 143x from about 3B a day in early September.
Performance
Live service metrics for GLM 5.3 Flash.
- Uptime
- 100.00% 1h window
- Time to first response
- 1.63 s 1h · P50
- Output speed
- 94.8 tok/s 1h · P50
- Cached input
- 97.4% 1h window
Pricing
Add credits with a card. All token usage comes from one prepaid balance.
| Token type | OpenRouter | Pareto | Pareto Promo |
|---|---|---|---|
| InputTokens you send | $0.15 | $0.09 | $0.06 |
| OutputTokens the model returns | $0.50 | $0.30 | $0.20 |
| Cached inputPrompt tokens reused from cache | $0.03 | $0.018 | $0.012 |
40% less than OpenRouter.
Setup
Connect your coding agent or chat app, bring Pareto to a router, or use our API directly. These are a few ways to get started.
Install the OpenAI Python SDK and set PARETO_API_KEY in your environment.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.paretoinference.com/v1",
api_key=os.environ["PARETO_API_KEY"],
)
response = client.chat.completions.create(
model="z-ai/glm-5.3-flash",
messages=[{"role": "user", "content": "Hello, Pareto."}],
stream=True,
)
for chunk in response:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")Install the OpenAI SDK and set PARETO_API_KEY in your server environment.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.PARETO_API_KEY,
baseURL: "https://api.paretoinference.com/v1",
});
const response = await client.chat.completions.create({
model: "z-ai/glm-5.3-flash",
messages: [{ role: "user", content: "Hello, Pareto." }],
stream: true,
});
for await (const chunk of response) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Set PARETO_API_KEY for OpenClaw and merge this into ~/.openclaw/openclaw.json. Start a new session.
{
"agents": {
"defaults": {
"model": {
"primary": "pareto/z-ai/glm-5.3-flash"
}
}
},
"models": {
"mode": "merge",
"providers": {
"pareto": {
"baseUrl": "https://api.paretoinference.com/v1",
"apiKey": "${PARETO_API_KEY}",
"api": "openai-completions",
"models": [
{
"id": "z-ai/glm-5.3-flash",
"name": "GLM 5.3 Flash (Pareto)"
}
]
}
}
}
}Run hermes model outside an active chat. Choose Custom endpoint, then enter these values.
- Command
hermes model- Endpoint
- Custom endpoint (self-hosted / VLLM / etc.)
- API compatibility
- Chat Completions or Auto-detect
- Base URL
https://api.paretoinference.com/v1- API key
- Your Pareto API key
- Model
z-ai/glm-5.3-flash
Set PARETO_API_KEY and merge this into ~/.config/opencode/opencode.json. Keep your existing providers and settings, then start OpenCode.
{
"$schema": "https://opencode.ai/config.json",
"model": "pareto/z-ai/glm-5.3-flash",
"small_model": "pareto/z-ai/glm-5.3-flash",
"provider": {
"pareto": {
"npm": "@ai-sdk/openai-compatible",
"name": "Pareto Inference",
"options": {
"baseURL": "https://api.paretoinference.com/v1",
"apiKey": "{env:PARETO_API_KEY}"
},
"models": {
"z-ai/glm-5.3-flash": {
"name": "GLM 5.3 Flash"
}
}
}
}
}Uses Claude Code Router. Back up ~/.claude/settings.json, save this as ~/.claude-code-router/config.json, and run ccr start. Set ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL to pareto/z-ai/glm-5.3-flash.
{
"HOST": "127.0.0.1",
"PORT": 3456,
"APIKEY": "your-local-ccr-key",
"Providers": [
{
"name": "pareto",
"api_base_url": "https://api.paretoinference.com/v1/chat/completions",
"api_key": "$PARETO_API_KEY",
"models": [
"z-ai/glm-5.3-flash"
]
}
]
}Uses LiteLLM. Run LiteLLM with this model, then add a Codex provider for http://127.0.0.1:4000/v1 with wire_api = "responses" and model pareto-flash.
model_list:
- model_name: pareto-flash
litellm_params:
model: openai/z-ai/glm-5.3-flash
api_base: https://api.paretoinference.com/v1
api_key: os.environ/PARETO_API_KEY
use_chat_completions_api: true
additional_drop_params: ["web_search_options"]Set PARETO_API_KEY and merge this provider into ~/.pi/agent/models.json. Select it with /model.
{
"providers": {
"pareto": {
"baseUrl": "https://api.paretoinference.com/v1",
"api": "openai-completions",
"apiKey": "$PARETO_API_KEY",
"models": [
{
"id": "z-ai/glm-5.3-flash",
"name": "GLM 5.3 Flash (Pareto)"
}
]
}
}
}Open Cline settings, select OpenAI Compatible, and save these connection values.
- API type
- OpenAI Compatible
- Base URL
https://api.paretoinference.com/v1- API key
- Your Pareto API key
- Model
z-ai/glm-5.3-flash
Set PARETO_API_KEY and save this as ~/.config/goose/custom_providers/pareto.json. Set GOOSE_PROVIDER=pareto and GOOSE_MODEL, then run Goose with the developer extension.
{
"name": "pareto",
"engine": "openai",
"display_name": "Pareto Inference",
"api_key_env": "PARETO_API_KEY",
"base_url": "https://api.paretoinference.com/v1/chat/completions",
"models": [
{
"name": "z-ai/glm-5.3-flash"
}
],
"requires_auth": true,
"supports_streaming": true
}Set PARETO_API_KEY, then start Aider in your repository. Keep the openai/ prefix and ignore the model names Aider suggests.
export OPENAI_API_BASE=https://api.paretoinference.com/v1
export OPENAI_API_KEY="$PARETO_API_KEY"
aider --model openai/z-ai/glm-5.3-flash --no-show-model-warningsPut PARETO_API_KEY in ~/.continue/.env and merge this into ~/.continue/config.yaml. The editor extension does not read shell variables.
name: Pareto Inference
version: 1.0.0
schema: v1
models:
- name: GLM 5.3 Flash (Pareto)
provider: openai
model: z-ai/glm-5.3-flash
apiBase: https://api.paretoinference.com/v1
apiKey: ${{ secrets.PARETO_API_KEY }}
roles:
- chat
- edit
- apply
capabilities:
- tool_useSet PARETO_API_KEY and merge this into ~/.omp/agent/models.yml. Start Oh My Pi with omp --model pareto/z-ai/glm-5.3-flash.
providers:
pareto:
baseUrl: https://api.paretoinference.com/v1
api: openai-completions
apiKey: PARETO_API_KEY
authHeader: true
models:
- id: z-ai/glm-5.3-flash
name: GLM 5.3 Flash (Pareto)Set PARETO_API_KEY and add this to ~/.config/crush/crushrc. Select pareto/z-ai/glm-5.3-flash in Crush. Use openai-compat, not openai.
provider add pareto --type openai-compat \
--name "Pareto Inference" \
--base-url "https://api.paretoinference.com/v1" \
--api-key "$PARETO_API_KEY"
model add pareto/z-ai/glm-5.3-flash --name "GLM 5.3 Flash"Set PARETO_API_KEY, then start Qwen Code. Use --auth-type openai, not openai-responses.
export OPENAI_API_KEY="$PARETO_API_KEY"
export OPENAI_BASE_URL=https://api.paretoinference.com/v1
export OPENAI_MODEL=z-ai/glm-5.3-flash
qwen --auth-type openaiOpen Settings → Providers → Custom provider. Save Pareto and select the model. Ask Pareto for the current context and output limits; Kilo needs accurate limits for automatic context compaction.
- Provider ID
pareto- Name
- Pareto
- API type
- OpenAI Compatible
- Base URL
https://api.paretoinference.com/v1- API key
- Your Pareto API key
- Model
z-ai/glm-5.3-flash
Settings → Model settings → Add provider → Create custom provider. Use Chat completions, the Pareto base URL, and model z-ai/glm-5.3-flash. Clear Image, Video and PDF, and paste this into the model's Reasoning parameter mapping. The default mapping sends fields Pareto rejects.
{
"reasoning_effort": reasoningLevel == "disabled" ? "none" : reasoningLevel == "enabled" ? "high" : reasoningLevel
}Add a custom OpenAI-compatible provider, then save a connection with your Pareto key.
- Name
- Pareto
- Prefix
pareto- API type
- Chat Completions
- Base URL
https://api.paretoinference.com/v1- Model
z-ai/glm-5.3-flash
Add this model to your existing config. Keep your proxy access settings.
model_list:
- model_name: pareto-flash
litellm_params:
model: openai/z-ai/glm-5.3-flash
api_base: https://api.paretoinference.com/v1
api_key: os.environ/PARETO_API_KEYAdd this entry to config.yaml and restart CLIProxyAPI. The base URL keeps /v1. Apps use a key from api-keys, not your Pareto key.
openai-compatibility:
- name: "pareto"
base-url: "https://api.paretoinference.com/v1"
api-key-entries:
- api-key: "YOUR_PARETO_API_KEY"
models:
- name: "z-ai/glm-5.3-flash"Open Channels → Add channel, save these values, and select Test. The base URL has no /v1. Apps use a new-api token, not your Pareto key.
- Type
- OpenAI
- Name
- Pareto
- Base URL
https://api.paretoinference.com- Key
- Your Pareto API key
- Models
z-ai/glm-5.3-flash
Open Channels (渠道) → Add channel, save these values, and select Test. The base URL has no /v1. Apps use a One API token, not your Pareto key.
- Type
- OpenAI
- Name
pareto- Base URL
https://api.paretoinference.com- Key
- Your Pareto API key
- Models
z-ai/glm-5.3-flash
Add a custom OpenAI-compatible provider, then save a connection with your Pareto key.
- Name
- Pareto
- Prefix
pareto- API type
- Chat Completions
- Base URL
https://api.paretoinference.com/v1- Model
z-ai/glm-5.3-flash
Run the gateway from the Docker image; the npx build breaks streaming. Apps send your Pareto key and both x-portkey headers.
docker run -d -p 127.0.0.1:8787:8787 portkeyai/gateway:latest
curl http://localhost:8787/v1/chat/completions \
-H "Authorization: Bearer $PARETO_API_KEY" \
-H "Content-Type: application/json" \
-H "x-portkey-provider: openai" \
-H "x-portkey-custom-host: https://api.paretoinference.com/v1" \
-d '{"model": "z-ai/glm-5.3-flash", "messages": [{"role": "user", "content": "Say hello."}], "stream": true}'Set PARETO_API_KEY and save this as config.json in the Bifrost app directory. The base URL has no /v1. Apps use the model pareto/z-ai/glm-5.3-flash.
{
"providers": {
"pareto": {
"keys": [
{
"name": "pareto",
"value": "env.PARETO_API_KEY",
"models": [
"*"
],
"weight": 1
}
],
"network_config": {
"base_url": "https://api.paretoinference.com",
"default_request_timeout_in_seconds": 300
},
"custom_provider_config": {
"base_provider_type": "openai",
"allowed_requests": {
"chat_completion": true,
"chat_completion_stream": true,
"list_models": true
}
}
}
}
}Open Groups → Add group, save these values, then create an access key for your apps. The upstream URL keeps /v1.
- Channel
- OpenAI Compatible
- Upstream URL
https://api.paretoinference.com/v1- API keys
- Your Pareto API key
- Models
z-ai/glm-5.3-flash
Open Channels → Add Channel and save these values. New channels start disabled, so enable the channel. Apps use an AxonHub API key.
- Provider
- OpenAI
- Name
pareto- Base URL
https://api.paretoinference.com/v1- API key
- Your Pareto API key
- Models
z-ai/glm-5.3-flash
Open Admin Panel → Settings → Connections and add an OpenAI API connection. Keep API Type on Chat Completions.
- URL
https://api.paretoinference.com/v1- Key
- Your Pareto API key
- API type
- Chat Completions
- Model
z-ai/glm-5.3-flash
Run NextChat with your key on the server. BASE_URL has no /v1, and CUSTOM_MODELS adds the Pareto model.
docker run -d -p 3000:3000 \
-e OPENAI_API_KEY="$PARETO_API_KEY" \
-e BASE_URL="https://api.paretoinference.com" \
-e CUSTOM_MODELS="-all,+z-ai/glm-5.3-flash@OpenAI" \
-e DEFAULT_MODEL="z-ai/glm-5.3-flash@OpenAI" \
yidadaa/chatgpt-next-web:v2.16.1Run AnythingLLM with the Generic OpenAI provider. Keep Max Tokens at 8192 or more; reasoning counts toward it.
docker run -d -p 3001:3001 --cap-add SYS_ADMIN \
-v anythingllm_storage:/app/server/storage \
-e STORAGE_DIR=/app/server/storage \
-e LLM_PROVIDER=generic-openai \
-e GENERIC_OPEN_AI_BASE_PATH=https://api.paretoinference.com/v1 \
-e GENERIC_OPEN_AI_API_KEY="$PARETO_API_KEY" \
-e GENERIC_OPEN_AI_MODEL_PREF=z-ai/glm-5.3-flash \
-e GENERIC_OPEN_AI_MAX_TOKENS=8192 \
mintplexlabs/anythingllm:1.16.2Open Settings → Model Provider → Add Provider, then Sync models. Also set the Quick Model to Pareto, or titles go to Cherry's own service.
- Name
- Pareto Inference
- API key
- Your Pareto API key
- OpenAI URL
https://api.paretoinference.com- Model
z-ai/glm-5.3-flash
Add this endpoint to librechat.yaml, set PARETO_API_KEY in .env, and restart. Keep dropParams: Pareto rejects the user field.
endpoints:
custom:
- name: "Pareto"
apiKey: ${PARETO_API_KEY}
baseURL: "https://api.paretoinference.com/v1"
models:
default: ["z-ai/glm-5.3-flash"]
fetch: true
titleConvo: true
titleModel: "current_model"
dropParams: ["user"]Open Settings → Model Provider → Add → Add Custom Provider, then Fetch models. Keep API Path empty.
- API mode
- OpenAI API Compatible
- API key
- Your Pareto API key
- API host
https://api.paretoinference.com/v1- Model
z-ai/glm-5.3-flash
Open API Connections, enter these values, and select Connect. Raise Max Response Length from 300, or reasoning uses up the answer.
- Source
- Custom (OpenAI-compatible)
- Endpoint
https://api.paretoinference.com/v1- API key
- Your Pareto API key
- Model
z-ai/glm-5.3-flash- Response
- Max Response Length 2048 or more


