Skip to content
Pareto Inference

Served on our GPUs. Pay per token. Connect your coding agent, router, or app.

Available on

OpenCodeMergeNanoGPT

Supported by

Recent news

Pareto raises $4M in seed funding to expand its GPU fleet.

Growth

Up 143x from about 3B a day in early September.

Performance

Live service metrics for GLM 5.3 Flash.

GLM 5.3 FlashChecking sample time
Uptime
100.00%
1h window
Time to first response
1.63 s
1h · P50
Output speed
94.8 tok/s
1h · P50
Cached input
97.4%
1h window
Uptime history12h

Pricing

Add credits with a card. All token usage comes from one prepaid balance.

GLM 5.3 Flash · USD per 1M tokens
Token typeOpenRouterParetoPareto Promo
InputTokens you send$0.15$0.09$0.06
OutputTokens the model returns$0.50$0.30$0.20
Cached inputPrompt tokens reused from cache$0.03$0.018$0.012
Pricing details

40% less than OpenRouter.

Setup

Connect your coding agent or chat app, bring Pareto to a router, or use our API directly. These are a few ways to get started.

Direct API
Coding agents
and more
Routers
and more
Chat apps

Install the OpenAI Python SDK and set PARETO_API_KEY in your environment.

example.pyPython
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.paretoinference.com/v1",
    api_key=os.environ["PARETO_API_KEY"],
)

response = client.chat.completions.create(
    model="z-ai/glm-5.3-flash",
    messages=[{"role": "user", "content": "Hello, Pareto."}],
    stream=True,
)
for chunk in response:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
Open setup guide