Moonshot AI
Operational· 100% uptime, 7 days

Kimi K3 API

The Kimi K3 API for long-horizon coding and agents, 58% below Moonshot's list price

Price per 1M tokens58%off
Input
$1.27$3.00
Output
$6.37$15.00
Cached input
$0.127
Model IDkimi-k3
chat.py
from openai import OpenAI
client = OpenAI(
base_url="https://api.zurelay.com/v1",
api_key="YOUR_ZURELAY_KEY",
)
stream = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Explain quicksort in two sentences."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")

Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.

Context window
1M tokens
Input
Text, Image
Output
Text
Released
Jul 16, 2026
Uptime, 7 days
100%

Pricing

Kimi K3 API pricing. A fraction of the list price.

Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.

RatezurelayMoonshot AI listYou save
Input
per 1M tokens
$1.27$3.0058%off
Output
per 1M tokens
$6.37$15.0058%off
Cached input
per 1M tokens
$0.127——

One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.

Savings calculator

Move the sliders to your monthly usage.

Input tokens per month50M
Output tokens per month10M
zurelay
$127.20
Moonshot AI list price
$300.00
You save every month$172.80

$2,074 a year

Overview

What is Kimi K3?

Kimi K3 is Moonshot AI's flagship model: 2.8 trillion parameters, native vision and a 1M-token context window, built for long-horizon coding and knowledge work. The zurelay Kimi K3 API serves the same model through one OpenAI-compatible endpoint at 58% below Moonshot's list price.

Kimi K3 is Moonshot AI's most capable model, released on July 16, 2026, with the full weights following on Hugging Face on July 27. The Kimi K3 API reads text and images, returns text and holds 1,048,576 tokens of context. Moonshot's max_completion_tokens defaults to 131,072 and can be set as high as 1,048,576.

It is a Mixture-of-Experts model with 2.8T total parameters and 104B active per token, routing each token to 16 of 896 experts. Moonshot built it on Kimi Delta Attention and Attention Residuals and reports about 2.5 times better scaling efficiency than Kimi K2. The weights use MXFP4 with MXFP8 activations and ship under the Kimi K3 License, which allows commercial use; very large Model-as-a-Service businesses need a separate agreement with Moonshot.

Moonshot built Kimi K3 for long-horizon coding, knowledge work and reasoning. It can run long engineering sessions with little supervision, move through large repositories and drive terminal tools, and it uses screenshots to refine frontend, game and CAD work. Moonshot says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, and reports it ahead of the other models it tested.

Thinking is always on. Moonshot's API takes a reasoning_effort of low, high or max, with max as the default, and fixes temperature at 1.0 and top_p at 0.95. The model supports tool calling, including tool_choice required, plus JSON schema output and automatic context caching. On zurelay you call it at https://api.zurelay.com/v1 with the model ID kimi-k3, for $1.27 per 1M input tokens and $6.37 per 1M output tokens against Moonshot's list price of $3.00 and $15.00.

Strengths

Where Kimi K3 shines. And what teams build with it.

01

Long-horizon coding

Moonshot built K3 to run long engineering sessions with little supervision, navigate large repositories and orchestrate terminal tools.

02

Vision in the loop

K3 reads screenshots natively, so it can check rendered output and refine frontend, game or CAD code in the same session.

03

Performance-critical code

In Moonshot's GPU kernel optimization tests, K3 performed competitively with Claude Fable 5 and well ahead of GPT-5.6 Sol.

04

Research and knowledge work

Moonshot reports consistent gains on its internal knowledge-work evaluations, from multi-source research to reports with interactive charts.

Use cases

  • Coding agents

    Run multi-hour refactors, feature builds and test loops over a large repository in one 1M-token session.

  • Frontend and UI fixes

    Send screenshots of a rendered page along with the code and ask the model to fix layout or visual bugs.

  • Kernel and systems tuning

    Profile, rewrite and benchmark GPU kernels or other hot paths where correctness and speed both matter.

  • Research reports

    Feed in papers, filings or data and get a structured analysis back, with JSON schema output when the next step needs clean data.

Get started

Call Kimi K3 in three steps.

No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.

  1. 1

    Create a key

    Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.

  2. 2

    Point your SDK at zurelay

    Change the base URL. Everything else in your code stays the same.

    https://api.zurelay.com/v1
  3. 3

    Use kimi-k3

    Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.

    kimi-k3

FAQ

Kimi K3 API questions.

Anything else, write to sales@zurelay.com and an engineer will answer.

How much does the Kimi K3 API cost on zurelay?

Kimi K3 costs $1.27 per 1M input tokens and $6.37 per 1M output tokens on zurelay, with cached input at $0.127. Moonshot's list price is $3.00 input and $15.00 output, so you save 58%. One rate applies at every prompt length, up to the full 1M-token window.

Is there a cheaper Kimi API than Moonshot's own?

Yes. zurelay serves the same kimi-k3 model at 58% below Moonshot's list price. You top up prepaid credit and pay only for the tokens you use, with no subscription.

How do I call the Kimi K3 API with the OpenAI SDK?

Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "kimi-k3" to chat.completions.create. Leave temperature and top_p unset, since Moonshot fixes them for Kimi K3, and use reasoning_effort to pick low, high or max.

What is the Kimi K3 context window?

Kimi K3 holds 1,048,576 tokens of context. Moonshot's max_completion_tokens defaults to 131,072 and can go as high as 1,048,576. Reasoning and the answer share that completion budget, so leave headroom on hard prompts.

What is Kimi K3 best at?

Moonshot built Kimi K3 for long-horizon coding, knowledge work and reasoning. It sustains long engineering sessions, works across large repositories and uses screenshots to refine frontend and game code. In Moonshot's GPU kernel tests it performed competitively with Claude Fable 5.

Can Kimi K3 read images?

Yes. Kimi K3 has native vision and takes images alongside text in the same request. Send them as base64 data URLs inside the content array, because Moonshot's API does not accept public image URLs. The model returns text only.

How is Kimi K3 different from Kimi K2.6?

Kimi K3 is a new model built on Kimi Delta Attention, with a 1M-token context against 256K on Kimi K2.6 and K2.7 Code. It adds low, high and max reasoning effort and supports tool_choice required, which the K2 models do not. Moonshot lists it at a higher price than K2.6.

Does the Kimi K3 API support streaming and per-key limits?

Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.

Is Kimi K3 on zurelay the same model Moonshot serves?

Yes. Requests go to kimi-k3 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Moonshot's own API. zurelay is an independent service and is not affiliated with Moonshot AI.

More models on zurelay

All models

Stop paying list price.
Start saving today.

Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.