Long-horizon coding
Moonshot built K3 to run long engineering sessions with little supervision, navigate large repositories and orchestrate terminal tools.
The Kimi K3 API for long-horizon coding and agents, 58% below Moonshot's list price
kimi-k3from openai import OpenAIclient = OpenAI( base_url="https://api.zurelay.com/v1", api_key="YOUR_ZURELAY_KEY",)stream = client.chat.completions.create( model="kimi-k3", messages=[{"role": "user", "content": "Explain quicksort in two sentences."}], stream=True,)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="")Base URL https://api.zurelay.com/v1. Works with any OpenAI SDK.
Pricing
Pay as you go from prepaid credit, with no subscription and no minimum. Requests that fail are never billed.
| Rate | zurelay | Moonshot AI list | You save |
|---|---|---|---|
Input per 1M tokens | $1.27 | 58%off | |
Output per 1M tokens | $6.37 | 58%off | |
Cached input per 1M tokens | $0.127 | — | — |
One rate at every prompt length, up to the full 1M context window. Streaming, tool calls and structured outputs cost nothing extra.
Move the sliders to your monthly usage.
$2,074 a year
Overview
Kimi K3 is Moonshot AI's flagship model: 2.8 trillion parameters, native vision and a 1M-token context window, built for long-horizon coding and knowledge work. The zurelay Kimi K3 API serves the same model through one OpenAI-compatible endpoint at 58% below Moonshot's list price.
Kimi K3 is Moonshot AI's most capable model, released on July 16, 2026, with the full weights following on Hugging Face on July 27. The Kimi K3 API reads text and images, returns text and holds 1,048,576 tokens of context. Moonshot's max_completion_tokens defaults to 131,072 and can be set as high as 1,048,576.
It is a Mixture-of-Experts model with 2.8T total parameters and 104B active per token, routing each token to 16 of 896 experts. Moonshot built it on Kimi Delta Attention and Attention Residuals and reports about 2.5 times better scaling efficiency than Kimi K2. The weights use MXFP4 with MXFP8 activations and ship under the Kimi K3 License, which allows commercial use; very large Model-as-a-Service businesses need a separate agreement with Moonshot.
Moonshot built Kimi K3 for long-horizon coding, knowledge work and reasoning. It can run long engineering sessions with little supervision, move through large repositories and drive terminal tools, and it uses screenshots to refine frontend, game and CAD work. Moonshot says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, and reports it ahead of the other models it tested.
Thinking is always on. Moonshot's API takes a reasoning_effort of low, high or max, with max as the default, and fixes temperature at 1.0 and top_p at 0.95. The model supports tool calling, including tool_choice required, plus JSON schema output and automatic context caching. On zurelay you call it at https://api.zurelay.com/v1 with the model ID kimi-k3, for $1.27 per 1M input tokens and $6.37 per 1M output tokens against Moonshot's list price of $3.00 and $15.00.
Strengths
Moonshot built K3 to run long engineering sessions with little supervision, navigate large repositories and orchestrate terminal tools.
K3 reads screenshots natively, so it can check rendered output and refine frontend, game or CAD code in the same session.
In Moonshot's GPU kernel optimization tests, K3 performed competitively with Claude Fable 5 and well ahead of GPT-5.6 Sol.
Moonshot reports consistent gains on its internal knowledge-work evaluations, from multi-source research to reports with interactive charts.
Use cases
Run multi-hour refactors, feature builds and test loops over a large repository in one 1M-token session.
Send screenshots of a rendered page along with the code and ask the model to fix layout or visual bugs.
Profile, rewrite and benchmark GPU kernels or other hot paths where correctness and speed both matter.
Feed in papers, filings or data and get a structured analysis back, with JSON schema output when the next step needs clean data.
Get started
No waitlist and no new SDK. If your code already talks to OpenAI, it already talks to zurelay.
Sign up, add credit and create an API key. Set a monthly budget or a rate limit per key if you like.
Change the base URL. Everything else in your code stays the same.
https://api.zurelay.com/v1Send chat completions as usual. Streaming, tool calls and usage reporting work as you expect.
kimi-k3Works with the tools you already use
Compare
$0.38 / $1.21 per 1M tokens
Pick GLM-5.3 for text-only coding and security work at a lower lab list price.
GLM-5.3 API$3.50 / $17.50 per 1M tokens
Pick Claude Fable 5 when you need the strongest overall results; Moonshot itself says K3 still trails it.
Claude Fable 5 API$0.18 / $0.36 per 1M tokens
Pick MiMo V2.6 Pro when you need audio or video input from an open-weight model with a 1M-token context.
MiMo V2.6 Pro APIFAQ
Anything else, write to sales@zurelay.com and an engineer will answer.
Kimi K3 costs $1.27 per 1M input tokens and $6.37 per 1M output tokens on zurelay, with cached input at $0.127. Moonshot's list price is $3.00 input and $15.00 output, so you save 58%. One rate applies at every prompt length, up to the full 1M-token window.
Yes. zurelay serves the same kimi-k3 model at 58% below Moonshot's list price. You top up prepaid credit and pay only for the tokens you use, with no subscription.
Install the official OpenAI SDK, set the base URL to https://api.zurelay.com/v1 and use your zurelay API key. Then pass model "kimi-k3" to chat.completions.create. Leave temperature and top_p unset, since Moonshot fixes them for Kimi K3, and use reasoning_effort to pick low, high or max.
Kimi K3 holds 1,048,576 tokens of context. Moonshot's max_completion_tokens defaults to 131,072 and can go as high as 1,048,576. Reasoning and the answer share that completion budget, so leave headroom on hard prompts.
Moonshot built Kimi K3 for long-horizon coding, knowledge work and reasoning. It sustains long engineering sessions, works across large repositories and uses screenshots to refine frontend and game code. In Moonshot's GPU kernel tests it performed competitively with Claude Fable 5.
Yes. Kimi K3 has native vision and takes images alongside text in the same request. Send them as base64 data URLs inside the content array, because Moonshot's API does not accept public image URLs. The model returns text only.
Kimi K3 is a new model built on Kimi Delta Attention, with a 1M-token context against 256K on Kimi K2.6 and K2.7 Code. It adds low, high and max reasoning effort and supports tool_choice required, which the K2 models do not. Moonshot lists it at a higher price than K2.6.
Yes. Set stream: true to receive tokens as they are generated, and add stream_options.include_usage to get token counts in the final chunk. You can give each zurelay API key its own monthly budget and requests-per-minute cap. Failed requests are retried on another route and never billed.
Yes. Requests go to kimi-k3 itself, never a smaller or substitute model. Output is sampled, so wording varies from run to run, just as it does on Moonshot's own API. zurelay is an independent service and is not affiliated with Moonshot AI.
Create a key in seconds, point the SDK you already use at zurelay, and every request costs up to 85% less from the first token.