Self-hosted Kimi K3 · Founding reservation

The cheapest Kimi K3 inference in the world.

Kimi K3 at half the public API price—served directly from GPU infrastructure we operate, without a router or third-party inference API.

Input$1.50/ 1M tokens
Output$7/ 1M tokens
No prompt logsPrompts and completions are not retained.
No hidden routingChoose Kimi K3 and your request goes to Kimi K3.
No extra down-quantizationNative MXFP4 weights and MXFP8 activations, as released.

$10,000 activates the inference cloud.

$850 reserved$10,000 target

Launches immediately at $10,000. Full refunds if not funded by August 30.

I was fed up with expensive Kimi K3 inference.

I’m Dhrumil—one heavy API user, not a funded startup. Kimi K3 is open-weight, and its license permits people to run and deploy the model. That means I can operate the weights on rented GPUs instead of buying Moonshot’s hosted API, so their API price does not have to become your price.

Why reserve instead of sign up?

Because a signup does not reserve a GPU.

Free signups are easy to collect, but they do not tell me whether enough people will actually use the service. Paid reservations measure real demand. If we reach $10,000, I will use the committed funds to reserve the inference capacity needed to launch.

What happens to my money?

It stays locked until we launch or refund everyone.

Your payment remains tied to your Kimi K3 reservation. It cannot be spent on existing Vynaris services. If the target fills, it converts 1:1 into account credit when inference launches. If we do not proceed, every reservation is refunded to its original payment method.

Frontier-scale inference without the frontier-scale bill.

Every price is per one million tokens. No blended router price and no separate platform markup.

ProviderInputOutput
VynarisIntroductory rate$1.50$7.00
Kimi APIPublic list price$3.00$15.00
Save50%on input53%on output
Quality promiseNative precision—not a cheaper quant hidden behind the same name.

Launch traffic will use Kimi K3’s native MXFP4 weights and MXFP8 activations, with no additional down-quantization. If people ask for a lower-precision version at a lower price, we will offer it later as a separate and clearly labelled tier.

Public-price comparison checked July 31, 2026 against the Kimi API list price. “Cheapest” will be reverified before reservations open. Cached input, taxes, and future promotional prices are excluded.

Private by architecture

Your prompts are not the product.

Specific-model requests bypass Vynaris Auto. They go straight to the Kimi K3 deployment running on GPU capacity we operate.

01
No third-party inference API

Your request is not forwarded to an external model provider.

02
No prompt or completion retention

Request content is processed for inference, then discarded.

03
Only necessary billing metadata

Token counts, costs, and operational records—not your content.

Your money stays yours—either as credit or as a refund.

01

Choose $50 or more

Make one reservation or add more later. There is no per-account limit.

02

We hold it for Kimi K3

You receive a payment receipt, but the amount is not spendable before launch.

03

Credit—or a full refund

If the target fills, we launch immediately and it becomes Vynaris credit. If not funded by August 30, it returns to your original payment method.

Direct Kimi K3, through the SDK you already use.

Keep your OpenAI-compatible client. Set the model explicitly and Vynaris skips automatic routing.

from openai import OpenAI

client = OpenAI(
    api_key="vyn_...",
    base_url="https://api.vynaris.com/v1",
)

response = client.chat.completions.create(
    model="vynaris/kimi-k3",
    messages=[{"role": "user", "content": "..."}],
)

Clear terms, before any payment.

What am I buying today?

A paid launch reservation. Your money is held as locked Kimi K3 credit, not added to your spendable Vynaris balance yet.

Why take reservations instead of free signups?

I am an individual, not a funded startup. A paid reservation shows real demand and gives me the money needed to reserve dedicated GPU capacity without raising outside funding.

What happens if the $10,000 target is not reached?

If the target is not reached by August 30, 2026, your full reservation is refunded to the original payment method. If it fills earlier, we begin the launch immediately.

Can I reserve more than once?

Yes. There is no per-account reservation limit. Every successful reservation is added to your total locked credit.

Can I spend the credit on other models?

After Kimi K3 launches, the reservation becomes normal Vynaris credit. You can use it on Kimi K3 or any other available model.

Is $1.50 input and $7 output a lifetime price?

It is the introductory launch rate, not a lifetime guarantee. Any future pricing change will be communicated before it takes effect.

Will you store my prompts or responses?

No. Kimi K3 API prompts and completions will not be retained. We keep only the account, token, cost, and operational records needed to bill and run the service.

Will this be a lower-quality quantized version?

No. The launch service will run Kimi K3 at its native released precision—MXFP4 weights with MXFP8 activations—with no additional down-quantization. If users later want a cheaper lower-precision option, it will be offered as a separate, clearly labelled tier.

Reserve Kimi K3 at the introductory rate.

$50 minimum. Immediate launch at $10,000—or a full refund after August 30.

Choose your amount →
$1.50 / $7per 1M tokens
Reserve from $50