Kimi K3 · Open Weights · INR Billing · GST Invoice

Kimi K3 API India
Frontier open weights, billed in rupees

Moonshot AI's Kimi K3 arrived in July 2026 as one of the strongest open-weight models available, with a trillion-scale mixture-of-experts design and a million-token context. Access it through Ogma on an Indian contract: INR invoicing, GST, purchase orders.

2.8T
Total parameters — mixture-of-experts
1M
Token context window
Jul 2026
Weights published on Hugging Face
INR
Billed in rupees — GST invoice every month

Kimi K3 — Available Models

Share your monthly token volume and we’ll quote in INR within 2 business hours.

Kimi K3 models available through Ogma
Model Shape Best For
Kimi K32.8T total · MoEThe flagship. Strong on agentic and terminal-style benchmarks, which is where long-horizon coding work actually lives. Weights are public.
Kimi K2.7 CodeCoding specialistThe preceding code-focused release, still available. Worth benchmarking against K3 on your own repository before assuming the newer model wins on your workload.
Licence Weights are published openly on Hugging Face. Confirm the exact licence terms on the model card before commercial deployment — Moonshot has used modified permissive licences across the K2 line.

Moonshot AI Kimi K3 — Why It Matters

Built for agents

K3's gains over the K2 line show up most on multi-step agentic and terminal benchmarks — tasks that run for many turns rather than answering in one.

Trillion-scale MoE

2.8T total parameters with a small fraction activated per token. Capacity where it helps, compute cost where it does not.

Open weights

Published on Hugging Face, so self-hosting stays an option if your data cannot leave your own infrastructure.

OpenAI-compatible access

Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.

An Indian contract

Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.

Spend controls that hold

Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Kimi K3 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.

Pricing — in rupees

Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.

Pricing per million tokens in INR
Model Input / 1M Output / 1M Cached input / 1M
Kimi K3
moonshotai/kimi-k3
₹302₹1,511₹30.22

Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.

Frequently Asked Questions

In principle yes — the weights are public. In practice a 2.8T-parameter model is a serious infrastructure commitment, and most teams should validate the workload against the hosted API first. If self-hosting turns out to be right, Ogma can scope the GPU footprint and deploy it; that is a different engagement from API access, and we will tell you which one you actually need.

In rupees, on a GST invoice, from an Indian registered company. You can run it on a purchase order with standard payment terms — no international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is GST-registered.

No. Access is through Ogma's AI Gateway, which speaks the OpenAI API. If your code already uses the OpenAI SDK, you change the base URL, the key and the model id. Streaming, tool calls and JSON mode behave as you expect.

On the model provider's infrastructure, reached through Ogma's gateway. We provide the Indian contracting, the billing, the spend controls and per-team usage attribution — we do not host these model weights ourselves, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess the provider's terms against them before sending regulated data, and talk to us about self-hosting the parts that need it.

Yes, and most teams should. One key and one bill across providers means routing high-volume work to a cheaper model and reserving a frontier model for the small share of traffic that needs it stays an engineering decision, not a procurement exercise.

Try Kimi K3 on an Indian invoice

One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.