GLM-5.2 · Open Weights · INR Billing · GST Invoice

GLM-5.2 API India
Long-horizon agents, billed in rupees

Z.ai's GLM-5.2 is built for long-horizon work — the multi-step coding and agent tasks that run for hours rather than seconds. Access it through Ogma on an Indian contract: INR invoicing, GST, purchase orders.

753B
Parameters — mixture-of-experts
Jul 2026
Latest weights published
SWE-bench
Among the strongest open models on software engineering tasks
INR
Billed in rupees — GST invoice every month

GLM-5.2 — Available Models

Share your monthly token volume and we’ll quote in INR within 2 business hours.

GLM-5.2 models available through Ogma
Model Shape Best For
GLM-5.2753B · MoEThe current flagship, positioned by Z.ai for long-horizon tasks. Consistently near the top of open-weight coding and agent leaderboards.
GLM-5.1754B · MoEThe preceding release, still published. Useful as a stability baseline if you are already in production on it.
Licence GLM weights are published on Hugging Face under a permissive licence. Confirm the exact terms on the model card for your release before commercial deployment.

Z.ai (Zhipu AI) GLM-5.2 — Why It Matters

Long-horizon by design

Z.ai positions GLM-5.2 specifically for tasks that span many steps — refactors, migrations, multi-file changes — rather than single-turn answers.

Strong on real software work

The model line scores well on SWE-bench-style benchmarks, which test changes against real repositories rather than isolated puzzles.

Open weights

Published on Hugging Face under a permissive licence, so self-hosting remains available if residency demands it.

OpenAI-compatible access

Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.

An Indian contract

Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.

Spend controls that hold

Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because GLM-5.2 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.

Pricing — in rupees

Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.

Pricing per million tokens in INR
Model Input / 1M Output / 1M Cached input / 1M
GLM-5.2
z-ai/glm-5.2
₹97.32₹306₹9.73

Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.

Frequently Asked Questions

They are close, and the honest answer is that it depends on your repository. GLM-5.2 is built around long-horizon tasks and tends to hold up over multi-step changes; DeepSeek V4 Flash is the cheaper per token and very strong on shorter tasks. Run both against a sample of your own tickets before committing — Ogma's gateway gives you one key across both, so that comparison costs you an afternoon rather than two procurement cycles.

In rupees, on a GST invoice, from an Indian registered company. You can run it on a purchase order with standard payment terms — no international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is GST-registered.

No. Access is through Ogma's AI Gateway, which speaks the OpenAI API. If your code already uses the OpenAI SDK, you change the base URL, the key and the model id. Streaming, tool calls and JSON mode behave as you expect.

On the model provider's infrastructure, reached through Ogma's gateway. We provide the Indian contracting, the billing, the spend controls and per-team usage attribution — we do not host these model weights ourselves, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess the provider's terms against them before sending regulated data, and talk to us about self-hosting the parts that need it.

Yes, and most teams should. One key and one bill across providers means routing high-volume work to a cheaper model and reserving a frontier model for the small share of traffic that needs it stays an engineering decision, not a procurement exercise.

Try GLM-5.2 on an Indian invoice

One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.