Gemma 4 · Open Weights · INR Billing · GST Invoice

Gemma 4 API India
Small models that fit, billed in rupees

Google DeepMind's Gemma 4 is the family you deploy rather than merely call — Apache 2.0 weights from 2B up to 31B, sized so a capable model fits on hardware you already own. Access it through Ogma on an Indian contract, or let us deploy it inside your network.

Apache 2.0
A genuine open-source licence, not a custom one
256K
Token context on the 12B, 26B and 31B sizes
2B–31B
Sizes, so you can match model to hardware
INR
Billed in rupees — GST invoice every month

Gemma 4 — Available Models

Share your monthly token volume and we’ll quote in INR within 2 business hours.

Gemma 4 models available through Ogma
Model Shape Best For
Gemma 4 E2B / E4BCompact · 128K contextThe small end. Built to run on modest hardware — edge deployments, on-device work, and high-volume classification where latency and unit cost dominate.
Gemma 4 12B / 26B-A4B / 31B256K contextThe capable end, still small enough to self-host on a single machine in most cases. The 26B is a mixture-of-experts variant activating roughly 4B per token.
Licence Apache 2.0 — a standard OSI-approved open-source licence with no additional conditions on scale, naming or field of use.

Google DeepMind Gemma 4 — Why It Matters

Apache 2.0, genuinely

Unlike most model licences worth reading carefully, Apache 2.0 needs no interpretation. No user thresholds, no naming conditions, no review before you ship.

Small enough to own

The point of Gemma is that you can run it yourself. For an Indian enterprise with data-residency obligations, a model that fits on your own GPU is a different compliance conversation entirely.

Cost at volume

For classification, extraction and routing at scale, a well-fitted small model beats a frontier model on cost by orders of magnitude, and often matches it on the task.

OpenAI-compatible access

Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.

An Indian contract

Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.

Spend controls that hold

Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Gemma 4 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.

Pricing — in rupees

Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.

Pricing per million tokens in INR
Model Input / 1M Output / 1M Cached input / 1M
Gemma 4 26B-A4B
google/gemma-4-26b-a4b-it
₹7.05₹34.25₹0.71
Gemma 4 31B
google/gemma-4-31b-it
₹10.07₹34.25₹1.01

Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.

Frequently Asked Questions

Often, yes — and this is the one family where we will say so plainly. Gemma exists to be run on your own hardware. If you have data-residency obligations under the DPDP Act or a sector regulator, a 12B Gemma on a GPU in your own rack answers the residency question completely, in a way that no hosted API can. Ogma can size the hardware, deploy it and support it. Use the API to prototype; move to your own infrastructure when the workload justifies it.

In rupees, on a GST invoice, from an Indian registered company. You can run it on a purchase order with standard payment terms — no international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is GST-registered.

No. Access is through Ogma's AI Gateway, which speaks the OpenAI API. If your code already uses the OpenAI SDK, you change the base URL, the key and the model id. Streaming, tool calls and JSON mode behave as you expect.

On the model provider's infrastructure, reached through Ogma's gateway. We provide the Indian contracting, the billing, the spend controls and per-team usage attribution — we do not host these model weights ourselves, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess the provider's terms against them before sending regulated data, and talk to us about self-hosting the parts that need it.

Yes, and most teams should. One key and one bill across providers means routing high-volume work to a cheaper model and reserving a frontier model for the small share of traffic that needs it stays an engineering decision, not a procurement exercise.

Try Gemma 4 on an Indian invoice

One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.