Llama 4 · Open Weights · INR Billing · GST Invoice

Llama 4 API India
The Llama 4 herd, billed in rupees

Meta's Llama 4 herd is natively multimodal and built on a mixture-of-experts design. Access it through Ogma on an Indian contract: INR invoicing, GST, purchase orders — no USD card, no forex spread you cannot see.

10M
Token context on Llama 4 Scout — the longest in the herd
17B
Active parameters per token on both Scout and Maverick
128
Experts in Maverick; Scout uses 16
INR
Billed in rupees — GST invoice every month

Llama 4 — Available Models

Share your monthly token volume and we’ll quote in INR within 2 business hours.

Llama 4 models available through Ogma
Model Shape Best For
Llama 4 Scout17B active · 16 expertsLong-context work. A 10M-token window and small enough to serve from a single H100, which makes it the practical choice for document-heavy pipelines.
Llama 4 Maverick17B active · 128 expertsThe stronger general model, with a 1M-token context. Meta positions it against GPT-4o and Gemini 2.0 Flash on general benchmarks.
Licence Llama 4 Community License — a custom Meta licence, not an OSI-approved open-source one. Read the terms before building a product on it; there are conditions on scale and on naming.

Meta Llama 4 — Why It Matters

Natively multimodal

Llama 4 was trained on text and vision together rather than having vision bolted on, so image understanding is part of the base model rather than a separate endpoint.

Mixture-of-experts

Both models activate 17B parameters per token. Scout draws from 16 experts, Maverick from 128 — the capacity differs, the per-token compute does not.

A 10M-token window

Scout's context is long enough to hold an entire codebase or a year of contracts in one call, which changes what a retrieval pipeline has to do.

OpenAI-compatible access

Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.

An Indian contract

Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.

Spend controls that hold

Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Llama 4 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.

Frequently Asked Questions

Open weights, not open source. The Llama 4 Community License is a custom Meta licence with conditions attached — including restrictions tied to monthly active users and requirements about how you name derivative models. It is permissive enough for most commercial use, but it is not Apache or MIT, and your legal team should read it rather than assume.

In rupees, on a GST invoice, from an Indian registered company. You can run it on a purchase order with standard payment terms — no international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is GST-registered.

No. Access is through Ogma's AI Gateway, which speaks the OpenAI API. If your code already uses the OpenAI SDK, you change the base URL, the key and the model id. Streaming, tool calls and JSON mode behave as you expect.

On the model provider's infrastructure, reached through Ogma's gateway. We provide the Indian contracting, the billing, the spend controls and per-team usage attribution — we do not host these model weights ourselves, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess the provider's terms against them before sending regulated data, and talk to us about self-hosting the parts that need it.

Yes, and most teams should. One key and one bill across providers means routing high-volume work to a cheaper model and reserving a frontier model for the small share of traffic that needs it stays an engineering decision, not a procurement exercise.

Try Llama 4 on an Indian invoice

One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.