Llama 4 API India
The Llama 4 herd, billed in rupees
Meta's Llama 4 herd is natively multimodal and built on a mixture-of-experts design. Access it through Ogma on an Indian contract: INR invoicing, GST, purchase orders — no USD card, no forex spread you cannot see.
Llama 4 — Available Models
Share your monthly token volume and we’ll quote in INR within 2 business hours.
| Model | Shape | Best For |
|---|---|---|
| Llama 4 Scout | 17B active · 16 experts | Long-context work. A 10M-token window and small enough to serve from a single H100, which makes it the practical choice for document-heavy pipelines. |
| Llama 4 Maverick | 17B active · 128 experts | The stronger general model, with a 1M-token context. Meta positions it against GPT-4o and Gemini 2.0 Flash on general benchmarks. |
Meta Llama 4 — Why It Matters
Natively multimodal
Llama 4 was trained on text and vision together rather than having vision bolted on, so image understanding is part of the base model rather than a separate endpoint.
Mixture-of-experts
Both models activate 17B parameters per token. Scout draws from 16 experts, Maverick from 128 — the capacity differs, the per-token compute does not.
A 10M-token window
Scout's context is long enough to hold an entire codebase or a year of contracts in one call, which changes what a retrieval pipeline has to do.
OpenAI-compatible access
Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.
An Indian contract
Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.
Spend controls that hold
Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.
Where your data goes — plainly
Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Llama 4 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.
Pricing — in rupees
Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.
| Model | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
Llama 4 Scoutmeta-llama/llama-4-scout | ₹10.07 | ₹30.22 | ₹1.01 |
Llama 4 Maverickmeta-llama/llama-4-maverick | ₹20.15 | ₹80.59 | ₹2.01 |
Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.
Frequently Asked Questions
Try Llama 4 on an Indian invoice
One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.