Llama 4 API India
The Llama 4 herd, billed in rupees
Meta's Llama 4 herd is natively multimodal and built on a mixture-of-experts design. Access it through Ogma on an Indian contract: INR invoicing, GST, purchase orders — no USD card, no forex spread you cannot see.
Llama 4 — Available Models
Share your monthly token volume and we’ll quote in INR within 2 business hours.
| Model | Shape | Best For |
|---|---|---|
| Llama 4 Scout | 17B active · 16 experts | Long-context work. A 10M-token window and small enough to serve from a single H100, which makes it the practical choice for document-heavy pipelines. |
| Llama 4 Maverick | 17B active · 128 experts | The stronger general model, with a 1M-token context. Meta positions it against GPT-4o and Gemini 2.0 Flash on general benchmarks. |
Meta Llama 4 — Why It Matters
Natively multimodal
Llama 4 was trained on text and vision together rather than having vision bolted on, so image understanding is part of the base model rather than a separate endpoint.
Mixture-of-experts
Both models activate 17B parameters per token. Scout draws from 16 experts, Maverick from 128 — the capacity differs, the per-token compute does not.
A 10M-token window
Scout's context is long enough to hold an entire codebase or a year of contracts in one call, which changes what a retrieval pipeline has to do.
OpenAI-compatible access
Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.
An Indian contract
Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.
Spend controls that hold
Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.
Where your data goes — plainly
Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Llama 4 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.
Frequently Asked Questions
Try Llama 4 on an Indian invoice
One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.