DeepSeek V4 Flash · V4 Pro · Vision · INR Billing · GST Invoice

DeepSeek API India
The V4 family, billed in rupees

DeepSeek’s current generation — V4 Flash, V4 Pro and the new vision model — with a one-million-token context window and OpenAI-compatible calls. Access them through Ogma on an Indian contract: INR invoicing, GST, purchase orders. No USD card, no forex spread you cannot see.

1M
Token context window across every V4 model
384K
Maximum output tokens in a single response
13B
Active parameters per token on V4 Flash, of 284B total
INR
Billed in rupees — GST invoice every month

DeepSeek API — Current Models

Three models are live. Share your monthly token volume and we’ll quote in INR within 2 business hours.

DeepSeek V4 models available through Ogma
Model ID Architecture Best For
deepseek-v4-flash 284B MoE — 13B activated The default. Coding, agents, extraction, summarisation and high-volume work. DeepSeek reports the 0731 release outperforming the V4-Pro preview on benchmarks.
deepseek-v4-pro 1.6T MoE — 49B activated The heavier model, at roughly 3× the per-token cost. Reach for it when you can measure Flash falling short, not by default.
deepseek-v4-flash-vision-exp Multimodal — experimental Document and screenshot understanding. Text quality on par with V4 Flash, with real gains on agent tasks that need to see. Images bill as input tokens by dimension.
Still calling V3 or R1? The legacy names deepseek-chat and deepseek-reasoner were retired on 24 July 2026 and no longer resolve. Migration is a model-id change, not a rewrite — the API shape is unchanged. We can move you across in an afternoon.
Cache-hit pricing
~3% of the miss rate
Repeated prefixes bill at a fraction of a fresh call
Off-peak discount
50% off list
Outside 01:00–04:00 and 06:00–10:00 UTC
INR billing + GST
Input tax credit
Recoverable where your business is GST-registered

Batch work that can run outside peak hours pays half. Combined with prefix caching on a stable system prompt, that is usually the largest single lever on an LLM bill — larger than switching model.

DeepSeek V4 — What Changed

A million tokens of context

Every V4 model takes a 1M-token context and can return up to 384K. A hybrid attention design combining compressed sparse and heavily compressed attention is what makes long context affordable rather than merely possible.

Mixture-of-experts, sharpened

V4 Flash activates 13B parameters of 284B; V4 Pro activates 49B of 1.6T. You pay for what is activated, not what is stored — which is why a model this capable does not price like one.

Thinking mode, built in

Reasoning is no longer a separate model. All three V4 models support thinking mode alongside JSON output and tool calls, so you switch behaviour with a parameter instead of routing to a different endpoint.

Vision, newly added

The experimental vision model matches V4 Flash on pure text while adding real gains on agent tasks that require reading a screen or a document. Treat the -exp suffix seriously: it is a preview, not a production commitment.

OpenAI-compatible API

The request and response shapes match the OpenAI SDK, so moving over is a base URL, a key and a model id. Existing Python, Node or LangChain code runs unchanged. An Anthropic-format endpoint is available too.

An Indian contract

Ogma Consulting is an Indian registered company. You get a rupee invoice with GST against a purchase order, on standard terms — the part that usually blocks an AI project in an Indian enterprise, rather than the model itself.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by DeepSeek’s own API infrastructure. We provide the Indian contracting, billing, spend controls and per-team usage attribution; we do not host DeepSeek model weights, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess DeepSeek’s terms against them before sending regulated data — and talk to us about which parts of the workload belong on a different model or a self-hosted one.

Pricing — in rupees

Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.

Pricing per million tokens in INR
Model Input / 1M Output / 1M Cached input / 1M
DeepSeek V4 Flash
deepseek/deepseek-v4-flash-0731
₹8.06₹18.13₹0.81
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
₹41.69₹83.37₹4.17

Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.

Frequently Asked Questions

Three: deepseek-v4-flash (the workhorse), deepseek-v4-pro (the larger model for the hardest tasks) and deepseek-v4-flash-vision-exp (multimodal, experimental). All three carry a 1M-token context window and 384K maximum output, and all three support thinking mode, JSON output and tool calls.

They are gone. The legacy API names deepseek-chat and deepseek-reasoner were retired on 24 July 2026 and are no longer reachable. Until that date they pointed at the non-thinking and thinking modes of deepseek-v4-flash respectively. Any integration still calling those names has already stopped working and needs to move to a V4 model id.

Start on Flash. It is a 284B-parameter mixture-of-experts model activating 13B per token, and DeepSeek states the 0731 release outperforms the V4-Pro preview on benchmarks despite the far smaller activation. Move a workload to Pro (1.6T total, 49B activated) only when you can measure Flash falling short — Pro costs roughly three times as much per token.

On DeepSeeks API infrastructure. Ogma provides the gateway, the Indian contract and the rupee billing — we do not host DeepSeek model weights ourselves, and we will not tell you otherwise. If your workload carries data-residency obligations under DPDPA or a sector regulator, assess that against DeepSeeks own terms before sending regulated data, and talk to us about which parts of the workload belong on a different model.

It is not cheaper per token — DeepSeeks list price is DeepSeeks list price. What changes is that you can actually pay for it: an INR invoice with GST, on a purchase order, from an Indian company. No international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is eligible.

Yes, and most teams should. Route high-volume, cost-sensitive work — classification, summarisation, extraction — to V4 Flash, and reserve a frontier model for the small share of traffic where quality is worth the premium. Ogmas AI Gateway gives you one key and one bill across providers, so the routing decision stays an engineering choice rather than a procurement exercise.

On V4 in an afternoon

DeepSeek V4 Flash, Pro and Vision through an Indian contract — INR billing, GST invoice, purchase orders. If you are still calling a retired V3 or R1 model name, that is a model-id change we can make with you today. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.