Qwen 3.7 & 3.8 API India
Both Qwen lines, billed in rupees
Alibaba ships Qwen as two lines: the hosted 3.7 series, and the 3.8 series published as open weights. Both are available through Ogma on an Indian contract — INR invoicing, GST, purchase orders. Every model in the range carries a one-million-token context.
Qwen 3.7 & 3.8 — Available Models
Share your monthly token volume and we’ll quote in INR within 2 business hours.
| Model | Shape | Best For |
|---|---|---|
| Qwen3.7-Flash | Hosted · 1M context | The cheapest capable model we can offer. At roughly a fifth of DeepSeek V4 Flash's input price, it changes what is economic at volume — classification, routing, extraction. |
| Qwen3.7-Plus | Hosted · 1M context | The general-purpose middle of the hosted line. Text, image and video input. |
| Qwen3.7-Max | Hosted · 1M context | Alibaba's flagship hosted model, for work where capability outweighs cost. |
| Qwen3.8-27B | Open weights · 1M context | A dense vision-language model small enough to self-host on a single professional GPU. Strong on coding and long-horizon agent tasks. |
| Qwen3.8-2.4T-A95B | Open weights · 1M context | Sparse mixture-of-experts, 95B active of 2.4T total — the open-weight counterpart to Qwen3.8-Max. |
| Qwen3.8-Max | Hosted · 1M context | The hosted flagship of the 3.8 generation. |
Alibaba Cloud Qwen 3.7 & 3.8 — Why It Matters
Two lines, one decision
The 3.7 series is hosted only; the 3.8 series publishes weights. Start on either through one key, and move to your own hardware later only if the workload justifies it.
Genuinely cheap at the bottom
Qwen3.7-Flash prices input at a small fraction of the models it competes with, which makes high-volume classification and extraction economic rather than merely possible.
Self-hosting stays open
Qwen3.8-27B fits on a single professional GPU. For a workload under data-residency obligations, that is a different compliance conversation from any hosted API.
OpenAI-compatible access
Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.
An Indian contract
Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.
Spend controls that hold
Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.
Where your data goes — plainly
Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Qwen 3.7 & 3.8 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.
Pricing — in rupees
Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.
| Model | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
Qwen3.7-Flashqwen/qwen3.7-flash | ₹3.02 | ₹13.1 | ₹0.3 |
Qwen3.7-Plusqwen/qwen3.7-plus | ₹32.24 | ₹129 | ₹3.22 |
Qwen3.7-Maxqwen/qwen3.7-max | ₹149 | ₹446 | ₹14.86 |
Qwen3.8-27Bqwen/qwen3.8-27b | ₹45.33 | ₹322 | ₹4.53 |
Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.
Frequently Asked Questions
Try Qwen 3.7 & 3.8 on an Indian invoice
One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.