Qwen 3.7 & 3.8 · Open Weights · INR Billing · GST Invoice

Qwen 3.7 & 3.8 API India
Both Qwen lines, billed in rupees

Alibaba ships Qwen as two lines: the hosted 3.7 series, and the 3.8 series published as open weights. Both are available through Ogma on an Indian contract — INR invoicing, GST, purchase orders. Every model in the range carries a one-million-token context.

1M
Token context window across all six models
₹3
Per million input tokens on Qwen3.7-Flash, in rupees
2.4T
Parameters on the largest open-weight model
INR
Billed in rupees — GST invoice every month

Qwen 3.7 & 3.8 — Available Models

Share your monthly token volume and we’ll quote in INR within 2 business hours.

Qwen 3.7 & 3.8 models available through Ogma
Model Shape Best For
Qwen3.7-FlashHosted · 1M contextThe cheapest capable model we can offer. At roughly a fifth of DeepSeek V4 Flash's input price, it changes what is economic at volume — classification, routing, extraction.
Qwen3.7-PlusHosted · 1M contextThe general-purpose middle of the hosted line. Text, image and video input.
Qwen3.7-MaxHosted · 1M contextAlibaba's flagship hosted model, for work where capability outweighs cost.
Qwen3.8-27BOpen weights · 1M contextA dense vision-language model small enough to self-host on a single professional GPU. Strong on coding and long-horizon agent tasks.
Qwen3.8-2.4T-A95BOpen weights · 1M contextSparse mixture-of-experts, 95B active of 2.4T total — the open-weight counterpart to Qwen3.8-Max.
Qwen3.8-MaxHosted · 1M contextThe hosted flagship of the 3.8 generation.
Licence The Qwen 3.8 open-weight models are published on Hugging Face; confirm the exact terms on the model card for your release. The 3.7 line is hosted only — weights are not published, and community re-uploads claiming otherwise are unofficial and should not be used commercially.

Alibaba Cloud Qwen 3.7 & 3.8 — Why It Matters

Two lines, one decision

The 3.7 series is hosted only; the 3.8 series publishes weights. Start on either through one key, and move to your own hardware later only if the workload justifies it.

Genuinely cheap at the bottom

Qwen3.7-Flash prices input at a small fraction of the models it competes with, which makes high-volume classification and extraction economic rather than merely possible.

Self-hosting stays open

Qwen3.8-27B fits on a single professional GPU. For a workload under data-residency obligations, that is a different compliance conversation from any hosted API.

OpenAI-compatible access

Reached through Ogma’s AI Gateway, which speaks the OpenAI API. Change a base URL, a key and a model id — existing Python, Node or LangChain code runs unchanged.

An Indian contract

Ogma Consulting is an Indian registered company. Rupee invoice with GST against a purchase order, on standard terms — usually the part that blocks an AI project here, rather than the model.

Spend controls that hold

Per-team monthly caps, per-minute rate limits and per-user cost attribution. A runaway loop hits a ceiling in seconds rather than appearing on an invoice a month later.

Where your data goes — plainly

Calls you make through Ogma are proxied by our gateway and served by the model provider’s infrastructure. We provide the Indian contracting, billing, spend controls and usage attribution; we do not host these model weights, and we will not claim to. Because Qwen 3.7 & 3.8 is published as open weights, self-hosting inside your own network is a genuine option where residency demands it — ask us to scope it rather than assuming a hosted API is your only route.

Pricing — in rupees

Per million tokens, inclusive of our margin. Billed in INR on a GST invoice — never in dollars.

Pricing per million tokens in INR
Model Input / 1M Output / 1M Cached input / 1M
Qwen3.7-Flash
qwen/qwen3.7-flash
₹3.02₹13.1₹0.3
Qwen3.7-Plus
qwen/qwen3.7-plus
₹32.24₹129₹3.22
Qwen3.7-Max
qwen/qwen3.7-max
₹149₹446₹14.86
Qwen3.8-27B
qwen/qwen3.8-27b
₹45.33₹322₹4.53

Rates convert at the mid-market USD/INR rate refreshed every six hours, and the rate applied is recorded against every call. Cached input applies automatically to repeated prompt prefixes. GST is charged on top and is recoverable as input tax credit where your business is registered.

Frequently Asked Questions

They are two parallel lines, not successive versions. The 3.7 series (Flash, Plus, Max) is hosted by Alibaba with closed weights. The 3.8 series publishes open weights for the 27B and 2.4T-A95B models, alongside a hosted 3.8-Max. If you may want to self-host later, start on 3.8; if you want the cheapest capable model today, start on 3.7-Flash.

Qwen3.7-Flash for high-volume, cost-sensitive work — its input pricing is a fraction of comparable models. Qwen3.8-27B if self-hosting is on your roadmap, because it is the one that fits on hardware you can actually buy. Both are reachable on the same key, so testing both costs you an afternoon.

No. Alibaba has not published weights for the 3.7 line. Files circulating on model hubs that claim to be Qwen3.7 are unofficial third-party conversions, often with safety behaviour modified, and carry no licence you could show an auditor. For self-hosting, use Qwen3.8-27B.

In rupees, on a GST invoice, from an Indian registered company. You can run it on a purchase order with standard payment terms — no international card, no foreign-entity receipt your finance team cannot book, and input tax credit where your business is GST-registered.

No. Access is through Ogma's AI Gateway, which speaks the OpenAI API. If your code already uses the OpenAI SDK, you change the base URL, the key and the model id. Streaming, tool calls and JSON mode behave as you expect.

On the model provider's infrastructure, reached through Ogma's gateway. We provide the Indian contracting, the billing, the spend controls and per-team usage attribution — we do not host these model weights ourselves, and we will not claim to. If your workload carries data-residency obligations under the DPDP Act or a sector regulator, assess the provider's terms against them before sending regulated data, and talk to us about self-hosting the parts that need it.

Yes, and most teams should. One key and one bill across providers means routing high-volume work to a cheaper model and reserving a frontier model for the small share of traffic that needs it stays an engineering decision, not a procurement exercise.

Try Qwen 3.7 & 3.8 on an Indian invoice

One key, INR billing, GST invoice, purchase orders. Tell us your token volume and we’ll come back with live rupee pricing within 2 business hours.