AI Cost Intelligence, Verified Daily.
Skip to content
Model releases5 min readUpdated Aug 18, 2026

Grok 4.6 Pricing Analysis | What xAI Charges and When to Pay

xAI has added Grok 4.6 to its API catalog at $2.00 per million input tokens and $6.00 per million output tokens. Here is how its unit economics stack up against cheaper high-speed alternatives.

Grok 4.6 Pricing cost chart in the AIOPLY house style
Grok 4.6 Pricing cost chart in the AIOPLY house style

Key takeaways

  • Grok 4.6 carries an API list price of $2.00 per 1M input tokens and $6.00 per 1M output tokens, recorded on 2026-08-12.
  • Prompt caching reduces Grok 4.6 input costs by 75 percent down to $0.50 per 1M tokens.
  • The model offers a 500k token context window and a maximum output capacity of 128k tokens.
  • At $6.00 per 1M output tokens, Grok 4.6 sits 60 percent above the $3.75 median output rate across 35 tracked current models.
  • Flash competitors like Qwen3.7 Flash offer input rates of $0.03 per 1M tokens and output rates of $0.13 per 1M tokens.

What does Grok 4.6 cost on the xAI API?

xAI introduced Grok 4.6 to its production API on 2026-08-12 at a rate of $2.00 per 1M input tokens and $6.00 per 1M output tokens. These rates put the model squarely in xAI flagship pricing tier, offering an expansive 500k token context window alongside a maximum output limit of 128k tokens.

Analyzing grok 4.6 pricing requires looking closely at how request architectures interact with xAI token billing rules. Standard prompt inputs are priced at $2.00 per million tokens, but xAI supports prompt caching, which drops the input token price down to $0.50 per 1M tokens. This represents a 75 percent reduction below standard input rates, allowing developers to optimize prompt engineering strategies for recurring background context.

The output rate of $6.00 per 1M tokens remains static regardless of generation volume or context reuse. With a 128k maximum output capacity, a single fully extended request generating maximum tokens will cost $0.768 in output fees alone. Teams planning long text generation tasks must budget carefully for these output costs, as output tokens account for the largest share of API expenditure in heavy synthesis pipelines.

How does Grok 4.6 compare to standard xAI pricing and market medians?

Grok 4.6 matches the exact rate card of its direct predecessor, Grok 4.5, which also costs $2.00 per 1M input tokens and $6.00 per 1M output tokens. However, its $6.00 output charge sits significantly higher than the median output rate of $3.75 per 1M tokens across 35 tracked current models in our dataset.

When examining xai api pricing over time, the decision to keep rates identical between Grok 4.5 and Grok 4.6 indicates a stabilization in xAI flagship unit costs. The baseline cost structure has not increased despite the integration of enhanced capabilities recorded in our pricing database, such as long context handling.

Compared to the broader market, Grok 4.6 sits above the median baseline. Across 35 tracked current models in our database, the median output rate is $3.75 per million tokens. At $6.00 per million output tokens, Grok 4.6 commands a 60 percent premium over that market median. Buyers evaluating grok 4.6 pricing must assess whether its long context performance justifies paying well above the central tendency of contemporary AI endpoints.

Model NameProviderInput Rate / 1MOutput Rate / 1MCached Input / 1MContext Limit
Grok 4.6xAI$2.00$6.00$0.50500k tokens
Grok 4.5xAI$2.00$6.00Not recordedNot recorded
DeepSeek V4 Flash 0731DeepSeek$0.14$0.28Not recordedNot recorded
Nemotron 3.5 LightningNVIDIA$0.10$0.25Not recordedNot recorded
Qwen3.7 FlashQwen$0.03$0.13Not recordedNot recorded
Token Pricing Comparison Across Selected Current Models

How much can prompt caching cut your effective cost per 1m tokens?

Prompt caching cuts the Grok 4.6 input rate from $2.00 down to $0.50 per 1M tokens, delivering a 75 percent discount on cached context. For applications with static system prompts, large technical documentation, or long conversation histories, this mechanism dramatically lowers the effective cost per 1m tokens.

To see how this math plays out in production, consider a system processing 100 million input tokens and 20 million output tokens per month. Under standard uncached rates, 100 million inputs cost $200.00 (100 multiplied by $2.00), while 20 million outputs cost $120.00 (20 multiplied by $6.00), resulting in a total monthly bill of $320.00.

If that same system achieves a 100 percent prompt cache hit rate on its inputs, the 100 million input tokens cost just $50.00 (100 multiplied by $0.50). Adding the $120.00 output charge brings the total monthly cost down to $170.00. That reflects an overall expenditure reduction of nearly 47 percent on the total API bill. Even with partial cache utilization, understanding grok 4.6 pricing with caching enabled is essential for high-throughput production planning.

How does Grok 4.6 pricing stack up against high-speed flash models?

Grok 4.6 is priced substantially higher than high-speed flash models, which offer unit rates that are a fraction of xAI charges. Evaluating grok 4.6 vs budget alternatives shows that models from NVIDIA, DeepSeek, and Qwen run at vastly lower price points for both input and output processing.

For instance, Qwen3.7 Flash by Qwen charges $0.03 per 1M input tokens and $0.13 per 1M output tokens. Nemotron 3.5 Lightning by NVIDIA charges $0.10 per 1M input tokens and $0.25 per 1M output tokens. DeepSeek V4 Flash 0731 by DeepSeek charges $0.14 per 1M input tokens and $0.28 per 1M output tokens.

To highlight the financial disparity, consider running the same monthly volume calculation of 100 million input tokens and 20 million output tokens across these alternatives:

  • Qwen3.7 Flash: (100 multiplied by $0.03) plus (20 multiplied by $0.13) equals $3.00 plus $2.60 for a total of $5.60 per month.
  • Nemotron 3.5 Lightning: (100 multiplied by $0.10) plus (20 multiplied by $0.25) equals $10.00 plus $5.00 for a total of $15.00 per month.
  • DeepSeek V4 Flash 0731: (100 multiplied by $0.14) plus (20 multiplied by $0.28) equals $14.00 plus $5.60 for a total of $19.60 per month.
  • Grok 4.6 (uncached): (100 multiplied by $2.00) plus (20 multiplied by $6.00) equals $200.00 plus $120.00 for a total of $320.00 per month.

Which specific production workloads justify paying for Grok 4.6?

Workloads that require continuous generation up to 128k output tokens and massive contextual ingestion up to 500k tokens justify paying for Grok 4.6. The model capabilities recorded in our pricing feed enable deep analysis across expansive datasets that exceed smaller context limits.

While GPT-5.6 Luna Pro currently holds the widest context window in our dataset at 1050k tokens, Grok 4.6 offers a 500k context window paired with a generous 128k maximum output token threshold. This makes Grok 4.6 suitable for generating whole codebases, drafting long legal documents, or summarizing dense technical books in a single pass.

Teams running enterprise pipelines where complete single-pass synthesis matters more than unit token price will find Grok 4.6 viable. However, to maintain budget control, these pipelines must aggressively utilize prompt caching to keep input token charges near $0.50 per million tokens rather than $2.00.

When should enterprise engineering teams avoid Grok 4.6?

Enterprise teams should avoid Grok 4.6 for high-volume utility tasks, simple data extraction, or lightweight routing pipelines. Paying $2.00 per million input tokens and $6.00 per million output tokens for simple classification or short conversational queries wastes infrastructure budget when alternative models deliver similar utility at pennies on the dollar.

When analyzing grok 4.6 pricing for utility tasks, the financial gap is stark. If your application averages 200 input tokens and 50 output tokens per request, routing millions of those queries through Grok 4.6 quickly accumulates thousands of dollars in spend that could be avoided by selecting models like Qwen3.7 Flash or Nemotron 3.5 Lightning.

Furthermore, if your application requires context beyond 500k tokens, Grok 4.6 cannot fulfill the request regardless of price. Applications needing context approaching 1 million tokens must look to models like GPT-5.6 Luna Pro, which provides 1050k tokens of context.

Frequently asked questions

What is the standard grok 4.6 pricing on xAI?
Grok 4.6 costs $2.00 per 1M input tokens and $6.00 per 1M output tokens. Standard cached input tokens are billed at $0.50 per 1M tokens, offering a 75 percent discount on input fees.
How does grok 4.6 api cost compare to Grok 4.5?
Grok 4.6 maintains the exact same baseline pricing as Grok 4.5. Both models are listed at $2.00 per 1M input tokens and $6.00 per 1M output tokens in our pricing tracking database.
What is the maximum context window and output generation for Grok 4.6?
Grok 4.6 features a context window of 500k tokens and can generate up to 128k output tokens in a single request. This makes it specialized for heavy long context tasks.
How does Grok 4.6 pricing compare to Qwen3.7 Flash and Nemotron 3.5 Lightning?
Grok 4.6 is significantly more expensive than flash tier alternatives. Qwen3.7 Flash charges $0.03 input and $0.13 output per 1M tokens, while Nemotron 3.5 Lightning charges $0.10 input and $0.25 output per 1M tokens.
What model currently holds the widest context in the AIOPLY dataset?
GPT-5.6 Luna Pro currently holds the widest context window in our database at 1050k tokens, compared to the 500k token context window provided by Grok 4.6.

About the author

AIOPLY Pricing Desk

Independent AI cost research, verified against live provider rates

How we work

Quality standards

Prices were checked against provider documentation on Aug 18, 2026. Rates change without notice, so confirm current figures with the provider before committing a budget. We publish list prices only and take no payment for placement.

Report a correction

Sources and further reading

More from the blog