The Fastest-Depreciating Purchase You'll Ever Make
There has never been a stranger market than the one for artificial intelligence in 2026. In most industries, when you buy the newest thing, you pay a premium for being early and the price drifts down slowly over years. AI has inverted that rule. The capability you pay top dollar for today is frequently available for a fraction of the cost within a single quarter, and often from the very same provider.
Consider the headline number that frames this entire discussion: the price of GPT-4-level quality has fallen from $30 per million tokens in March 2023 to under $0.50 by mid-2026, a decline of roughly 98%. By some benchmark measures, the same tier of intelligence now costs over 280 times less than it did in early 2023.
~98% cheaper in three years The cost of GPT-4-class capability collapsed from $30 to under $0.50 per million tokens between March 2023 and 2026 (TokenCost AI Price Index / BenchLM). |
That single fact creates the central tension of this article. On one hand, waiting almost always means paying less for more. On the other, the business world is full of warnings that waiting is the most expensive decision of all, that competitors who adopt now accumulate months of compounding advantage that latecomers can never fully claw back. Both statements are true. The skill is knowing which one applies to your specific decision.
This guide breaks the question into its component parts, backed by current market data, and gives you a practical framework for deciding, purchase by purchase, whether to pay now or wait three months.
The Data: Why Waiting Usually Means Paying Less
Before deciding anything, it helps to internalise just how fast the ground is moving. Three trends dominate the economics of AI pricing right now.
Token prices are in freefall
The raw cost of AI computation measured in price per million tokens has been dropping at a pace that resembles a steeper version of Moore's Law. Analysts estimate that the cost of a given level of capability now halves roughly every six to eight months, and unlike Moore's Law there is no obvious physical limit yet slowing it down.
In the ten weeks leading into August 2026 alone, average inference costs fell about 43%, from roughly $2.04 to between $1.16 and $1.18 per million tokens. Two forces drove that: aggressive price cuts from Western labs (one provider slashed a flagship model's pricing by 80% in a single month), and a relentless wave of low-cost open-weight models from Chinese labs, some of which achieved price reductions of up to 99% earlier in the year.

Figure 1 — The price of frontier-level quality has collapsed by roughly 98% since GPT-4's launch.
New and better models arrive constantly
The second reason waiting pays is capability, not just price. The release calendar has become almost absurdly dense. Major labs now ship a new flagship roughly every 6–12 months, with smaller point upgrades landing every few weeks. Across the whole industry, a notable model now arrives on average every two days, and 2025 alone saw more than 25 significant releases.
What this means in practice: the tool that frustrates you today with a limitation may simply not have that limitation in ninety days at the same price or lower. Paying a premium to be an early tester of a rough feature often means paying to endure bugs that the next version quietly fixes for free.

Figure 2 — Average gap between model releases by provider. A meaningful upgrade is rarely more than a quarter away.
The current consumer price landscape
For individuals and small teams, the sticker prices have converged tightly. The three leading assistants now cost within a dollar of one another for their standard paid tier, while cheaper entry tiers and much pricier power-user tiers bracket that midpoint.
| Plan / Tier | Approx. price (2026) | Who it's for |
| ChatGPT Plus | $20 / month | Standard individual use |
| Claude Pro | $20/mo ($17 annual) | Individual; includes Claude Code & Cowork |
| Google AI Pro | $19.99 / month | Individual; deep Google integration |
| Entry tiers | $5–$8 / month | Light users (Gemini AI Plus, ChatGPT Go) |
| Power tiers | $100–$300 / month | Heavy / professional workloads |
| Team seats | $25–$30 / user/mo | Collaborative business use |
Table 1 — Representative 2026 consumer pricing. Verify current figures on providers' official pages before buying, as prices change multiple times a year.
The Counter-Data: Why Waiting Can Be the Costly Choice
If prices only fall and models only improve, the naive conclusion would be to wait forever. That conclusion is wrong, and the data shows why.
Cheaper per unit does not mean cheaper overall
Here is the paradox that pins CFOs to their chairs: even as blended token prices fell roughly 67% year-over-year from $18.40 to $6.07 per million tokens between Q1 2025 and Q1 2026, total AI bills went up, not down. Enterprise spend on large language models more than doubled in six months, from $3.5 billion in late 2024 to $8.4 billion by mid-2025.
The reason is that usage explodes faster than prices fall. Google reported going from 9.7 trillion tokens a month to over 3.2 quadrillion in two years, roughly a 330x jump. When capability gets cheaper, people don't spend less; they do far more. Notably, 73% of enterprises exceeded their original AI cost projections last fiscal year, and one major company reportedly burned its entire annual AI coding budget in just four months.

Figure 3 — Prices per token fell sharply while total spend more than doubled: consumption grows faster than prices drop.
The invisible cost of waiting
The most dangerous cost of waiting never appears on an invoice. Business analysts are near-unanimous that delay carries a real, if hidden, price: competitors who adopt now accumulate months of learning, workflow refinement, and data advantage that late adopters must rebuild from scratch. One widely cited survey found early adopters seeing 20–30% operational efficiency gains, with late adopters struggling to close the gap even with accelerated investment.
As one strategist put it, the person who waits for certainty is always outrun by the person who starts messy and learns in motion. Generative AI adoption jumped from under 1% of firms in 2022 to more than 11% by 2024, the window to be meaningfully early is closing, even as the tools get cheaper.
73% of enterprises blew their AI budget Per-token prices fell 67% year-over-year, yet nearly three-quarters of organisations still exceeded their AI cost projections because usage grew faster than prices fell. |
A Framework: The Pay-Now vs. Wait Decision Matrix
The way to reconcile these opposing truths is to stop asking "is AI getting cheaper?" (it is) and start asking two better questions about each specific purchase:
- How much measurable value does this deliver to me today? Not in theory, in dollars saved, hours reclaimed, or revenue enabled this month.
- How fast is this specific thing depreciating? A month-to-month subscription and a three-year enterprise contract carry wildly different obsolescence risk.
Plotting any AI decision against those two axes produces four clear zones.

Figure 4 — Map any AI purchase onto value-today versus depreciation risk to see which zone it falls in.
Pay now (high value, durable)
If a tool delivers real, measurable value today and isn't at high risk of being obsoleted or drastically repriced against you, pay now. Core workflow automation and niche vertical tools with no real substitute belong here. The savings from waiting are trivial compared to the value you forfeit each month you delay.
Pay, but stay flexible
The $20 flagship chat subscription and a daily-use coding assistant are high-value but sit in fast-moving categories. The answer isn't to wait, it's to pay month-to-month rather than locking into annual commitments, so you can switch providers the moment a better or cheaper option appears. The price convergence around $20 means switching costs are low and your leverage is high.
Wait about three months (overpaying to be a beta tester)
Bleeding-edge feature add-ons and premium tiers priced far above the mainstream fall here. When you're paying a steep premium for a capability that is improving weekly and dropping in price monthly, waiting a quarter is often the rational choice, the next release cycle frequently delivers the same feature, more stable, for less. Long multi-year enterprise contracts signed at 2024–2025 rates are especially suspect; those rates are almost certainly now above market, and providers will renegotiate to keep you.
Okay to wait (no downside yet)
Speculative "we might use this someday" tools with no measurable value today cost you nothing to postpone. Waiting here isn't indecision, it's discipline. Let early adopters shake out the bugs while you watch.
Practical Rules of Thumb
Translating the matrix into everyday habits, a few rules hold up well against the 2026 data:
| Situation | Recommended move |
| A tool saves you real hours every week | Pay now, the value dwarfs any waiting discount. |
| You're tempted by an annual plan to save 15% | Prefer monthly; flexibility is worth more than the discount in a fast-moving market. |
| A shiny new premium feature just launched | Wait one release cycle (~3 months); expect it cheaper and more stable. |
| A vendor wants a 2–3 year commitment | Push back hard or decline; today's rates will look expensive in a year. |
| You signed an AI contract in 2024/early 2025 | Renegotiate now, you're very likely above current market pricing. |
| A competitor is visibly pulling ahead with AI | Don't wait; the compounding-advantage cost of delay is real. |
| You can't name the dollar value a tool adds | Wait , run a small, cheap pilot before committing budget. |
Table 2 — Quick-reference decisions grounded in current market behaviour.
The 90-day pilot: the best of both worlds
The single most useful tactic that resolves the pay-versus-wait dilemma is the small, time-boxed pilot. Pick one bottleneck that costs you money, the most expensive problem, not the most interesting one. Measure what it costs you today. Apply an AI tool to just that one thing, give it ninety days, and check the number. If it moved, expand and pay confidently. If it didn't, you've spent one quarter and a modest budget learning something concrete, far cheaper than either a year of deliberation or a rushed enterprise-wide rollout.
Conclusion: Time the Purchase, Not the Market
The instinct to wait for AI to get cheaper is grounded in real, dramatic data, prices really are collapsing, and better models really do arrive within weeks. But that same data contains the warning: the money is not in pocketing the savings, it is in reinvesting the falling costs into doing more, sooner, than your competitors.
So the honest answer to "when should I pay for AI and when should I wait three months?" is not a date. It is a discipline:
- Pay now for anything delivering measurable value today that you'll keep using regardless of which model wins.
- Stay flexible on fast-moving subscriptions, month-to-month, never locked in.
- Wait a quarter on premium features and long contracts where you'd be paying a premium to beta-test something that will be cheaper and better by the next release.
Trying to time the bottom of AI pricing is a losing game, there is no bottom in sight. But timing each individual purchase against its value and its depreciation risk is entirely winnable. Do that consistently, and you capture the upside of falling prices without paying the invisible, compounding cost of standing still.
Comments 0
Join the discussion and share your perspective.
Sign in to post a comment and reply to other readers.
No comments yet
Be the first to share your perspective on this article.