A follow-up to Stop Limiting Your Thinking to Columns and Rows.

Most “public AI vs private AI” cost comparisons online are shallow: they put a $20/month subscription next to a $3,000 box and declare the subscription cheaper. That’s not a fair comparison, and it skips the two questions that actually matter to a business handling real customer and contract data — where’s the genuine break-even point, and what does it cost you if something goes wrong along the way?

This analysis takes one realistic task, prices it against the actual published numbers for both sides — down to the per-token API cost of the models involved — and models the point at which volume tips the economics from public to private. It also looks at the legal and regulatory precedent around what happens when sensitive business data goes into a public chatbot, because that risk has a real, quantifiable cost of its own.

The test case

“Review this 300-page legal rules and regulations document and highlight any legal gotchas that could risk the relationship with our top ten customers.”

A 300-page document is roughly 150,000 words, or about 200,000 tokens — comfortably inside the context window of every current frontier model, public or private. That makes it a genuine apples-to-apples test, unlike tasks that require live access to internal systems (CRM, order history, contracts) that no public chatbot can reach at all.

What it actually costs, token by token

Subscription apps (ChatGPT, Claude.ai, Gemini) don’t publish a per-token rate — that’s the point of a flat monthly fee. But every one of them is a wrapper around models that do have published API pricing, and that’s the real, verifiable cost of the underlying compute:

Model (used inside plan) Input $/1M tokens Output $/1M tokens Plan tier it’s included in
GPT-5 $1.25 $10.00 ChatGPT Plus and above
GPT-5.6 “Luna” (light) $0.20 $1.20 Plus (lower-cost fallback model)
GPT-5.6 “Terra” / “Sol” (mid-high reasoning) $2.00–$5.00 $12.00–$30.00 Higher reasoning modes gated to Pro
Claude Sonnet 5 $2.00–$3.00* $10.00–$15.00* Claude Pro and above
Claude Opus 5 $5.00 $25.00 Claude Max only
Gemini 3.1 Pro $2.00 (up to 200K tokens) $12.00 Google AI Pro and above
Gemini 3.5/3.6 Flash $1.50 $7.50–$9.00 Included across most Google AI tiers

*Sonnet 5’s $2/$10 rate is introductory pricing in effect through 31 August 2026, after which the standard $3/$15 applies.

At those rates, running our 300-page document through the API directly (~200,000 input tokens + ~5,000 output tokens) costs roughly $0.45–$0.70 per document, regardless of which frontier provider you use. On raw compute cost alone, public AI is extremely cheap — that has to be said plainly, because it’s the honest number.

Where the plans actually differ is not price per token, it’s which model you’re allowed to call and how fast you burn your usage allowance. Plus-tier plans default you to lighter, cheaper models and gate the heaviest reasoning models (the ones best suited to a dense 300-page legal read) behind Pro/Max/Ultra tiers. A model like Opus 5 or GPT-5.6’s top reasoning mode consumes usage allowance far faster than a lightweight model — so the same document, run on a top-tier model, can eat a large share of a daily/five-hourly quota in one sitting. In practice, a single heavy document review at the quality level this task needs can be enough to hit Plus/Pro-tier limits, pushing regular users of this kind of task toward the $100–$200/month tiers.

Where does volume tip the balance?

This is the real question, and it doesn’t have a single universal answer because none of the providers publish an exact “N documents per month” cap — usage limits are dynamic, based on compute consumed. But modelling it from the published tier language and the token costs above gives a reasonable picture:

Monthly volume of this task Likely plan tier needed Annualised public AI cost Annualised DGX Spark cost*
1–10 documents/month Plus / Pro ($20/mo) ~$240/yr ~$1,716/yr
10–40 documents/month Max / Ultra ($100/mo) ~$1,200/yr ~$1,716/yr
40+ documents/month, heavy/regular Max 20x / Ultra ($200/mo) ~$2,400/yr ~$1,716/yr

*DGX Spark: $4,699 hardware (NVIDIA’s revised Founders Edition MSRP as at February 2026, up from $3,999) amortised over 3 years + ~$150/yr power ≈ $1,716/yr, fixed regardless of volume, with zero marginal cost per document and nothing leaving the building. A ~$3,000 workstation-class alternative works out closer to $1,150/yr on the same basis.

On pure compute cost, the crossover sits at the point where usage limits push you past the $100/month tier and into the $200/month bracket — around 40 heavy documents a month, or roughly two a working day. Below that, a public subscription is genuinely the cheaper option and there’s no honest way to argue otherwise. Above it, private hardware wins on cost too, on top of removing the data-exposure question entirely. And it’s worth being clear about the shape of the two curves: the public cost steps up each time volume forces a tier upgrade, while the hardware cost is flat — so every document past the crossover widens the gap.

But cost-per-token was never really the full story for this kind of task — which brings us to the part of the analysis that a pure pricing table leaves out.

The part a token-cost comparison misses: what happens when data leaves the building

This task means feeding unredacted contracts, regulatory text and top-customer risk analysis into a third party’s servers. There’s real, documented precedent for what can go wrong:

  • Samsung, 2023: engineers in Samsung’s semiconductor division uploaded confidential source code and internal meeting notes to ChatGPT on three separate occasions. Samsung banned generative AI tools company-wide as a result. (Bloomberg, TechCrunch, May 2023).
  • Italy’s data protection authority, 2023: the Garante temporarily blocked ChatGPT nationwide over a data breach and unlawful collection of personal data, forcing OpenAI to change its data practices before being allowed back. (AP News, March 2023).
  • Google indexing “shared” ChatGPT conversations, 2025: a short-lived sharing feature resulted in private ChatGPT conversations — some containing sensitive personal and business details — being indexed and surfaced in public search results before OpenAI pulled the feature. (TechCrunch, Fast Company, July–August 2025).
  • United States v. Heppner, SDNY, February 2026: in a ruling directly relevant to this task, a federal court held that documents a defendant generated using a consumer version of a public AI chatbot (Claude) were not protected by attorney-client privilege or work-product doctrine — because feeding case information into a public AI tool was treated the same as disclosing it to a third party. Legal commentators have since warned that any confidential or legally sensitive material — including exactly the kind of contract and compliance analysis in our test case — could lose its legal protection the moment it’s entered into a public AI tool. (Proskauer Rose, Orrick, Perkins Coie, Law.com — March–April 2026).

That last precedent matters specifically for this article’s test case: a “highlight legal gotchas in top-customer contracts” task is, by definition, legally sensitive work. Running it through a public chatbot doesn’t just carry a generic privacy risk — under current case law, it can strip away legal protections the business may need later.

The practical takeaway

For light, occasional use — summarising a public document, drafting routine correspondence — a $20/month public AI plan remains the right, and cheaper, choice. That’s a straight, honest reading of the token economics.

For a recurring task like this one — legal/contract review touching the customers who matter most to revenue — three things point the same direction at once: usage limits push regular users into the $100–200/month tiers within a normal month’s volume, private hardware reaches cost parity or better once you’re at the top of that range, and the legal/regulatory precedent means the “free” option carries a real, demonstrated cost if something goes wrong.

That’s the actual decision framework worth applying before choosing a side: not “is private AI cheaper than ChatGPT,” but “at what volume, and at what sensitivity, does the balance genuinely tip” — and for businesses handling real customer and contract data on a recurring basis, both the cost curve and the legal precedent suggest that point arrives sooner than most assume.

Convergence works with businesses evaluating Private AI infrastructure for exactly this kind of decision — get in touch if you’d like the numbers modelled against your own volume.

Social media & sharing icons powered by UltimatelySocial