Pega's high-profile 'No Token Tax' announcement represents a significant shift in enterprise AI commercial models: bundling variable generative AI inference costs into a fixed business transaction price per resolved case in Infinity 26. However, public evidence indicates that Pega has not eliminated underlying model inference costs or replaced frontier foundation models; rather, it has wrapped third-party LLMs inside a governed workflow and decisioning framework.

Core Architectural Finding: Pega's offering is governed business automation with external AI costs bundled into a business transaction unit. It offers commercial billing predictability, but underlying model economics, integration boundaries, and total operating costs still require rigorous validation.

Executive Evaluation: Published Claims vs. Technical Reality

Key Buyer QuestionTechnical Research FindingArchitectural Boundary
Is the no-token-pricing offer real?Pega charges a flat fee per resolved case in Infinity 26 across Pega Cloud and customer-managed cloud.High certainty for published offer; individual terms depend on contract schedules.
Does fixed cost mean unlimited usage?The billing unit is a completed case; total spend still scales directly with business case volume.Unit price predictability is distinct from annual budget predictability.
Does Pega train its own LLMs?Pega routes requests through a GenAI Gateway to third-party models (GPT, Claude, Gemini).Pega provides workflow orchestration, not a proprietary frontier foundation model.
Are external models open-source?The documented catalog includes commercial frontier models via Azure OpenAI, Bedrock, and Vertex AI.Commercial API costs are absorbed into the case fee, not replaced by free models.
Does it eliminate hallucinations?Structured execution constrains outputs, but residual LLM hallucination risks remain.Pega documentation explicitly notes customer responsibility for output validation.

Dissecting the Model Architecture: Gateway, Decisioning, and LLMs

Pega Academy documentation describes GenAI as a managed bridge connecting Infinity workflows to external model services. The platform handles prompt construction, context retrieval, gateway routing, and response parsing:

  • OpenAI GPT Series: Accessed via Azure OpenAI Service within customer or Pega-managed tenants.
  • Anthropic Claude Series: Routed through Amazon Bedrock for enterprise reasoning tasks.
  • Google Gemini Series: Integrated via Google Cloud Vertex AI for multimodal processing.
  • Pega Adaptive Decisioning: Proprietary Bayesian and gradient-boosting models within Customer Decision Hub (CDH) used for next-best-action selection, distinct from generative text models.

The Architectural Argument: Deterministic Logic vs. Probabilistic LLMs

The economic rationale behind case-based pricing is eliminating wasteful model calls. In a well-architected business workflow, generative AI is restricted to unstructured interpretation, while deterministic rules handle calculations and state transitions:

Workflow StepAppropriate MechanismArchitectural Rationale
Customer Intent IdentificationBounded LLM ClassificationNatural language inputs are unstructured and ambiguous.
Account & Order RetrievalAuthenticated API & System of RecordDeterministic lookup prevents hallucinated data.
Policy Eligibility CheckValidated Business RulesFixed policy rules should not be re-interpreted by LLMs on each run.
Document ExtractionDocument AI with Confidence ScoringUnstructured invoices or claims require OCR and semantic extraction.
Exception AuthorizationHuman-in-the-Loop Approval GateFinancial authorities must remain auditable and deterministic.
Disbursement & PaymentIdempotent Transaction APITransactional execution must strictly prevent duplicate payments.

Total Cost of Ownership: Platform Case Billing vs. Custom API Architecture

Enterprise buyers must evaluate total cost of ownership across the complete operating model rather than comparing raw token invoices to enterprise platform licenses:

TCO Formula: Annual Cost = Contracted Platform Fees + Non-included Tooling + Annualized Delivery/Engineering + Ongoing Operations + Human Review & Rework.
Annual Cost ComponentCase-Priced Platform (e.g., $1/case)Custom API Architecture (100k cases)
Platform / Case License Fee$100,000$0 (Self-managed)
Raw LLM Inference TokensIncluded in Case Fee$4,500 (5 calls/case @ $2/M tokens)
Hosting, Observability & Gateway$10,000$30,000 (Cloud infra, vector DBs, APM)
Annualized Delivery & Maintenance$40,000$140,000 (Dedicated engineering team)
Human Handling & Exception Review$100,000$100,000
Total Annual Expenditure$250,000$274,500

While raw API tokens represent a minor expense ($4,500), building and maintaining custom governance, security guardrails, and audit trails adds significant operational overhead ($170,000). However, if an enterprise already maintains existing custom platforms, the comparison can swiftly reverse.

Industry Comparison: Beyond Raw Token Billing

Vendor / PlatformPublished Billing UnitStrategic Implications
Pega Infinity 26Flat fee per resolved caseExamine definitions of parent/child cases, reopenings, and model catalog limits.
Salesforce AgentforceFlex Credits per action ($2/conversation option)Fine-grained attribution required; underlying Data 360 queries consume separate credits.
Fin AI$0.99 per resolution + base platform feeRequires strict contractual definition of automated resolution vs human handoff.
Custom Model GatewayDirect token consumption (Input/Output/Cache)Full architectural flexibility; shifts maintenance and compliance burden to internal IT.

The 22-Point Procurement Due Diligence Checklist

  1. Order Form & SKU Alignment: Exactly which SKU, license tier, and order form clauses govern the case-based pricing offer?
  2. Baseline vs. Add-On: Does the pricing replace current core seat licenses, or is it structured as an incremental add-on?
  3. Trigger Event Definition: What precise technical state transition marks a case as 'resolved' and billable?
  4. Parent & Child Case Hierarchy: Are sub-cases, nested tasks, or child workflows billed as separate individual transactions?
  5. Reopenings & Exceptions: How are reopened cases, customer retries, timeouts, and cancellations billed?
  6. Human Handoff Accounting: Is a case that escalates to a human agent billed at the full automated rate?
  7. Batch & High-Volume Processing: How are asynchronous batch workloads and recurring scheduled runs counted?
  8. Commitments & Expirations: What minimum annual volume commitments, overage rates, or expiration terms apply?
  9. Included Model Catalog: Which specific model families and versions (e.g., GPT-4o, Claude 3.5 Sonnet) are included without surcharges?
  10. Fair-Use & Concurrency Caps: What token-per-case, payload size, or concurrency limits apply before throttling occurs?
  11. BYOM & Third-Party APIs: Who absorbs costs for customer-hosted models, custom MCP tools, or third-party web services?
  12. Auxiliary Service Inclusion: Are OCR, document ingestion, vector embeddings, and speech transcription covered in the fee?
  13. Non-Production Environments: How are Dev, QA, Staging, Performance Testing, and DR environments licensed?
  14. Specialized AI Capabilities: Are Pega Blueprint, Coach, Knowledge Agent, and Customer Decision Hub covered by the same rate?
  15. Model Deprecation & Substitution: What notification periods, quality floors, and rollback mechanisms apply if Pega replaces a model?
  16. Model Version Pinning: Can enterprises lock an approved, audited model version for compliance repeatability?
  17. Data Residency & Processing Locations: Exactly which regions process and store prompt telemetry and audit logs?
  18. Security & Zero-Retention Terms: Do all underlying model provider agreements enforce zero-day retention for enterprise data?
  19. End-to-End SLAs: What latency, availability, and throughput commitments cover the entire workflow, not just the LLM endpoint?
  20. Invoice Dispute & Audit Telemetry: How are incorrect automated actions, incomplete runs, and invoice discrepancies reconciled?
  21. Renewal Escalation Caps: What contract caps govern annual price escalations upon multi-year renewal?
  22. Exit Rights & Asset Portability: Can case definitions, business rules, prompts, and evaluation datasets be exported cleanly upon platform exit?

Decision Matrix: Who Benefits Most from Case-Based Pricing?

  • High Fit: Enterprises with existing, mission-critical Pega workflows running high-volume, repeatable, rule-bound operations (e.g., claims intake, dispute resolution).
  • High Fit: Regulated organizations requiring centralized decision audits, predictable unit economics, and unified case lifecycle management.
  • Moderate Fit: Organizations requiring extensive integration with external agent ecosystems (MCP) and third-party SaaS APIs, where bundled models cover only part of the stack.
  • Low Fit: Low-volume ad-hoc document drafting or open-ended analytical research where simple pay-as-you-go API consumption is orders of magnitude cheaper.

How YuniQ Accelerates Pega Modernization and AI Strategy

YuniQ's AI-Driven Enterprise Modernization practice helps organizations evaluate, modernize, and extract business logic from legacy BPM platforms:

  • Pega Architecture & TCO Audits: Benchmarking case-based pricing vs. custom API architectures on verified production workloads.
  • Workflow & Rule Extraction: Reverse-engineering undocumented rules, case lifecycles, and data transforms for modern cloud architectures.
  • Proof-of-Value Benchmarking: Running three-way comparisons across legacy Pega, upgraded Infinity 26, and modern cloud microservices.
  • Hybrid AI Governance: Designing deterministic policy guardrails and audit logging across multi-model enterprise gateways.
Executive Recommendation: Base your Pega AI decision on empirical unit economics: verify your contract's billable-case definition, audit the included model catalog, and measure the cost per verified successful business outcome.