Pega's high-profile 'No Token Tax' announcement represents a significant shift in enterprise AI commercial models: bundling variable generative AI inference costs into a fixed business transaction price per resolved case in Infinity 26. However, public evidence indicates that Pega has not eliminated underlying model inference costs or replaced frontier foundation models; rather, it has wrapped third-party LLMs inside a governed workflow and decisioning framework.
Core Architectural Finding: Pega's offering is governed business automation with external AI costs bundled into a business transaction unit. It offers commercial billing predictability, but underlying model economics, integration boundaries, and total operating costs still require rigorous validation.
Executive Evaluation: Published Claims vs. Technical Reality
| Key Buyer Question | Technical Research Finding | Architectural Boundary |
|---|---|---|
| Is the no-token-pricing offer real? | Pega charges a flat fee per resolved case in Infinity 26 across Pega Cloud and customer-managed cloud. | High certainty for published offer; individual terms depend on contract schedules. |
| Does fixed cost mean unlimited usage? | The billing unit is a completed case; total spend still scales directly with business case volume. | Unit price predictability is distinct from annual budget predictability. |
| Does Pega train its own LLMs? | Pega routes requests through a GenAI Gateway to third-party models (GPT, Claude, Gemini). | Pega provides workflow orchestration, not a proprietary frontier foundation model. |
| Are external models open-source? | The documented catalog includes commercial frontier models via Azure OpenAI, Bedrock, and Vertex AI. | Commercial API costs are absorbed into the case fee, not replaced by free models. |
| Does it eliminate hallucinations? | Structured execution constrains outputs, but residual LLM hallucination risks remain. | Pega documentation explicitly notes customer responsibility for output validation. |
Dissecting the Model Architecture: Gateway, Decisioning, and LLMs
Pega Academy documentation describes GenAI as a managed bridge connecting Infinity workflows to external model services. The platform handles prompt construction, context retrieval, gateway routing, and response parsing:
- OpenAI GPT Series: Accessed via Azure OpenAI Service within customer or Pega-managed tenants.
- Anthropic Claude Series: Routed through Amazon Bedrock for enterprise reasoning tasks.
- Google Gemini Series: Integrated via Google Cloud Vertex AI for multimodal processing.
- Pega Adaptive Decisioning: Proprietary Bayesian and gradient-boosting models within Customer Decision Hub (CDH) used for next-best-action selection, distinct from generative text models.
The Architectural Argument: Deterministic Logic vs. Probabilistic LLMs
The economic rationale behind case-based pricing is eliminating wasteful model calls. In a well-architected business workflow, generative AI is restricted to unstructured interpretation, while deterministic rules handle calculations and state transitions:
| Workflow Step | Appropriate Mechanism | Architectural Rationale |
|---|---|---|
| Customer Intent Identification | Bounded LLM Classification | Natural language inputs are unstructured and ambiguous. |
| Account & Order Retrieval | Authenticated API & System of Record | Deterministic lookup prevents hallucinated data. |
| Policy Eligibility Check | Validated Business Rules | Fixed policy rules should not be re-interpreted by LLMs on each run. |
| Document Extraction | Document AI with Confidence Scoring | Unstructured invoices or claims require OCR and semantic extraction. |
| Exception Authorization | Human-in-the-Loop Approval Gate | Financial authorities must remain auditable and deterministic. |
| Disbursement & Payment | Idempotent Transaction API | Transactional execution must strictly prevent duplicate payments. |
Total Cost of Ownership: Platform Case Billing vs. Custom API Architecture
Enterprise buyers must evaluate total cost of ownership across the complete operating model rather than comparing raw token invoices to enterprise platform licenses:
TCO Formula: Annual Cost = Contracted Platform Fees + Non-included Tooling + Annualized Delivery/Engineering + Ongoing Operations + Human Review & Rework.
| Annual Cost Component | Case-Priced Platform (e.g., $1/case) | Custom API Architecture (100k cases) |
|---|---|---|
| Platform / Case License Fee | $100,000 | $0 (Self-managed) |
| Raw LLM Inference Tokens | Included in Case Fee | $4,500 (5 calls/case @ $2/M tokens) |
| Hosting, Observability & Gateway | $10,000 | $30,000 (Cloud infra, vector DBs, APM) |
| Annualized Delivery & Maintenance | $40,000 | $140,000 (Dedicated engineering team) |
| Human Handling & Exception Review | $100,000 | $100,000 |
| Total Annual Expenditure | $250,000 | $274,500 |
While raw API tokens represent a minor expense ($4,500), building and maintaining custom governance, security guardrails, and audit trails adds significant operational overhead ($170,000). However, if an enterprise already maintains existing custom platforms, the comparison can swiftly reverse.
Industry Comparison: Beyond Raw Token Billing
| Vendor / Platform | Published Billing Unit | Strategic Implications |
|---|---|---|
| Pega Infinity 26 | Flat fee per resolved case | Examine definitions of parent/child cases, reopenings, and model catalog limits. |
| Salesforce Agentforce | Flex Credits per action ($2/conversation option) | Fine-grained attribution required; underlying Data 360 queries consume separate credits. |
| Fin AI | $0.99 per resolution + base platform fee | Requires strict contractual definition of automated resolution vs human handoff. |
| Custom Model Gateway | Direct token consumption (Input/Output/Cache) | Full architectural flexibility; shifts maintenance and compliance burden to internal IT. |
The 22-Point Procurement Due Diligence Checklist
- Order Form & SKU Alignment: Exactly which SKU, license tier, and order form clauses govern the case-based pricing offer?
- Baseline vs. Add-On: Does the pricing replace current core seat licenses, or is it structured as an incremental add-on?
- Trigger Event Definition: What precise technical state transition marks a case as 'resolved' and billable?
- Parent & Child Case Hierarchy: Are sub-cases, nested tasks, or child workflows billed as separate individual transactions?
- Reopenings & Exceptions: How are reopened cases, customer retries, timeouts, and cancellations billed?
- Human Handoff Accounting: Is a case that escalates to a human agent billed at the full automated rate?
- Batch & High-Volume Processing: How are asynchronous batch workloads and recurring scheduled runs counted?
- Commitments & Expirations: What minimum annual volume commitments, overage rates, or expiration terms apply?
- Included Model Catalog: Which specific model families and versions (e.g., GPT-4o, Claude 3.5 Sonnet) are included without surcharges?
- Fair-Use & Concurrency Caps: What token-per-case, payload size, or concurrency limits apply before throttling occurs?
- BYOM & Third-Party APIs: Who absorbs costs for customer-hosted models, custom MCP tools, or third-party web services?
- Auxiliary Service Inclusion: Are OCR, document ingestion, vector embeddings, and speech transcription covered in the fee?
- Non-Production Environments: How are Dev, QA, Staging, Performance Testing, and DR environments licensed?
- Specialized AI Capabilities: Are Pega Blueprint, Coach, Knowledge Agent, and Customer Decision Hub covered by the same rate?
- Model Deprecation & Substitution: What notification periods, quality floors, and rollback mechanisms apply if Pega replaces a model?
- Model Version Pinning: Can enterprises lock an approved, audited model version for compliance repeatability?
- Data Residency & Processing Locations: Exactly which regions process and store prompt telemetry and audit logs?
- Security & Zero-Retention Terms: Do all underlying model provider agreements enforce zero-day retention for enterprise data?
- End-to-End SLAs: What latency, availability, and throughput commitments cover the entire workflow, not just the LLM endpoint?
- Invoice Dispute & Audit Telemetry: How are incorrect automated actions, incomplete runs, and invoice discrepancies reconciled?
- Renewal Escalation Caps: What contract caps govern annual price escalations upon multi-year renewal?
- Exit Rights & Asset Portability: Can case definitions, business rules, prompts, and evaluation datasets be exported cleanly upon platform exit?
Decision Matrix: Who Benefits Most from Case-Based Pricing?
- High Fit: Enterprises with existing, mission-critical Pega workflows running high-volume, repeatable, rule-bound operations (e.g., claims intake, dispute resolution).
- High Fit: Regulated organizations requiring centralized decision audits, predictable unit economics, and unified case lifecycle management.
- Moderate Fit: Organizations requiring extensive integration with external agent ecosystems (MCP) and third-party SaaS APIs, where bundled models cover only part of the stack.
- Low Fit: Low-volume ad-hoc document drafting or open-ended analytical research where simple pay-as-you-go API consumption is orders of magnitude cheaper.
How YuniQ Accelerates Pega Modernization and AI Strategy
YuniQ's AI-Driven Enterprise Modernization practice helps organizations evaluate, modernize, and extract business logic from legacy BPM platforms:
- Pega Architecture & TCO Audits: Benchmarking case-based pricing vs. custom API architectures on verified production workloads.
- Workflow & Rule Extraction: Reverse-engineering undocumented rules, case lifecycles, and data transforms for modern cloud architectures.
- Proof-of-Value Benchmarking: Running three-way comparisons across legacy Pega, upgraded Infinity 26, and modern cloud microservices.
- Hybrid AI Governance: Designing deterministic policy guardrails and audit logging across multi-model enterprise gateways.
Executive Recommendation: Base your Pega AI decision on empirical unit economics: verify your contract's billable-case definition, audit the included model catalog, and measure the cost per verified successful business outcome.