Dreamforce 2026 expanded the Agentforce menu. Salesforce announced job-ready agents, long-horizon agents, Agent Optimizer, and additional ways for agents to work across enterprise systems. Some capabilities are generally available, while others are in pilot or planned for later release.

That creates momentum, but it does not answer the question facing many Salesforce leaders: Which business workflow is ready to become a production agent?

A polished demonstration is not enough. It may prove that an agent can answer a scripted question or call an action under controlled conditions. Production requires something different: a repeatable outcome, trustworthy inputs, bounded authority, a working handoff, observable performance, and an owner who will improve the service after launch.

Recent evidence makes the distinction important. In Salesforce's August 2026 account of a company-sponsored study of more than 2,000 AI decision-makers, clean and accessible data and a narrowly scoped use case were the two most-cited success factors, at 36% each. Defined escalation paths followed at 35%. Model quality, platform choice, and unified orchestration ranked lower in that analysis. Salesforce also reported that only 1%–2% of deployed organizations had agents operating without any human involvement. Read the Salesforce analysis .

Independent reporting provides a useful counterweight. TechRadar Pro's August coverage of a TD Cowen partner survey described subdued adoption and persistent concerns about data readiness and agent maturity. Because the underlying research is not available in the article, those figures should be treated as reported market signals, not a verdict on every Agentforce program. Read the TechRadar Pro report .

The practical conclusion is not to wait for perfect data or a flawless model. It is to choose a production job carefully enough that the organization can improve it with evidence.

Start with a job, not an Agentforce feature

“Deploy an Agentforce service agent” is a technology project. “Resolve address-change requests without an agent rekeying data” is a job.

The second statement tells the team who is served, what completion means, which systems must change, and what can be measured. It also exposes the hard questions early. Can the requester be authenticated? Which record is authoritative? Are there exceptions? Must a person approve some changes? What should happen if an integration fails after the conversation appears complete?

Before estimating or configuring anything, write a one-sentence production contract:

For [person or team], when [trigger] occurs, the agent will [complete a business outcome] using [approved data and actions], will escalate when [defined boundary], and will be judged by [baseline and target measures].

For example:

For policyholders, when a routine claim-status inquiry arrives through authenticated chat, the agent will retrieve the current status and next required document from approved systems, escalate disputed or ambiguous records to a claims representative, and be judged by verified completion, repeat-contact rate, handling time, and customer effort.

If a team cannot write this contract, the use case is not ready for an implementation estimate.

Score the use case before funding the build

Use a 0–3 score for each factor: 0 means absent, 1 means materially weak, 2 means workable with defined remediation, and 3 means ready. Multiply the rating by the listed weight and divide by three.

FactorWeightWhat a production-ready score looks like
Business value and baseline20A named outcome has a current cost, delay, error, revenue, experience, or capacity baseline. The expected benefit can be measured without relying only on “deflection.”
Volume and process suitability15Demand is frequent enough to matter, the job has recognizable patterns, and exceptions can be identified. Variation is understood rather than hidden in tribal knowledge.
Data and knowledge readiness15Required sources are named, permitted, sufficiently accurate, current, and retrievable. Conflicting sources have an explicit precedence rule.
Action and integration readiness15Required reads and writes have supported interfaces, appropriate identities, idempotency or duplicate protection, error handling, and transaction confirmation.
Risk boundaries and human handoff15The agent's authority is explicit. Prohibited actions, escalation triggers, queue ownership, context transfer, and response expectations are defined.
Measurement and feedback10The team can observe task completion, failure, correction, escalation, customer outcome, latency, and consumption at the use-case level.
Operating ownership and change capacity10A business owner and technical owner have time, authority, release discipline, and a backlog process for ongoing improvement.

Interpret the total cautiously:

  • 75–100: A reasonable candidate for a limited production release, assuming the mandatory gates below are satisfied.
  • 55–74: Redesign or remediate before committing to production. The score should point to a short, specific readiness backlog.
  • Below 55: Do not compensate with a larger model or more prompt engineering. Choose a narrower outcome or a different workflow.

The score is a decision aid, not a statistical prediction of ROI. Adjust the weights for the organization's risk, economics, and operating model, but do not remove weak categories simply to make a favored project pass.

Apply five mandatory gates

A high score cannot override a missing safety or operating control. Before a limited production release, require a clear “yes” to each gate.

1. Is the data use authorized and appropriate?

Confirm which records, fields, documents, conversation history, and external sources the agent may access. Test effective access through the same identity, channel, and integration path the production agent will use. Data classification, privacy, retention, residency, and sector-specific requirements need qualified review; a Salesforce permission alone does not establish that every use is appropriate.

2. Is the agent's authority bounded?

Separate probabilistic work from deterministic control. An agent may interpret an intent or draft a response, while Flow, Apex, an API, or another governed service enforces eligibility, required fields, approval, transaction limits, and audit behavior. The system should be able to explain which controlled action ran and whether it completed.

3. Does the human handoff work as a service?

“Escalate to a human” is incomplete. Specify the destination, hours of coverage, priority, response target, transcript or summary, record context, and ownership if the customer disconnects. Test whether the human can continue the work without asking the person to start again.

4. Can the organization detect and recover from partial failure?

Plan for the awkward middle: the agent confirms an action, but the downstream system times out; a case is created twice; a source is stale; an appointment is reserved but not committed. Define confirmation, retry, reconciliation, rollback where possible, and human recovery before launch.

5. Is someone accountable for the production outcome?

The owner should control more than the agent configuration. They need authority to change the process, fix knowledge, prioritize integration defects, coordinate security review, and decide when to expand, pause, or retire the use case.

If any gate is “no,” the correct status is redesign, not “pilot extended.”

Measure completed outcomes, not attractive conversations

Conversation count, answer rate, and containment can be useful operational measures, but none proves that the intended job was completed correctly.

Build a measurement tree with four levels:

  1. Business outcome: Did the workflow reduce avoidable effort, delay, error, leakage, or unmet demand? Did it create usable capacity or improve an experience that matters?
  2. Task completion: Did the requested action finish in the system of record? Was it correct, durable, and free of duplicate or compensating work?
  3. Experience and safety: How often did users repeat contact, abandon, request a person, receive a correction, reopen a case, or encounter a policy exception?
  4. Economics and operations: What did each verified completion cost in platform consumption, integration activity, monitoring, human review, and support? How much time is spent maintaining knowledge and controls?

Establish the baseline before launch. Otherwise, an impressive dashboard can show that the agent is busy without showing that the organization is better off.

Choose a narrow first release without creating a dead end

Narrow scope is not the same as a toy use case. The best first production job is bounded but valuable, and its architecture can support adjacent work later.

A useful pattern is:

  • one defined audience;
  • one entry channel;
  • a controlled set of intents;
  • a small number of authoritative sources;
  • a limited action set;
  • explicit exclusions;
  • a tested handoff;
  • and one outcome dashboard.

Salesforce's September Dreamforce recap describes a similar progression: select and deploy a use case, customize and extend it, then observe and optimize. That sequence is credible only when expansion follows measured performance rather than the excitement of a successful demo. Review Salesforce's Dreamforce takeaways .

What good first use cases look like by industry

Nonprofit foundations

Promising starting points include authenticated grant-application status, document-completeness guidance, donor service requests, or internal research support based on an approved knowledge set. These jobs can reduce repetitive inquiry while keeping award decisions, beneficiary eligibility, safeguarding concerns, and sensitive exceptions with authorized people.

A weak use case is “help applicants with grants.” A stronger contract names the grant program, application stage, approved data, actions, excluded advice, escalation team, and a measure such as verified self-service completion or reduced repeat contact.

Healthcare organizations and health insurers

Digital-front-door navigation, provider-directory questions, appointment or referral status, administrative document intake, and authenticated member-service inquiries may be suitable starting points when sources and handoffs are dependable. Clinical advice, emergency situations, coverage determinations, and other high-consequence decisions require much stronger boundaries and qualified review.

Salesforce's September 25 healthcare article highlights trust, context, and action and describes organizations applying agents to navigation and administrative workflows. Those examples are useful patterns, but the reported results belong to the named customers and should not be treated as a performance forecast for another organization. Read Salesforce's healthcare examples .

Insurers

First-notice-of-loss intake, missing-document follow-up, claim-status questions, adjuster scheduling, and agent-assist knowledge retrieval can have clear volume and completion measures. Coverage interpretation, claim adjudication, fraud conclusions, pricing, and underwriting decisions carry greater legal, financial, and fairness implications and should not be casually folded into the same autonomy level.

For every insurance use case, distinguish information retrieval from a decision, and distinguish a drafted recommendation from an executed transaction.

Use a 90-day path with decision points

Days 1–15: Select and baseline

Score three to five candidate workflows, interview frontline teams, map exceptions, and validate the operational baseline. Choose one production contract and document why the others were deferred.

Days 16–30: Prove data, actions, and boundaries

Trace every required field and knowledge source. Run representative reads and writes through the intended identities and interfaces. Define deterministic controls, prohibited actions, escalation rules, and recovery paths.

Days 31–50: Build and test the complete journey

Test normal, ambiguous, adversarial, stale-data, permission, integration-failure, and handoff scenarios. Evaluate whole conversations and completed transactions, not isolated answers. Confirm logging and outcome attribution.

Days 51–70: Release to a limited cohort

Use a restricted audience, channel, volume, or operating window. Staff the handoff and incident paths. Review individual failures frequently enough to change the design while the exposure remains bounded.

Days 71–90: Decide with evidence

Compare verified outcomes with the baseline. Include correction work, human review, platform consumption, support effort, and unresolved risk. Choose one of four decisions: expand, hold, redesign, or stop.

Stopping a weak use case is not a failed AI strategy. It is portfolio discipline that protects funds for a workflow with better economics and operating fit.

When a pilot should be stopped or redesigned

Pause and reconsider when any of the following persists:

  • No one can agree on the business outcome or baseline.
  • The agent answers questions but cannot complete the job.
  • The required source is not authoritative, current, or legally appropriate for the use.
  • Most interactions require exceptions that only experienced staff can resolve.
  • Human handoffs create more effort than the prior process.
  • The workflow depends on unsupported screen behavior or a fragile integration with no recovery path.
  • Success depends on excluding corrections, rework, consumption, or operating labor from the economics.
  • The team cannot identify a business owner after the demonstration succeeds.

Sometimes the right redesign is simple: reduce the action set, move from autonomous execution to agent assist, restrict the audience, or automate a deterministic step with Flow instead of using generative AI.

The production decision is an operating-model decision

Agentforce capabilities will continue to expand. Salesforce's September 28 recap alone included job-ready agents, long-running agents, and tools intended to analyze production sessions and improve agents. Each may shorten part of the technical path, but none removes the need to choose the job, prepare the right data, set authority, design recovery, and own the outcome. See the Dreamforce 2026 announcement recap .

The strongest program is not the one with the most pilots. It is the one that can explain why a particular workflow deserves production investment, show how it will be controlled, and decide from evidence whether to scale it.

Yuniq call to action: Bring Yuniq your top three Salesforce AI workflow candidates. Yuniq's Salesforce practice offers consulting, custom development, integration, implementation support, and managed support. Use the scorecard above as the agenda for a focused architecture and readiness discussion, then leave with a narrower production contract and an explicit remediation backlog. Discuss your Salesforce use case with Yuniq .

---