Einstein Copilot entered public beta in February 2024 with a compelling native advantage: it could use Salesforce data, metadata and actions inside the flow of work. That advantage deserved a better test than user enthusiasm after a demo.

A credible pilot should compare assisted and unassisted work, expose failure modes and determine whether the system changes an operational metric without creating hidden review labor.

Cloud Group point of view

Pilot the workflow, not the feature. If the control group, baseline and decision rule are vague, positive anecdotes will overwhelm inconvenient evidence.

A practical playbook

The strongest next step is narrow enough to govern and useful enough to produce evidence. We would structure the work around these moves:

  1. Select one role, one workflow and one primary success metric.
  2. Create a representative scenario set before tuning prompts.
  3. Run real permissions and approved data from the beginning.
  4. Capture edits, rejections, escalations and time to completion.
  5. Define the production threshold and stop conditions in advance.

The architecture and operating implication

Keep the pilot close to native Salesforce controls. Use a bounded set of fields and actions, expose citations where possible and make the handoff path explicit. Record prompt and configuration versions so an improvement can be distinguished from a changed test.

Measure what changes

Model activity is not a business result. Track a small set of indicators that connect behavior to accountable work:

  • Task completion time against baseline
  • Output acceptance with no material edit
  • Unsupported-claim and escalation rates
  • User confidence after repeated use, not first use

The best pilot result may be a disciplined “not yet.” What matters is leaving the organization with reusable evidence, a clearer architecture and a more valuable next decision.

Primary sources

This field note is grounded in the product and market context available at the time of publication.