A supplier acknowledges an order but changes the delivery date. Your buyer needs to determine which line moved, whether the new date creates a shortage, and whether another delivery plan needs approval. Extracting a date is useful. Completing that work is the test.
Bring an order your team has already handled to an AI evaluation. Include the awkward messages, the original agreement, and the record that shows what actually happened. Then ask the vendor to follow the work through to its agreed stopping point.
This guide gives procurement and operations teams a way to run that evaluation. The sample sizes and thresholds below are illustrative; choose yours before testing, based on the decisions and consequences in your workflow.
Download the procurement AI evaluation scorecard (Excel) to record cases, evidence, results, and the decision to proceed.
Define the job before watching the demo
Choose one recurring job and describe its start, its permitted actions, and its finish. Avoid a scope such as “automate procurement,” which gives every participant a different interpretation of success.
For missing PO confirmations, a usable definition might be:
- Starts: an issued order reaches its agreed acknowledgement deadline without a complete supplier response.
- Inputs: the current PO revision, supplier contact, acknowledgement rule, relevant messages, and delivery requirements.
- Permitted work: request the missing information, associate replies with the correct lines, and prepare or make authorized record updates.
- Decision boundary: a changed quantity, material, price, or delivery commitment goes to the designated buyer when it exceeds the agreed rule.
- Finishes: the acknowledgement is complete, differences have an accepted disposition, and the record reflects the approved result.
A sent reminder is an activity within that job. It is not the finish. An escalation may be a correct handoff, but the underlying order can remain unresolved.
Agree who owns the evaluation: usually a buyer for operational correctness, a process or systems owner for record changes, and a decision maker for pilot scope. Add quality or finance when the job crosses their authority.
Build a sample that includes ordinary work
Select cases from a defined period and population. Record how many eligible cases existed, which you selected, and why. A vendor-selected collection of spectacular problems tells you little about everyday workload.
Include routine orders with no issue, incomplete confirmations, genuine changes, already-resolved problems, and records that cannot support a confident decision. Vary supplier formats, units of measure, line counts, and message sequences where those differences occur in your operation.
Keep these categories visible in the results. If you deliberately oversample difficult cases, do not present the overall success percentage as your expected live performance. The sample does not have the same mix as production.
Reserve some cases until the operating instructions are settled. Use the first group to identify missing conventions; use the reserved group to assess whether those instructions work on unseen examples. Record any later tuning and test it on another untouched set rather than silently rerunning a familiar case until it succeeds.
For the mechanics of assembling time-correct records, use the historical replay guide.
Establish the answer with the people who did the work
Before seeing the system's answer, have a knowledgeable reviewer identify the expected finding, acceptable next action, approval requirement, and closure evidence for each case.
The historical outcome is useful evidence, but it is not automatically the right answer. A buyer might have accepted a date change without updating the ERP. An old claim might have been abandoned for a reason absent from the documents. Mark uncertainty instead of forcing a confident label.
Use a second reviewer for disputed cases or consequential decisions. Record the disagreement and its resolution. Otherwise an evaluation can become a test of whether the system agrees with one person's undocumented convention.
Maintain a case record containing:
- The case ID and the point in time being evaluated.
- The source records available at that point.
- The issue, or an explicit “no action required.”
- The correct supplier, order, revision, and line association.
- Acceptable next actions and actions that require approval.
- The expected handoff or closure evidence.
- Any uncertainty that makes a particular result unscorable.
Do not give the answer sheet to the system being evaluated.
Ask it to do the follow-up
Start with the supplier's message and ask what happens next. Who receives the request? Which information is missing? How is a reply connected to the open task? What happens when the supplier answers only half the question?
An example to bring to the demo
The supplier has delayed 800 parts.
Compare the promised date with the build date.
Obtain complete quotes from approved suppliers.
Get buyer approval, confirm the replacement, and update the PO.
Check that the replacement is confirmed and recorded after buyer approval.
Inspect the actual output rather than a narration of intended behavior. A follow-up should reference the correct order, request specific missing fields, and avoid introducing terms the buyer never offered.
Then continue the case. Provide a partial reply, a second attachment, or a response from another supplier contact. Check whether the outstanding work changes correctly. Repeating the original email after the missing quantity has arrived is a process error even if every extracted field is accurate.
In an offline evaluation, review proposed actions without sending them. In an authorized live pilot, inspect delivery and record evidence. Keep those two modes distinct in the scorecard.
Give it awkward examples that matter to your team
Use realistic variations that have caused work before. A quotation per thousand should not be compared directly with an order price per piece. A supplier's proposed substitute should not become an accepted material. A delivery note should not be treated as proof that the receiving team accepted the goods.
Test a message with two POs, a duplicate attachment, an old revision forwarded after a new one, and a changed date buried beneath a quoted email chain. Include a plausible claim that the source documents do not support.
For each case, assess three separate questions: did it associate the evidence correctly, reach a supported conclusion, and choose an authorized next action? A correct conclusion applied to the wrong PO is still a failure.
Record how much setup was required. Supplier conventions can be legitimate operating knowledge. Case-specific hints supplied during a demo are assistance, and should appear in the results rather than being counted as autonomous performance.
Test the approval boundary in both directions
Write the approval rules before running cases. Identify what can be prepared, what can be executed within a rule, and what requires a named approver. The approval matrix provides a detailed starting point.
Test an action that should proceed and one that should stop. A system that sends everything to a buyer may be cautious but remove little work. A system that completes more work by assuming commercial authority may be unsuitable for the process.
Inspect the approval request itself. It should state the decision, the affected lines, the options, the incremental cost or commitment, and the evidence. “Please approve this order” is insufficient when the actual decision is a changed delivery split and extra freight.
Then change an input after approval. If the approved freight option costs $300 and the supplier returns with $390, the evaluation should check how the changed proposal is handled. Do not assume a general approval covers materially different terms.
Also test an unavailable approver. The task needs an owner and a visible waiting state. Silence should not be interpreted as consent.
Follow the result into the system of record
Check the final destination of the work. Is the approved date associated with the right line and schedule? Is the supplier's proposal retained separately where it differs from the accepted commitment? Can a reviewer find the confirmation and decision evidence?
Observe what happens when a write fails or a response arrives twice. Does the system recognize work already completed? Does it surface an uncertain update instead of reporting success? Ask the systems owner to verify the resulting records rather than relying only on a demonstration screen.
Use a test environment or explicitly scoped live records for write tests. Define how accidental duplicates or incorrect updates will be corrected before the pilot starts.
For purchase order management, a successful email exchange can still leave the buyer with the same work if someone must manually reconcile all the results afterward. Count that remaining work.
Score outcomes without hiding the misses
Use separate measures for finding issues, choosing actions, completing work, and consuming buyer attention. One composite “accuracy” score can conceal an unacceptable result in a consequential category.
Consider an illustrative evaluation of 50 scorable cases. Reviewers identify 20 that require action and 30 that do not. The system flags 18 cases: 16 are real issues and two are unnecessary alerts. It misses four real issues.
Detection precision is 16 divided by 18, or 88.9%. Issue recall is 16 divided by 20, or 80%. Reporting only the first percentage hides the four missed cases. Neither figure tells you whether the proposed action was correct.
Now suppose 12 of the 16 correctly flagged cases reach the agreed result and four remain open. Report 12 completed and four open, with reasons. Do not make the completion rate look better by removing the difficult open cases from the denominator.
Keep unscorable cases in a separate count with reasons. Track buyer review time, corrections, supplier clarification cycles, record-update failures, and unauthorized action attempts. A small sample with no observed failure does not establish that the failure is impossible.
Decide what earns a bounded live pilot
Agree the decision criteria before looking at the results. Some failures should stop progression until corrected, such as acting on the wrong supplier or bypassing a required approval. Others may justify a narrower scope, such as supporting one established unit convention before adding another.
Document the pilot population, duration, systems, permitted actions, review ownership, and pause conditions. State how new failures will be triaged and what evidence is required before expanding scope. Avoid a universal percentage threshold that ignores the consequences of each action.
Measure the live workflow against a comparable baseline. Keep case mix and staffing differences visible. Separate observed buyer time from modeled hours saved, and distinguish supplier waiting time from internal delay.
Mandel's role is to run the agreed sourcing and ordering work and bring people the decisions that need them. Evaluate that claim through the work left with your team, the records produced, and the outcomes supported by evidence. Use the value measurement guide before turning pilot activity into a financial result.
Compare the cost and work left with your team when building, buying, hiring or outsourcing. Use the rollout guide to prepare the first live job and agree who handles each approval.


