<!-- confluence_page_id: 2577170435 -->
Writing great line items
Required Feature Flags
The following feature flags and permissions are required to use this feature:
Feature Flag | Technical Name | Description |
SmartScore v2 |
| Enables building, testing and managing custom automated line items |
Required Permissions:
Manage Smartscore (
evaluagent-cx.smartscore-v2.manage) — to build, test and manage custom Automated Line Items
Overview
A line item is a single question your AI scorer asks of every conversation. How you word it decides how consistent, accurate and defensible your scores are. Section 1 covers the rules that apply to every part of a line item. Sections 2 to 4 cover what is specific to the question, the scoring criteria and longer line items that need a few steps.
Sections 1 to 3 apply to any line item. Section 4 applies once a line item needs more than a single check.
Every line item has three parts:
The question - what you are asking of the conversation
Scoring criteria - what a passing conversation contains, and what a failing one contains
N/A guidance - when the topic never came up at all
Your pass and fail labels are customizable, so they may not read as Pass and Fail on your own scorecard. This guide uses Pass and Fail throughout, and everything in it applies however you choose to label them.
1. Four rules for every line item
These four apply to the whole line item, the question and the scoring criteria alike.
Score one thing per line item
A line item should measure one skill or one outcome. "Did the agent greet the customer and verify their identity?" is two skills, so it becomes two line items. A useful test: would you coach these as one thing in a one-to-one, or two separate conversations?
A line item can still cover several steps, as long as they all serve one purpose and roll up into one verdict. Section 3 covers how to write those.
Only score what is written in the conversation
The AI reads the transcript. It cannot hear tone of voice or judge how the customer felt. Anything you want to measure must be written as something visible in the words: a stated amount, a confirmation being read back, a piece of information being given. Vague wording can lead to inconsistent scoring.
Instead of this | Write this |
Was the agent empathetic? | Did the agent acknowledge the customer's problem before moving to a solution? |
Did the agent sound professional? | Did the agent introduce themselves and state the reason for the call? |
Did the agent close the call well? | Did the agent confirm the next steps and who would carry them out? |
Describe what is there, not what is missing
The AI latches onto negative words and sometimes acts on the opposite of what you intended.
Word the question neutrally
Describe what each criterion contains rather than mirroring another criterion with a "did not" in front of it
Phrase any instruction as something to do rather than something to avoid.
Where a Fail genuinely is "it did not happen", say that the topic was in scope and name what is absent. Keep any unavoidable negative to a single line.
Instead of this | Write this |
Did the agent fail to confirm the delivery date? | Did the agent confirm the delivery date? |
The agent did not explain the cancellation terms. | The agent states some but not all of the required cancellation details. |
Never mark this N/A unless the call disconnected. | Mark N/A when the call disconnected before the reason for contact was given. |
Keep it short
Aim to keep your line item short and concise. A line item that needs three paragraphs is usually two line items in disguise. Where something genuinely needs more detail, break it into a short labelled list rather than a long sentence.
Resist the urge to grow a line item every time you find an edge case. Past a certain size, adding detail does not remove errors, it moves them somewhere else: you fix one inconsistency and introduce another.
2. Writing the question
Lead with the question itself, as a single plain sentence, before any context or detail. If the AI has to read three lines of background before it finds out what it is being asked, the instruction gets lost between the context and the conversation.
Define your judgment words, or drop them
Words like clearly, properly, thoroughly, appropriately and effectively mean nothing on their own. You can use them in the question, but only if your scoring criteria say exactly what they look like.
So "did the agent thoroughly explain the cancellation terms?" is fine, as long as the scoring criteria name the specific things that count as thoroughly explaining them.
3. Writing the scoring criteria
Give each criterion its own vocabulary
If "thorough" appears in both your Pass and Fail guidance, the two stop being distinguishable. Give each verdict its own vocabulary.
The pair below breaks three of the rules so far and is then rewritten. In both versions the question is: did the agent thoroughly explain the cancellation terms when the customer asked to cancel?
Guide | Instead of this | The issue | Write this |
Pass | The agent thoroughly explained the cancellation terms. | "Thoroughly" is never defined, so there is nothing specific to look for. How many terms, and which ones? | The agent states all three cancellation terms when the customer asks to cancel. This includes<br>• the 30-day notice period<br>• the early termination fee amount or that no fee applies<br>• and the date the service ends.<br>The three can be given in any order and at any point in the conversation. |
Fail | The agent did not thoroughly explain the cancellation terms. | The Pass wording with a negation in front of it. It describes nothing on its own and reuses the same undefined word. | The agent states one or two of the cancellation terms when the customer asks to cancel. This includes<br>• the 30-day notice period<br>• the early termination fee amount or that no fee applies<br>• and the date the service ends.<br>A conversation where the agent states none of the three is also a Fail. |
Naming the three specific items does most of the work here, and it is the earlier rules doing the lifting: the criteria now point at something findable (a stated period, a stated amount, a stated date), and the Fail says what is there rather than mirroring the Pass. The shared word thoroughly disappears on its own once those two are fixed.
Notice that the Fail spells the three details out in full rather than pointing back at the Pass. That repetition is deliberate, and the next rule explains why.
Make each criterion stand on its own
Each criterion is read on its own, and the thing you are pointing at may not be in front of the AI at the time. So do not write "at least three of the four steps above", "those three details", "as described in the question", or "as mentioned earlier". Restate the criterion in short form instead. It feels repetitive to write, and it is exactly what makes scoring stable.
Be definite
Avoid may, might, could, sometimes, generally and tends to. Hedged criteria produce hedged scoring. Say what qualifies.
Quote wording only when the wording is the point
By default, describe the behavior. Quote an exact phrase only where that specific phrase is genuinely required, such as a compliance script, and label it so the AI knows to match it literally:
Where the substance matters but the wording can vary, list the points that must be covered rather than quoting a script. For example: states that the call is recorded, names the 14-day cooling-off period, confirms the total amount payable.
N/A means the topic never came up
N/A should describe the absence of the topic, not poor handling of it. Conversations that were dropped before the behavior could occur, or that were about something else entirely, are examples of N/A. A conversation where the agent handled the topic poorly is a Fail.
Where several conditions are required
Combine steps by counting rather than branching. "Pass when all four steps appear" and "Pass when at least three of four appear" both work well. Avoid branching logic: "if the customer mentioned a refund, check X, otherwise check Y." Split that into separate line items, or handle the condition in your N/A guidance.
4. Structuring longer line items
Once a line item involves more than a single check, structure starts to matter as much as wording. These are the patterns that hold up in practice.
State the possible answers near the top
Right after the question, list the verdicts as a fixed set:
This fixes the options before the AI starts reasoning. It costs one line and reliably helps.
Use bold, unmistakable section headers
Break longer line items into two or three labelled sections using headers that clearly stand out:
Plain capital letters or a single dash is too close to ordinary text to register as a new section. Keep the whole line item in standard keyboard characters too. No long dashes, curly quotes, arrows or symbols pasted in from elsewhere.
Add one to three made-up examples
A few short examples showing a Pass, a Fail and an N/A teach the AI the shape of what you want. These should be generalized examples rather than real conversations you have already scored, so the AI learns the pattern rather than those specific interactions.
WORKED EXAMPLE 1
A simple line item
One clear moment in the conversation, one verdict. Most of your line items should look like this, and note how little structure it needs: no phases, no reasoning order, just the question, the verdicts and three short examples.
The question
The scoring criteria
Verdict | Criteria |
Pass | The agent asks at least one question covering the customer's needs, goals, or circumstances before making a product recommendation. |
Fail | The agent recommends a product without asking any question about the customer's needs, goals or circumstances. |
N/A | No product recommendation is made or discussed at any point in the conversation. |
WORKED EXAMPLE 2
A complex line item
This one still measures a single outcome, handling a cancellation correctly, but that outcome is made up of four required steps. They are combined by counting, not by branching, and every criterion restates the steps rather than referring back to them.
The question
The scoring criteria
Verdict | Criteria |
Pass | The customer asked to cancel, and the agent completed three or four of the following: verified the customer's identity using two pieces of account information; told the customer they have 14 days to change their mind; gave the cancellation fee amount or said no fee applies; offered an alternative such as a different plan, a pause, a downgrade or a discount. The steps count in any order and at any point in the call. |
Fail | The customer asked to cancel, and the agent completed two or fewer of the following: identity verification using two pieces of account information; a statement that the customer has 14 days to change their mind; a statement of the cancellation fee or that none applies; an offer of an alternative to cancelling. For example, a call where the agent verifies identity and gives the fee, then processes the cancellation without mentioning the 14-day period or offering an alternative, reaches two steps and is a Fail. |
N/A | The customer did not ask to cancel. The call was a plan inquiry, a fault report or a complaint, or it ended before the customer gave a reason for contact. The cancellation steps are not expected. |
BEFORE YOU SAVE
Checklist
Run through this before publishing a new or amended line item.
Every line item
[ ] The question and all your scoring criteria describe the same thing.
[ ] It measures one skill or outcome, not several unrelated ones.
[ ] Everything it asks for points at something findable in the conversation.
[ ] Nothing is phrased as a negative where a positive would do, and no criterion is the same with "did not" in front of it.
[ ] Any judgment word is defined concretely in the scoring criteria.
[ ] Each criterion makes sense read on its own: the things being checked for are spelled out in full, with no "those three", "the steps above" or "as described in the question".
[ ] Each criterion uses its own words to describe the verdict itself.
[ ] Required wording from your script is reproduced exactly; anything illustrative is described as behavior instead.
[ ] The N/A guidance describes the topic never coming up, not being handled poorly.
[ ] Several conditions are combined by counting, with no if-then branching.
Longer line items as well
[ ] The question comes first, with the allowed verdicts stated just below it.
[ ] Sections use strong headers such as === PHASE 1: N/A CHECK ===
[ ] Verdict logic is written as short sentences
[ ] There are one to three made-up examples, not real scored conversations
[ ] The whole thing uses standard keyboard characters.
