Accuracy you can defend.
Every vendor claims their AI understands sales calls. Ask how they know and the answer is usually a demo. Ours is the number, how it was tested, and the evidence behind both.
98.2% accuracy — 214 of 218 checks across 50 deliberately difficult sales conversations, scored against answer keys written before the test was run; two were corrected on review, and the ledger names them. The analysis has changed since the run. When the test is re-run and the number moves, this page moves with it — including downward.
Zero invented figures and zero invented commitments remain after review. Of the four misses still counted, three said too little and one filed a real reversal under the wrong heading. The failure that would reach your forecast as agreed money is the one we do not have.
How it was tested.
- On the calls that actually go badly. Buyers who interrupt mid-sentence. Hedged answers that sound like commitments. A price said once, then revised ten minutes later. Danish, Swedish and Norwegian dropped into English mid-thought. Threads that start and never finish. And the "we’ll see" ending that every rep has mistaken for a yes. A test built on tidy calls tells you nothing, because your calls are not tidy.
- Answer keys written in advance. What each conversation actually established was written down before the test was run. Two keys were corrected on review, when re-reading showed the product was right and the sheet was wrong; the ledger names both.
- Against the product, not a raw model. Every conversation went through the same system your calls go through — same processing, same conservative judgements — as it stood on the day of the run. Testing a raw model instead would measure something we do not sell.
- Both directions counted. Missing something a call established costs a rep a few minutes. Stating something it did not establish travels into a forecast, into a board pack, out of the building. Those are not the same failure and they are not scored as though they were.
Where the checks land.
The 218 checks are four kinds of question, and they are not equally hard. The table is generated from the score file; a test fails the build if any cell drifts from the measurement.
| What was checked | Checks | Correct | |
|---|---|---|---|
| The deal value — right figure, or correctly none | 50 | 50 | 100% |
| The acceptance state — confirmed, likely, or none | 50 | 47 | 94% |
| Declined to state what the call did not support | 111 | 111 | 100% |
| Caught a buyer reversing themselves, on the right fact | 7 | 6 | 86% |
| All checks | 218 | 214 | 98.2% |
Who wrote the test.
We did, and that is worth being plain about. The conversations and the answer keys are ours; the keys are written before every run, two have been corrected on review, and the ledger names them. Every conversation is then read back by the founder — ten years selling B2B — and judged against one question: does a real salesperson behave this way? Conversations that fail that reading are removed from the test entirely, with their names and the reasoning on the record in our repository. Six have been.
Where the number really comes from
Most of this test is less "did it catch the fact" than "did it stay quiet when a careless reader would not have". 117 times it declined to state something the conversation did not support: a buyer’s off-hand guess at their own costs ("that’s a feeling, not a number"); a price revised down later in the same call; an unnarrowed 350–450k range; a procurement threshold that looked like a deal size; a competitor’s spend offered as a comparison; a rep’s own delivery promise; and every "we’ll see" ending. Across all 50 conversations it never once recorded a commitment the buyer had not made.
A whole messy call, and everything Talqo said about it.
The transcript below is one of the fifty — the hardest kind: a procurement conversation that circles, a budget thread interrupted by someone walking into the room and never picked up again, and an ending where both people say out loud that nothing has been agreed. On the right is Talqo’s output, unedited, including the fields it left empty.
- Deal value
- Not stated — no price was quoted on this call, so none was recorded.
- Price acceptance
- None — nothing to accept.
- Expected close date
- Not stated — the budget round "is normally October" and may have moved. That is not a date.
- Stage
- Needs Analysis
- Summary
- Rep spoke with Kjell Værnes about ~40 field service engineers who informally quote customers but whose on-site intelligence never reaches the business — a clear revenue leakage problem. No price was given at Kjell's request because procurement requires a formal requirement document before any number can be introduced. Rep agreed to send a template of requirement-document categories instead, and Kjell committed to checking the budget-round timeline with a colleague named Solveig.
- Next steps
- Rep sends Kjell a requirement document category framework (not vendor-specific). Kjell reads it and asks Solveig about the revised budget round date, then feeds that back to the rep.
- Qualification
-
budget Unknown
validated: falseNo budget figure discussed; Kjell stopped the rep before pricing and explained procurement process gates any spend. authority Medium Kjell owns the requirement document and is clearly the business sponsor, but procurement and Solveig are involved in budget approval. need High Kjell explicitly confirmed the core problem — field intelligence lost in transit — as 'exactly it' and 'that's the sentence'. timeline Low Budget round is normally October but may have moved; Kjell needs to check with Solveig and has made no firm commitment.
Why it prefers to be cautious.
A false "unproven" costs a rep two minutes: they look at the deal, see we did not score something, and either supply the fact or ignore it. A false "validated" costs a sales leader their credibility, because it travels — and the first anyone hears of it is when a defended number turns out not to have been said by anybody.
So every ambiguous case resolves downward, and that rule lives in the code rather than in a prompt: an acceptance we cannot quote back to you is not an acceptance; a price met with silence is not a price agreed; an answer we do not recognise resolves to the least claiming state, never the friendliest. Across all 50 conversations it never once recorded a commitment the buyer had not made — no invented figure, no agreement nobody gave. That is worth more than the percentage, because it is the failure that reaches your forecast.
Two states, precisely
Contradicted. Not a worse grade of agreement — off the scale entirely. The buyer accepted a price and then took it back, in their own words. Talqo keeps both quotes and shows the reversal beside the thing it reversed, because a fact that quietly disappears is worse than one never captured.
Not asked. The four words run strongest claim to weakest. “Not asked” means no price has been put to the buyer — it is not where Talqo puts a deal it is unsure about. A deal the buyer pushed back on reads not agreed, which is a different fact and ranks differently.
No score without a source.
A Talqo claim is the item, the reason, and the passage from the call with the cited sentence marked in place. This one is from the call above, so nobody in it is real; the analysis is.
Four words, and it earns each one.
A price is not “agreed” because a rep felt good about the call. Every deal’s price sits in one of these, and the sentence beside it is the one Talqo prints.
- agreed
The customer said yes on a call.
- not agreed
Talked about, no yes yet.
- contradicted
A yes, then taken back. The bar is deliberately high: an explicit reversal, a verbatim quote, and a named fact. Hesitation is not reversal.
- not asked
No price has been put to the buyer.
Absent, never invented. An item nobody discussed reads Missing, with the question to ask next. Talqo would rather show you a gap than fill it.
No score without a source. If we cannot show you the sentence, we do not say the buyer said it.
Run it on your own call.
Want this run against your own call instead? That is the offer on the front page — send one, and you get the same output, including everything it declines to state.
Full methodology and test data — the conversations, the answer keys and the scorer — are available for review on request.