Can AI models count?
Short answer: not reliably.
When Sorinai adds a line to your notes, it backs it up with evidence: a word-for-word quote from the call. Click it and you jump straight to that part of the transcript. Our servers check every quote against the transcript, so if you can see a quote, it's real.
But a real quote can still be attached to the wrong line, vouching for something it doesn't say. That's harder to catch, because it looks exactly as trustworthy as a correct one.
This week I caught Sorinai doing exactly that. The culprit? I'd asked an AI model to count from zero to four…
How Sorinai edits your notes
You can ask Sorinai to write or edit your notes directly: add a line you forgot, expand a point, fix a number. Behind the scenes, the model doesn't rewrite the whole page. It sends small edits, like "replace this block" or "insert after that one", and our server applies them.
If you've already written something, Sorinai merges its additions with yours: your points stay on top, and the details from the call go underneath as sub-bullets. Here's a block from a mock interview I test with:
- Senior Backend Engineer at Northwind Payments
- Six years experience total (4 at Northwind, 2 at a startup before)
- Engineering org ~80 people, his team is 6 engineers plus a PM and EM
- Senior IC title, but de facto tech lead on the ledger for ~2 years
To add a sub-bullet, the model rewrites the whole block with the new line in it, and every new line needs a quote from the transcript. So when a block gains several lines at once, each quote has to say which line it backs.
Attempt one: just ask for the line number
The obvious approach: a required field saying which line each quote backs, 0 indexed lines (counting from 0). On our server, that number was just an index into the block's lines:
// Find the line by the number the model gave
const target = lines[item.line ?? defaultLine];
// Attach the quote. If that line doesn't exist, it's dropped.
if (target) addEvidence(target, buildBlockEvidence([item], textByIndex));
Here's what came back:
- 0Senior Backend Engineer at Northwind Payments
- 1Six years experience total (4 at Northwind, 2 at a startup before)
- 2Engineering org ~80 people, his team is 6 engineers plus a PM and EM
- 3Senior IC title, but de facto tech lead on the ledger for ~2 years← quote went here
- 4Mentors two juniors; one was recently promoted to mid-level← it belonged here
The model counted the sub-bullets from 0 and forgot the point on top. A good old off-by-one error: programmers have been making it for as long as there have been programmers, and it turns out AI models make it too.
Was it a fluke?
To find out, I gave the two lightest models Sorinai uses the same system prompt, the same meeting and the same edit tools, with the same settings to analyse the worst case. Here's how they did:
- Claude Haiku 4.5: wrong line for 23 out of 43 quotes. More than half.
- OpenAI's GPT-5.6 Luna: wrong line for 6 out of 38.
And they weren't even wrong in a consistent way. On one point the model was a line short; on the next it was a line or two too far, past the end of the block, where our code quietly threw the quote away. So there was no simple "just add one to everything" fix.
Attempt two: explain it better
Next thing I tried was simply editing the prompt. I spelled out the counting rule in full: "count every line from 0, the point is 0, its first sub-bullet is 1, its second is 2, and so on".
To my surprise, it made no real difference. Re-running three of the requests, 16 quotes out of 40 still landed on the wrong line, against 17 out of 38 on those same three before. Apparently you can't explain your way out of an off-by-one.
Attempt three: copy, don't count
So I stopped asking for a number altogether. Now each quote names its line by copying it, word for word:
// What the model fills in for each quote: before (-), after (+)
line: {
- type: "integer",
- description: "Which line of your markdown this quote backs, counting from 0."
+ type: "string",
+ description: "The line of your markdown this quote backs, copied word for word."
}
On our server, finding the line is now a simple text match that ignores case and punctuation:
// Ignores case and punctuation, so a copy that kept its "- " still matches.
function looseText(text: string): string {
return text.toLowerCase().replace(/[^\p{L}\p{N}]+/gu, " ").trim();
}
// Tidy up every line in the block the same way...
const names = lines.map((line) => looseText(collectText(line)));
// ...then pick the one the model copied
const target = lines[names.indexOf(looseText(item.line))];
Same requests, same models:
- Claude Haiku 4.5: right line for 49 out of 49 quotes.
- GPT-5.6 Luna: right line for 40 out of 40.
Across both models: 29 of 81 quotes on the wrong line with numbers, 0 of 89 with copies.
The only cost is a few extra words: each quote the model wants to use as evidence now carries the whole line of notes it's supposed to back, a dozen or two tokens where the number was one.
One more catch
With copying in place, running requests through the real edit code turned up one more problem. Now and then, Haiku would add a new line with its quote, then hang a second quote on a line it hadn't touched, usually the point right above its edit. It did this in 10 of 58 runs.
The line it named was real, but it was the wrong one to cite. So I added one more rule: while an edit is adding lines, a quote aimed at a line it didn't change gets ignored.
After both changes, I ran the edits another 64 times through the real edit code. Not one quote landed on the wrong line.
Why copying works
Looking at what I was asking the model to do, nothing in front of notes says which line is line 4, which means the model has to count the lines it's writing as it's editing, inside a tool call where the whole block is one long string and every line break is a \n. And it has to decide where the count starts. That's exactly where it slipped: 26 of the 29 misses in the first round were off by exactly one.
Copying is the opposite. The line it need to reference is the one it wrote just a moment ago, sitting right there in its context window. That's about the easiest thing you can ask of an AI model.
So numbers aren't the problem. Counting things are. Every block the model sees is tagged [b0], [b1] and so on, and every transcript segment is numbered, so numbers are perfectly fine when it's used as an id and its labelled clearly, because just means treating it as text and copying what's printed right beside it.
Conclusion
If you're building with AI and need a model to point at something, have it copy the thing and let your code work out where it is. Things like line numbers, list positions, character offsets, "the third bullet" should be avoided, unless the data is labelled explicitly with that number.
Coding tools figured this out first. Anthropic's text editor tool is literally called str_replace_based_edit_tool: its main way of editing finds the text to change by exact match. And Aider, an open-source coding assistant, tells models to leave line numbers out of the code edits it asks for. Its 2023 write-up puts it bluntly: "GPT is terrible at working with source code line numbers." Seems like that is still true today on cheaper models.
So, can AI models count? Still not reliably. Luckily, they're very good at copying.
The feature is live: ask Sorinai to fill in something you missed, then click the quote to see exactly which moment in your conversation it came from.