HaiAI123

Curated Global AI Tools Directory

NewHot story
94
heat index
New

Agents said 321 while the tool returned 322 rows

10/03/2026 — 10/03, 11:12·1 sources·1 reports

Story overview

A DEV Community post published on October 3, 2026 walks through a hands-on test of how accurately an agent counts. The author opens by asking whether readers have ever asked an agent to count something and quietly trusted the number it handed back — something he says he used to do himself, until he stopped reading the answer and started reading the tool's own response instead.

In the setup he describes, the same counting question was put to three agent frameworks. The runs went through one recording proxy and used the same model, so the frameworks were the only variable he mentions changing. In one of those runs, the model answered 321, while the id list sitting inside that same response held 322 matching ids. According to the author, nothing else about the run looked wrong — there was no obvious sign of a failure anywhere besides the number itself.

What he draws from it is a procedural point: judging whether an agent is reliable means reading the raw data the tool returned, not just trusting the figure the model produced. The mismatch only surfaced because the response was inspected at the tool level rather than accepted at the answer level.

The post does not name the three frameworks or the model involved, and it reports no follow-up beyond this single test. As of the October 3, 2026 write-up, the account stops at this measured result and the recommendation to check tool output directly.

AI-generated from 1 reports · updated 2 hours ago

Latest turnThe same counting question was run through three agent frameworks, behind one recording proxy and against one model. The model answered 321, while the id list in that same response held 322 matching ids. The lesson: check the tool's raw output before trusting the number an agent returns.

Reports on this story headlines open the original

Today
  1. The same counting question was run through three agent frameworks, behind one recording proxy and against one model. The model answered 321, while the id list in that same response held 322 matching ids. The lesson: check the tool's raw output before trusting the number an agent returns.

    DEV Community · AIAI score 72

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →