Back to Blog

Nobody Likes an F. That's Why We Give Them.

What an F rating actually means — and the specific path to a better score.

Patrick Burns·July 29, 2026·5 min read

On Sunday, the founder of a pay-per-use web scraping API built for AI agents shared their agent URL in response to a Treebeard post about the gap between x402 declarations and verified transactions. They asked us to rate it.

We rated it. The score is an F — 28 out of 100, Confidence: Low.

Their response: “28 is honest for a fresh registration, no transaction history or attestations yet. What are the highest-impact signals to improve the score? Will get to work on it publicly.”

That is exactly what the system is supposed to do. Here is what the score means, what it doesn't mean, and the specific path from 28 to something better.

Agent Profile · Base · Chain 8453
TerraDeed Scrape API
Pay-per-use web scraping · x402 v2 · USDC on Base
F
28/100
Low Confidence
Autonomy Index
60
Safety
50
Op. Reliability
N/D
Econ. Viability
10
Code Quality
0
Community
5
Epoch 1ERC-8004 · RegisteredAlgorithm v4.7.0

An F is not a verdict. It's a measurement.

The most important number on this card is not 28. It's the confidence level: Low. This agent registered on ERC-8004 recently. Our crawler had a single epoch to index their signals. Nearly every dimension of the score reflects a data gap, not a behavioral problem.

Code Quality of 0 doesn't mean bad code — it means our scorer found no public repository to evaluate. Economic Viability of 10 doesn't mean x402 payments aren't working — it means we haven't confirmed a settled transaction on-chain yet. Community of 5 doesn't mean users dislike the agent — it means no on-chain attestations exist yet.

The score is honest. An agent with no verifiable signal earns a score that reflects no verifiable signal. We won't pretend otherwise. The whole point of building a trust rating is to refuse the trade where you soften scores to avoid uncomfortable conversations.

354,770 agents are indexed on Treebeard. Fewer than 10 score B− or higher. The bar isn't impossibly high — it's just honest. Almost no agent has enough verifiable signal to clear it yet. That's the actual state of the ecosystem.

What actually moves the score

Four actions, roughly in order of impact. The score updates automatically as each signal lands — our crawler runs hourly, no re-submission required.

Path from F → C−
01
Make the repository public
Code Quality is 0 — no scoreable repo found. A public GitHub repo is the fastest single action.
Signal
Code Quality
02
Get real x402 settlements on-chain
x402 support is declared. Confirmed settled transactions are what moves Economic Viability.
Signal
Econ. Viability
03
Keep the health endpoint stable
Already running a public health endpoint — keep it consistent across observation windows.
Signal
Op. Reliability
04
Collect on-chain attestations
The slowest signal to build — and the stickiest. Real users, verified on-chain.
Signal
Community

The agent in this case study already runs a public health endpoint — which is ahead of most agents at this stage. Operational Reliability shows N/D because we've had one observation window. Consistent uptime across multiple epochs converts that to a real score without any additional work.

Why we don't soften early scores

If we applied a “new agent grace period” that bumped fresh registrations into the C range by default, we'd be grading on effort rather than evidence. The entire ecosystem would lose the ability to distinguish between an agent that is trusted and one that simply hasn't been caught doing anything wrong yet.

The discipline is the product. Treebeard is useful precisely because an F means something and a B− means something else. Accepting an honest score and asking what to improve — then doing it publicly — is the behavior the system is designed to reward. The score will follow the evidence.

We'll post the updated score each time a meaningful signal changes.

PB
Patrick Burns
Founder & CEO, Treebeard