Method

Read this before quoting a number

How this was measured.

Two records, built two different ways, with two different failure modes. Both are stated here so you can decide how much weight either deserves.

Feature record

Every feature cell was taken from the vendor's own documentation, pricing page, changelog, or repository, read on the date shown beside the cell. Third-party writing was used only where no first-party page states the fact, and is marked as such in the underlying data.

Where a vendor does not publicly state something, the cell reads unknown. We do not infer, and we do not fill a gap with a competitor's description. Each product page ends with a list of exactly what we could not source and why.

Vendors change pricing and features faster than any index updates, so the as-of date on each cell is the bound on how much you should trust it.

Sentiment

For each product we collected substantive public opinions from Hacker News, Reddit, GitHub issues and discussions, developer blogs, and YouTube: comments and posts where someone described actually using the thing, rather than headlines and announcements. Each opinion was classified positive, mixed, or negative. The score is the net balance of those classifications, mapped to a −100 to +100 scale, and the exact wording of how each product's score was produced is printed on its own page.

Quotes are verbatim excerpts, trimmed only for length, each linking to the comment itself so you can check the context we read it in.

People who post about a tool are not a random sample of people who use it. Forums skew toward the frustrated and toward early adopters, and they skew differently from each other. That is why every product prints how big its sample is, which platforms it came from, how it skews, and how confident we are.

The GitHub gap

The clearest pattern in this data is not about any product. Every product with a public issue tracker scores between 75 and 93 points lower on GitHub than on Hacker News. That gap is structural. An issue tracker only ever collects defects, so the channel has no positive register at all.

Cursor is the useful control. It runs its own forum and has issues disabled on its public repo, so it has no GitHub score and cannot take the penalty. A product looks better on this axis purely for not having a public bug queue.

We publish the split anyway, for two reasons. It shows which channel a headline number is really made of, and the GitHub sample is the best available read on what actually breaks in practice, as long as you treat it as a defect list.

Counting

One honest limit on all of this: the quotes are exact and machine-verified against their live sources, but the counts are our best read of how many distinct people were talking across the threads. These are not survey numbers.

Samples
ProductOpinionsQuotesConfidencePlatforms
Cursor 46 14 medium hackernews 30, reddit 9, youtube 5, blog 2
Claude Code 69 17 medium hackernews 38, reddit 9, youtube 4, blog 3, github 15
Codex 58 18 medium hackernews 27, reddit 14, youtube 6, blog 1, github 10
Buzz 65 20 low hackernews 46, blog 1, github 18
Hermes Agent 110 21 medium hackernews 38, reddit 62, blog 1, github 9
OpenClaw 167 24 high hackernews 92, reddit 66, blog 1, github 8
Not measured

This site does not benchmark anything. We never run these products against a task set. We do not measure code quality, speed, real token cost, or success rate, and we do not rank them. Anyone claiming a clean benchmark across six products of four different shapes is measuring their own setup.

Enterprise readiness, security posture, support quality and company viability are all outside what this measures. And we take no money from anyone listed: no affiliate links, no sponsorships.

Updates

The dataset lives as versioned JSON and is regenerated by a single script, so every refresh re-reads the vendor pages and re-samples the discussion from scratch. The generation date is printed on the index, and corrections are welcome.