Guides / ChatGPT vs the record

Why pay for this when ChatGPT is right there?

5 min readCites 1 record

Fig. 1. One question, put to a chat window and to the record.

In brief

Because the chat window and the record are different instruments. A chatbot returns the average of what was said, smooth and undated. A record returns the fracture: 56 quotes, every one linked to its source, the denial mechanism first, a confidence label, a trend, and a signature. On 9 July 2026 we put the same question to ChatGPT and to the pipeline, and published both answers.

This is the objection we hear most, so it gets the first guide and a fair fight. On 9 July 2026 I asked one question twice: once to ChatGPT, in a logged-out chat with nothing configured, and once to the Nolemy pipeline. Same question, same day. Both answers are below, and both are reproduced exactly.

The experiment#

The question was “Does Uber Eats refund missing items?”, a claim thousands of people live out every week. ChatGPT got it exactly as typed: one message on chatgpt.com, default model, no account. The pipeline got the same claim and did what it always does: searched nine community platforms, killed the marketing, kept 56 admissible quotes, and compressed them into a signed record.

ChatGPTchatgpt.com · 9 Jul 2026

“Yes. Uber Eats often refunds missing items, but it depends on the circumstances and the outcome of their review.”

“Refund eligibility can depend on factors such as the order details, the evidence available, and your account history.”

“If you tell me: your country, what item was missing, and whether the restaurant bag was sealed, I can help you understand what outcome is most likely.”

Dates cited: 0 · Sources cited: 0 · Vanishes when the tab closes

The record56 reports · 9 platforms

“Uber Eats denies refunds for missing items after multiple claims, contradicting its promise to refund incorrect or missing orders.”

Refunds routinely $2 to $3 short of the missing item. Aggressive language reaches a human; polite requests get the automated rejection.

All 56 quotes linked, 49 dated · Strong evidence · Deteriorating · Signed, fingerprint be14 0a84

What each artifact knew#

Read the two columns again, slowly. The ChatGPT answer is competent. It is also unfalsifiable: zero dates, zero sources, zero named platforms, and a confidence expressed as maybes (“often”, “may”, “can depend”). You can act on it. You can never check it. The record is the opposite kind of object. Here are three of its 56 quotes, exactly as people typed them:

“Uber denied my refund even though the driver confirmed the items were missing.”

“Customer Support is saying that Uber cannot issue a refund "because we have submitted missing item reports in the past." This is beyond frustrating!”

“My Uber eats orders were constantly missing items and they eventually stopped issuing me refunds when I kept submitting requests.”

Those three quotes span June 2026 back to 2021. Same mechanism every time: the refund system works until you have used it, then an account flag turns every future claim into not eligible, evidence be damned. That pattern, repeated across 56 reports on nine platforms, is what the record leads with. It is the single most useful fact about this claim, and the chat window knew it too. Which brings us to the interesting part.

The buried wall#

To be fair to ChatGPT, the flag mechanism is in its answer. It sits mid-list, as one factor among several: “the order details, the evidence available, and your account history.” Four words, no explanation, equal billing with paperwork. The 56 people in the record describe “your account history” as the whole game: the wall that ends refunds for anyone who has used the system more than a few times.

That is the fracture this whole guide exists to show. A chatbot weighs text by how often the internet repeated it, and the internet repeats the policy far more often than the outcome. A record weighs text by testimony density: 56 people, sorted with the pain pattern first, because the pain pattern is what you were actually asking about. Same fact, opposite placement: one buried it under a reassurance, the other led with it, because 56 people put it first.

From the public record

“Uber Eats refunds missing items”

56 reports · 9 platforms · strong evidence · deteriorating · signed be14 0a84 dcc7 b842

The published numbers#

My experiment is one claim. The published tests point the same direction. In April 2026, WIRED asked ChatGPT to cite WIRED’s own product picks in three categories: TVs, headphones, laptops. All three answers were wrong, including headphones WIRED had never tested. OpenAI’s own announcement of its shopping research feature, November 2025, scored the purpose-built tool at 52% product accuracy on multi-constraint queries, against 37% for ordinary ChatGPT search. Those are the vendor’s numbers, about the vendor’s flagship.

The academic result is harsher. ByteDance’s ShoppingComp benchmark (November 2025, 558 scenarios, 35 domain experts) put GPT-5.2’s product-retrieval F1 at 17.76%. Human experts scored 30.02%. One scope note before you quote any of this: these numbers measure shopping tasks, and the Columbia Tow Center’s March 2025 finding that AI search tools botched more than 60% of citation queries is about news sourcing, a different task. Each number counts only for what it tested.

The objection#

Where the chat window wins#

Now the limits, because this entry owes you its own. For settled knowledge, a chat window is the better tool: syntax, recipes, the boiling point of things. Use it there; we do. The instrument changes the moment a question has a counterparty, someone who profits from your belief in a particular answer. Refund windows. Cancellation flows. “Rated Excellent.” On those questions the average of the internet is the marketing budget talking, and you need the fracture instead: the measured gap between the promise and the 56 people who lived it.

One of those 56 might be wrong. All of them together, dated and linked and agreeing on the mechanism, are a fact about the world, and no model produces that fact from memory. Ours is rented, same as everyone’s. The 56 receipts are what you pay for.

ChatGPT capture: chatgpt.com, logged out, default model, 9 July 2026, verbatim. Record: 56 reports across 9 platforms, published and author-signed the same day. External tests: WIRED (Apr 2026), OpenAI shopping research announcement (Nov 2025), ShoppingComp (arXiv 2511.22978), Tow Center (Mar 2025). Every number above links to its source in place.