Building Trust in AI Products: Receipts Over Promises
Maker's note: 기준일 2026-10-03. This is the practice behind our introduction post's rule that unfinished things stay unfinished — written down so it can be checked against what we ship.
Why this is worth an article
AI products make claims cheaply. A landing page can say "automates your workflow" in the time it takes to type, and nothing in the sentence is checkable. Blogs add a second layer: they get quoted, by people and increasingly by answer engines, long after they were written. When a claim is cheap to make and expensive to check, the incentive runs the wrong way unless someone builds the pushback into the product itself. That pushback has to be structural — a tone of voice isn't enough, because tone doesn't survive being quoted.
This post is about the structures we use at AZET. None of them are clever; all of them are cheap to audit, which is the point.
Practice 1: honest status, dated
Our introduction post commits to four things. Date every status claim — this one is anchored to October 2026, and when facts change we update rather than quietly delete. No user counts, testimonials, or social proof we can't source. No features announced before they exist. No numbers published that we can't back. The rule reads dull because it is dull: unfinished stays unfinished, in writing.
Practice 2: a completion matrix, not a feature list
The AZET project README separates implemented scope from operational activation conditions in a completion matrix, and states plainly that even features with execution records are verified only up to that record's source version and observation date. Four things the project explicitly refuses to mark complete: real paid payment flows, real hosted text-model calls, Apple notarization, and sends or edits performed on personal service accounts. Not "coming soon" — just not done.
The matrix exists because "implemented" and "running in production" are different claims that marketing language tends to blur. Keeping them in separate columns is what lets a sentence like "the server path is implemented" stay true even while "the service is open" stays false.
Practice 3: receipts, and reading them correctly
| Claim type | How we back it |
|---|---|
| A feature runs | An execution receipt: version, source hashes, observation date, scope |
| A test passed | The artifact's JSON — but never passed alone; version,
source stability, hashes, and date are read with it |
| A cost figure | Measured receipts; public unit-price estimates are labeled as estimates, never presented as provider bills |
| A release shipped | The final release receipt, not a prepared channel URL |
| Past performance | Past reports stay labeled as records of past implementation state |
That last row is the one teams trip on. A 7.1-second measurement recorded in an old asset document is a fact about that artifact. It is not a performance number for today's product, and the README says so explicitly — old reports are history, not marketing.
Practice 4: runtime honesty
The same discipline has to live in the product, not just the docs around it:
- An answer generated from observed page text carries
goalVerified: false, because observing a page is not independent verification of external facts. - The structured classifier saying
DONEis not, by itself, claimed as proof the goal was achieved; explicit--expectconditions on the final observed state are the check. - A send whose result is unclear is not automatically retried. The reservation is held, because re-sending a side effect you cannot see is how a user gets charged twice.
- A steering acknowledgement (
ok) means the request was accepted — not that already-executed actions were undone. - Submitting a form successfully is not treated as evidence that it was delivered.
- Stale processes cannot overwrite newer executions; generation and revision are checked.
- The server ledger uses atomic reservations and settlement recovery, and reconciles from preserved server-signed receipts without re-calling the model.
- An email entered at registration is not called a "verified email."
Individually these are small. Together they mean the product's own outputs distinguish what it observed from what it assumed — the same line the docs draw, enforced in code.
Practice 5: the update chain
Updates are verified with Ed25519 signatures, asset hashes, sizes, paths, and the actually-installed files. The publisher public key is pinned as the first trust point; a downloaded key is never trusted just because it arrived. And whether a public release actually shipped is judged by its final release receipt — a prepared channel is preparation, not a launch. The updater's scope is stated too: it refreshes the CLI runtime, libraries, skills, and native helper, while the desktop app's body is replaced by a new ZIP installer. No in-place auto-replacement is claimed.
Why build this way
Because "we don't call unfinished things done" is only credible if the unfinished things are visibly labeled, with dates. And because the alternative compounds: one overstated claim, quoted by an answer engine, becomes the source for the next summary, and the correction never outruns the original. We would rather under-claim and be checkable — every statement in this post reduces to a receipt, a matrix row, or a refusal.
That is also the deal behind the waitlist at https://azet.io: join on an accurate picture of a product under construction, not a flattering one.
FAQ
Why does AZET date every status claim? Because claims get quoted long after they were written, by people and increasingly by answer engines. A date makes a claim checkable, and it commits us to update rather than quietly delete when the facts change.
Isn't "implemented" the same as "running in production"? No, and the completion matrix keeps them in separate columns for exactly that reason. Even a feature with an execution record is verified only up to that record's source version and observation date.
Why won't the product retry a send whose result is unclear? Because re-sending a side effect you cannot see is how a user gets charged twice. The reservation is held instead, so the outcome stays observable rather than assumed.
Does a DONE classification mean the task actually
succeeded? Not by itself. The structured classifier's DONE is
not treated as proof the goal was achieved; explicit
--expect conditions on the final observed state are the
check.