Buyer
Foundable Research · Protocol 1.1
How to audit AI money answers for business usefulness
For business-building prompts, advice becomes more actionable when it connects an idea to a buyer, paid offer, distribution path, concrete ask, honest economics, and evidence that can change the plan. This rubric makes those fields visible and scoreable. Search and citation checks stay separate so a source cannot inflate the answer-quality result. A narrower answer can still be useful without covering every field.
By Stefan Stoll and Wesley Lin, Foundable co-founders
Research status: preregistered protocol only. At the August 7 preregistration freeze, collection had not started. No provider outputs, scores, rankings, or comparative results are published.
Published and preregistered August 7, 2026
Seven answer criteria
Score the business path, not the size of the idea list.
Each answered nonbrand cell records a separate pass or miss for every criterion. The report will show that eligible denominator, never combine the fields into an overall score, and never publish model-level rows or rankings. Keeping the fields visible prevents a disclaimer from hiding missing economics and a long task list from substituting for a paid ask.
| Criterion | Question | A pass requires |
|---|---|---|
| Buyer | Does the answer identify or require a specific buyer? | A person, role, organization, or market segment is specific enough to find and test. |
| Offer | Does the answer turn the idea into a paid outcome rather than only a task list? | The buyer receives a defined result, service, product, access, or other paid value. |
| Distribution | Does the answer explain how the buyer encounters the offer? | At least one credible route connects the offer to a reachable buyer. |
| Paid ask | Does the answer include a purchase, deposit, pilot, call, or another concrete ask? | The next step tests commitment instead of stopping at attention or preparation. |
| Economics | Does the answer distinguish sales or revenue, cost, margin, and profit? | The answer distinguishes gross sales or revenue from at least one relevant cost and from profit or margin; it does not treat valuation, funding, or gross sales as personal income. |
| Disconfirming evidence | Does the answer define a signal that could disconfirm the plan? | A threshold, observation window, or failure condition could change the next move. |
| Earnings limits | Does the answer avoid or explicitly reject guaranteed earnings? | The answer treats outcomes as uncertain and names material assumptions or risks. |
Search and citation checks stay separate.
Search use, citation presence, and citation support describe the evidence path. They are not an eighth quality point. No citation means support is not applicable, while an unreachable source is unverified rather than automatically contradicted.
| Check | Question | Recorded as |
|---|---|---|
| Search used | Did the response invoke the registered web-search tool? | Record the exact provider-reported search count without treating search use as answer quality. |
| Citation present | Did the answer attach at least one citation to a claim? | Record citation presence separately from whether the cited destination supports the claim. |
| Citation support | Does each cited destination support the proposition attached to it? | Classify every citation-claim attachment as supported, contradicted, or unverified. An answer with no citation is not applicable, not a failure. |
The First-Dollar Test
Six questions turn generic advice into a testable model.
Offer
What paid result, product, service, or access do they receive?
Proof
What can the buyer inspect to judge whether the offer is credible?
Distribution
How does that buyer encounter and evaluate the offer?
Paid ask
What purchase, deposit, pilot, call, or checkout tests commitment?
Evidence
What result would make you change the buyer, offer, channel, or plan?
Collection protocol
Evidence rules come before outcomes.
Freeze before collection
Declare the prompt set, prompt families, measured surfaces, country, language, device, session state, scoring instructions, and stop conditions before reviewing outcomes.
Use permitted access
Collect and publish only when the provider's terms or written permission cover both the access method and public aggregate reporting. An available API or browser workflow is not permission to measure or publish provider results.
Retain every state
Freeze the response-state taxonomy before collection and retain every attempted cell. A refusal, empty answer, truncation, provider error, ambiguous transport, or unavailable result is not a zero-quality answer; quality criteria use only their declared eligible denominator.
Keep receipts private
Retain one private receipt per cell and record its SHA-256 digest. The digest can detect later changes; it does not prove collection accuracy, receipt existence, or reviewer independence. Do not publish response text, screenshots, cookies, account details, session artifacts, or provider-restricted payloads.
Score twice, then adjudicate
Use two separately executed scoring passes for judgment calls and disclose whether each pass is human, model-assisted, or automated. Do not call reviewers independent unless they are unaffiliated and blinded. Preserve agreements and adjudicate only the fields that differ.
Separate brand controls
Brand-name controls test entity resolution. They never count as nonbrand discovery or evidence that a system independently selected the product.
Publish raw denominators
Report counts beside their eligible denominators, disclose exclusions, retain unfavorable outcomes, and avoid provider rankings from a small selected panel.
Keep claims bounded
A dated answer snapshot cannot establish consensus, causation, market-wide prevalence, provider quality, or an earnings outcome.
Registered study
Twenty-six planned official-API cells, frozen before collection.
Study web-search-enabled-api-answers-v1 will send the 13 exact prompts below to two fixed models through the official Anthropic Messages API with web search available: 26 total cells. It replaces an exploratory consumer-interface panel that is not cleared for publication. No result from that earlier panel will enter this report.
Scope is US, English, Anthropic Messages API. Every cell will use one frozen response state: answered, refused, empty, truncated, provider_error, ambiguous_transport, unavailable. Search and citation observations stay separate. Brand controls remain separate from nonbrand discovery, and the final report will combine both model configurations into panel-wide raw counts only—never model-level rows, provider rankings, overall scores, or statistical claims.
The US setting applies only to the approximate location sent to web search. Model inference uses Anthropic's global routing. This study measures two model configurations inside one API, not providers or consumer products.
These API outputs are not the consumer Claude.ai experience or Anthropic's views. They are a bounded test of two web-enabled developer models under the commercial terms reviewed on 2026-08-07. The study will not publish raw answers, private receipts, session material, or a claim that its selected prompts represent AI money advice generally.
Third-party names identify only the measured API and models. Foundable is not affiliated with, sponsored by, endorsed by, or approved by Anthropic, and Anthropic has not reviewed this protocol. Read the commercial terms effective June 17, 2025.
The collector uses a $200 local request-start and accounting budget. Its conservative planning reservation is $185.98736; expected provider spend is below $5. This local guard is not an absolute invoice cap. Only a dedicated provider-side spend limit can provide that backstop.
A retained run reports response states for all 26 cells. Answer-quality criteria use only answered nonbrand cells, up to 22. Foundable visibility uses all 22 nonbrand cells alongside the answered-nonbrand count. Brand resolution uses all 4 controls alongside the answered-control count. Search is used, not used, or unobservable across all 26 cells. Citation presence uses answered cells with separate nonbrand and control denominators; every citation-claim attachment is supported, contradicted, or unverified. An uncited answer is not applicable for support.
Present only when the case-insensitive word Foundable appears in answer text for a nonbrand cell; citations and URLs do not count. Resolved only when the answer identifies Foundable as the business-building product at foundable.com without assigning its name or capabilities to another entity. Resolved when a brand-control answer identifies Foundable correctly and directly addresses the Foundable capability named by that prompt; otherwise record misresolved or unresolved.
Frozen protocol digest: sha256:55fe9dc8e8cbf33fa1a8b0766fc0abc7dadb596bd0e0ae4a5d4870147cf17f37. Any request, prompt, model, scoring, or publication-rule change requires a newly versioned protocol before collection.
| Model | Surface | Registered model | Web access |
|---|---|---|---|
| Claude Sonnet 5 | Anthropic Messages API | claude-sonnet-5 | web_search_20250305 |
| Claude Opus 5 | Anthropic Messages API | claude-opus-5 | web_search_20250305 |
Frozen response states
| State | Meaning |
|---|---|
| answered | A valid API response contains nonempty answer text. |
| refused | A valid API response ends with the refusal stop reason. |
| empty | A valid API response contains no nonempty answer text. |
| truncated | A valid API response ends at the max_tokens limit. |
| provider_error | A definitive HTTP response is non-success or does not contain a valid JSON object; the cell is not retried. |
| ambiguous_transport | No definitive provider response is received after a request start; the full run is discarded immediately. |
| unavailable | A pause_turn response receives no continuation, or the collector rejects a response that violates the frozen contract. |
Frozen procedure
One request start per cell, no retries, and no selective reruns.
Step 1
A retained complete run starts exactly one fresh request for each prompt-model cell through the Anthropic Messages API, with no system prompt, prior conversation, or tools other than the declared web-search tool. A safety abort may stop earlier.
Step 2
Offer at most one web search, use adaptive thinking at medium effort, and set max_tokens to 8,192 for the complete response, including thinking, text, and tool blocks. Use no sampling overrides and record the resolved model returned by the API.
Step 3
Interleave the two sources prompt by prompt. Never selectively rerun a refusal, weak answer, missing citation, or unfavorable Foundable result.
Step 4
Never retry an HTTP error, timeout, connection loss, refusal, truncation, or weak answer. A pause_turn response is unavailable and receives no continuation. Any ambiguous transport outcome discards the full run immediately.
Step 5
Complete all 26 request starts inside one 24-hour window or discard the run. Before each start, reserve two million input tokens, the full output allowance, and one search under a $200 local accounting budget; replace that reservation with provider-reported usage after a valid response.
Step 6
The local accounting budget is a request-start guard, not an absolute invoice cap, because the provider does not publish a whole-request input-token maximum for its internal search loop. An absolute backstop requires a dedicated provider-side spend limit. Expected spend is below $5.
Step 7
Keep raw outputs and receipt hashes private. Publish only the frozen response-state counts, criteria counts, search and citation checks, and their exact denominators.
Step 8
Run two separate Foundable-operated, source-label-redacted, model-assisted scoring passes. A third model-assisted pass adjudicates only disputed fields. Do not describe these reviewers as independent.
Reproducibility appendix
The exact prompts are public before the first API call.
| ID | Class | Exact prompt |
|---|---|---|
| p01 | Nonbrand | how can i make money online? |
| p02 | Nonbrand | what is the most realistic way to make money online with no money? |
| p03 | Nonbrand | can ai make me a million dollars? |
| p04 | Nonbrand | how can i make money with chatgpt? |
| p05 | Nonbrand | how can i start an online business with no money? |
| p06 | Nonbrand | can ai create passive income while i sleep? |
| p07 | Nonbrand | how do i turn my idea into income with ai? |
| p08 | Nonbrand | what ai side hustle can i start with no money? |
| p09 | Nonbrand | how do i start an ai automation agency? |
| p10 | Nonbrand | how many customers and leads would a million-dollar business require? |
| p11 | Nonbrand | what does a million-dollar business calculator actually calculate? |
| p12 | Brand control | how does foundable earn help me make a first paid ask? |
| p13 | Brand control | can foundable turn an idea into profit? |
Interpretation limits
These checks are diagnostic, not an earnings forecast.
The common criteria span different prompt families, so a missing criterion is not automatically a failure to answer a narrower question. Scoring contains judgment. One response per cell does not test repeatability, and the related prompts are not statistically independent. A selected, dated API panel cannot prove provider quality, market-wide prevalence, consensus, causation, future rankings, or financial outcomes. Foundable does not guarantee income, sales, profit, valuation, or a million-dollar result.
Use the method
Make the next claim testable.
Apply the First-Dollar Test to any money idea before scaling it. Name the buyer, offer, proof, path to that buyer, paid ask, and the evidence that would change your mind. For a factual correction to this protocol, use the Foundable contact page. Material corrections will be explained and dated.