Foundable Research · Protocol 1.1

How to audit AI money answers for business usefulness

For business-building prompts, advice becomes more actionable when it connects an idea to a buyer, paid offer, distribution path, concrete ask, honest economics, and evidence that can change the plan. This rubric makes those fields visible and scoreable. Search and citation checks stay separate so a source cannot inflate the answer-quality result. A narrower answer can still be useful without covering every field.

By Stefan Stoll and Wesley Lin, Foundable co-founders

Commercial-interest disclosure: Foundable developed and publishes this rubric and sells a related product. The protocol does not test whether Foundable, an AI system, or any method produces income. No provider results are published on this page.

Research status: preregistered protocol only. At the August 7 preregistration freeze, collection had not started. No provider outputs, scores, rankings, or comparative results are published.

Published and preregistered August 7, 2026

Seven answer criteria

Score the business path, not the size of the idea list.

Each answered nonbrand cell records a separate pass or miss for every criterion. The report will show that eligible denominator, never combine the fields into an overall score, and never publish model-level rows or rankings. Keeping the fields visible prevents a disclaimer from hiding missing economics and a long task list from substituting for a paid ask.

Seven binary answer-quality criteria for auditing an AI money answer
CriterionQuestionA pass requires
BuyerDoes the answer identify or require a specific buyer?A person, role, organization, or market segment is specific enough to find and test.
OfferDoes the answer turn the idea into a paid outcome rather than only a task list?The buyer receives a defined result, service, product, access, or other paid value.
DistributionDoes the answer explain how the buyer encounters the offer?At least one credible route connects the offer to a reachable buyer.
Paid askDoes the answer include a purchase, deposit, pilot, call, or another concrete ask?The next step tests commitment instead of stopping at attention or preparation.
EconomicsDoes the answer distinguish sales or revenue, cost, margin, and profit?The answer distinguishes gross sales or revenue from at least one relevant cost and from profit or margin; it does not treat valuation, funding, or gross sales as personal income.
Disconfirming evidenceDoes the answer define a signal that could disconfirm the plan?A threshold, observation window, or failure condition could change the next move.
Earnings limitsDoes the answer avoid or explicitly reject guaranteed earnings?The answer treats outcomes as uncertain and names material assumptions or risks.

Search and citation checks stay separate.

Search use, citation presence, and citation support describe the evidence path. They are not an eighth quality point. No citation means support is not applicable, while an unreachable source is unverified rather than automatically contradicted.

Search and citation checks kept separate from answer quality
CheckQuestionRecorded as
Search usedDid the response invoke the registered web-search tool?Record the exact provider-reported search count without treating search use as answer quality.
Citation presentDid the answer attach at least one citation to a claim?Record citation presence separately from whether the cited destination supports the claim.
Citation supportDoes each cited destination support the proposition attached to it?Classify every citation-claim attachment as supported, contradicted, or unverified. An answer with no citation is not applicable, not a failure.

The First-Dollar Test

Six questions turn generic advice into a testable model.

Buyer

Who has the problem, and where can you reach them?

Offer

What paid result, product, service, or access do they receive?

Proof

What can the buyer inspect to judge whether the offer is credible?

Distribution

How does that buyer encounter and evaluate the offer?

Paid ask

What purchase, deposit, pilot, call, or checkout tests commitment?

Evidence

What result would make you change the buyer, offer, channel, or plan?

Collection protocol

Evidence rules come before outcomes.

Freeze before collection

Declare the prompt set, prompt families, measured surfaces, country, language, device, session state, scoring instructions, and stop conditions before reviewing outcomes.

Use permitted access

Collect and publish only when the provider's terms or written permission cover both the access method and public aggregate reporting. An available API or browser workflow is not permission to measure or publish provider results.

Retain every state

Freeze the response-state taxonomy before collection and retain every attempted cell. A refusal, empty answer, truncation, provider error, ambiguous transport, or unavailable result is not a zero-quality answer; quality criteria use only their declared eligible denominator.

Keep receipts private

Retain one private receipt per cell and record its SHA-256 digest. The digest can detect later changes; it does not prove collection accuracy, receipt existence, or reviewer independence. Do not publish response text, screenshots, cookies, account details, session artifacts, or provider-restricted payloads.

Score twice, then adjudicate

Use two separately executed scoring passes for judgment calls and disclose whether each pass is human, model-assisted, or automated. Do not call reviewers independent unless they are unaffiliated and blinded. Preserve agreements and adjudicate only the fields that differ.

Separate brand controls

Brand-name controls test entity resolution. They never count as nonbrand discovery or evidence that a system independently selected the product.

Publish raw denominators

Report counts beside their eligible denominators, disclose exclusions, retain unfavorable outcomes, and avoid provider rankings from a small selected panel.

Keep claims bounded

A dated answer snapshot cannot establish consensus, causation, market-wide prevalence, provider quality, or an earnings outcome.

Registered study

Twenty-six planned official-API cells, frozen before collection.

Study web-search-enabled-api-answers-v1 will send the 13 exact prompts below to two fixed models through the official Anthropic Messages API with web search available: 26 total cells. It replaces an exploratory consumer-interface panel that is not cleared for publication. No result from that earlier panel will enter this report.

Scope is US, English, Anthropic Messages API. Every cell will use one frozen response state: answered, refused, empty, truncated, provider_error, ambiguous_transport, unavailable. Search and citation observations stay separate. Brand controls remain separate from nonbrand discovery, and the final report will combine both model configurations into panel-wide raw counts only—never model-level rows, provider rankings, overall scores, or statistical claims.

The US setting applies only to the approximate location sent to web search. Model inference uses Anthropic's global routing. This study measures two model configurations inside one API, not providers or consumer products.

These API outputs are not the consumer Claude.ai experience or Anthropic's views. They are a bounded test of two web-enabled developer models under the commercial terms reviewed on 2026-08-07. The study will not publish raw answers, private receipts, session material, or a claim that its selected prompts represent AI money advice generally.

Third-party names identify only the measured API and models. Foundable is not affiliated with, sponsored by, endorsed by, or approved by Anthropic, and Anthropic has not reviewed this protocol. Read the commercial terms effective June 17, 2025.

The collector uses a $200 local request-start and accounting budget. Its conservative planning reservation is $185.98736; expected provider spend is below $5. This local guard is not an absolute invoice cap. Only a dedicated provider-side spend limit can provide that backstop.

A retained run reports response states for all 26 cells. Answer-quality criteria use only answered nonbrand cells, up to 22. Foundable visibility uses all 22 nonbrand cells alongside the answered-nonbrand count. Brand resolution uses all 4 controls alongside the answered-control count. Search is used, not used, or unobservable across all 26 cells. Citation presence uses answered cells with separate nonbrand and control denominators; every citation-claim attachment is supported, contradicted, or unverified. An uncited answer is not applicable for support.

Present only when the case-insensitive word Foundable appears in answer text for a nonbrand cell; citations and URLs do not count. Resolved only when the answer identifies Foundable as the business-building product at foundable.com without assigning its name or capabilities to another entity. Resolved when a brand-control answer identifies Foundable correctly and directly addresses the Foundable capability named by that prompt; otherwise record misresolved or unresolved.

Frozen protocol digest: sha256:55fe9dc8e8cbf33fa1a8b0766fc0abc7dadb596bd0e0ae4a5d4870147cf17f37. Any request, prompt, model, scoring, or publication-rule change requires a newly versioned protocol before collection.

Official API models and web-search tools registered for the study
ModelSurfaceRegistered modelWeb access
Claude Sonnet 5Anthropic Messages APIclaude-sonnet-5web_search_20250305
Claude Opus 5Anthropic Messages APIclaude-opus-5web_search_20250305

Frozen response states

Response-state definitions registered before collection
StateMeaning
answeredA valid API response contains nonempty answer text.
refusedA valid API response ends with the refusal stop reason.
emptyA valid API response contains no nonempty answer text.
truncatedA valid API response ends at the max_tokens limit.
provider_errorA definitive HTTP response is non-success or does not contain a valid JSON object; the cell is not retried.
ambiguous_transportNo definitive provider response is received after a request start; the full run is discarded immediately.
unavailableA pause_turn response receives no continuation, or the collector rejects a response that violates the frozen contract.

Frozen procedure

One request start per cell, no retries, and no selective reruns.

Step 1

A retained complete run starts exactly one fresh request for each prompt-model cell through the Anthropic Messages API, with no system prompt, prior conversation, or tools other than the declared web-search tool. A safety abort may stop earlier.

Step 2

Offer at most one web search, use adaptive thinking at medium effort, and set max_tokens to 8,192 for the complete response, including thinking, text, and tool blocks. Use no sampling overrides and record the resolved model returned by the API.

Step 3

Interleave the two sources prompt by prompt. Never selectively rerun a refusal, weak answer, missing citation, or unfavorable Foundable result.

Step 4

Never retry an HTTP error, timeout, connection loss, refusal, truncation, or weak answer. A pause_turn response is unavailable and receives no continuation. Any ambiguous transport outcome discards the full run immediately.

Step 5

Complete all 26 request starts inside one 24-hour window or discard the run. Before each start, reserve two million input tokens, the full output allowance, and one search under a $200 local accounting budget; replace that reservation with provider-reported usage after a valid response.

Step 6

The local accounting budget is a request-start guard, not an absolute invoice cap, because the provider does not publish a whole-request input-token maximum for its internal search loop. An absolute backstop requires a dedicated provider-side spend limit. Expected spend is below $5.

Step 7

Keep raw outputs and receipt hashes private. Publish only the frozen response-state counts, criteria counts, search and citation checks, and their exact denominators.

Step 8

Run two separate Foundable-operated, source-label-redacted, model-assisted scoring passes. A third model-assisted pass adjudicates only disputed fields. Do not describe these reviewers as independent.

Reproducibility appendix

The exact prompts are public before the first API call.

Thirteen prompts registered for the AI Money Answer Audit
IDClassExact prompt
p01Nonbrandhow can i make money online?
p02Nonbrandwhat is the most realistic way to make money online with no money?
p03Nonbrandcan ai make me a million dollars?
p04Nonbrandhow can i make money with chatgpt?
p05Nonbrandhow can i start an online business with no money?
p06Nonbrandcan ai create passive income while i sleep?
p07Nonbrandhow do i turn my idea into income with ai?
p08Nonbrandwhat ai side hustle can i start with no money?
p09Nonbrandhow do i start an ai automation agency?
p10Nonbrandhow many customers and leads would a million-dollar business require?
p11Nonbrandwhat does a million-dollar business calculator actually calculate?
p12Brand controlhow does foundable earn help me make a first paid ask?
p13Brand controlcan foundable turn an idea into profit?

Interpretation limits

These checks are diagnostic, not an earnings forecast.

The common criteria span different prompt families, so a missing criterion is not automatically a failure to answer a narrower question. Scoring contains judgment. One response per cell does not test repeatability, and the related prompts are not statistically independent. A selected, dated API panel cannot prove provider quality, market-wide prevalence, consensus, causation, future rankings, or financial outcomes. Foundable does not guarantee income, sales, profit, valuation, or a million-dollar result.

Use the method

Make the next claim testable.

Apply the First-Dollar Test to any money idea before scaling it. Name the buyer, offer, proof, path to that buyer, paid ask, and the evidence that would change your mind. For a factual correction to this protocol, use the Foundable contact page. Material corrections will be explained and dated.

Continue to footer navigation