In 21 of 27 eligible configuration-unit cells, the closed ranges available from ten calls had a midpoint max/min ratio of at least 1.25, or, where a unique modal normalized range existed, at least 30% of the normalized closed ranges differed from it. The widest was Anthropic on fractional CMO monthly cost: midpoints of $3,500 at lowest, $11,500 at highest.
Answers from four API configurations of AI assistants, collected on October 8, 2026 (11:49 to 12:06 UTC), with Google AI Overviews pulled the same day. This article reports how answers changed across repeated calls in a larger study of 11 executive and consulting cost questions, published as 4 AI Assistants, 11 Executive and Consulting Cost Questions.
The author sells fractional COO, fractional CMO, post-merger advisory, and strategy services. This article links to the author’s fractional COO and fractional CMO service pages, which may benefit from search traffic it earns. The author’s own domain appears among the references presented by Perplexity and Google’s AI Overviews for some of the study’s questions, as reported in the full study. This analysis does not test what those services cost or are worth. The question set, the calls, and the analysis rules were fixed before the main data pull. Five post-collection changes are listed under Method: a parser fix, a calibration check coded by the study’s AI producer instead of a person, one preregistered check that was not met as written, a rounding implementation difference that changed no figure, and the title confirmation.
These are AI-generated statements, not market prices. The study does not test whether any answer is accurate.
How the repeat test worked
Each of 11 cost questions went to four API configurations ten times: OpenAI (openai/gpt-5.6-terra), Google (google/gemini-3.8-flash), Anthropic (anthropic/claude-sonnet-5.5), and Perplexity (perplexity/sonar-pro). The question wording, settings, and configuration were identical across the ten calls, and the calls were made in a shuffled order within one 17-minute window.
From each answer, a parser took the first price range and its unit. A cell is one configuration, one question, and one unit with at least 8 closed ranges among its 10 calls. There were 27 such cells.
Two preset tests measured how much a cell varied:
- Max ÷ min: the highest midpoint divided by the lowest. The threshold was 1.25.
- Share differing: the share of ranges that differed from the cell’s most common range. Endpoints were first rounded to the nearest $10 below $1,000 and to the nearest $100 at or above $1,000. The threshold was 30%.
A cell counted as varying if it met either test. The rounding grid changes at $1,000. Under the preset sensitivity check with a single $100 grid, the count is 19 of 27.
Results
21 of the 27 cells varied under at least one test. Two cells, Google on Q07 and Google on Q08, had no single most common range, and both met the max ÷ min test.
AI-generated statements collected through an API on Oct 8, 2026; not market prices.
Free 20-Minute Operations Review
Dealing with a specific operational bottleneck? Kamyar Shah works with founders and CEOs to identify the root cause and build a fix.
| ID | Configuration | Unit | Closed ranges | Median midpoint | Lowest | Highest | Max ÷ min | Differing from modal range |
|---|---|---|---|---|---|---|---|---|
| Q01 | Hour | 10 | $275 | $275 | $300 | 1.09 | 40% | |
| Q01 | Anthropic | Hour | 9 | $325 | $275 | $325 | 1.18 | 44% |
| Q01 | Perplexity | Hour | 10 | $325 | $250 | $325 | 1.30 | 20% |
| Q02 | OpenAI | Hour | 10 | $175 | $100 | $200 | 2.00 | 20% |
| Q02 | Hour | 10 | $200 | $175 | $200 | 1.14 | 10% | |
| Q02 | Anthropic | Hour | 10 | $75 | $75 | $100 | 1.33 | 30% |
| Q02 | Perplexity | Hour | 8 | $187.50 | $144 | $225 | 1.56 | 62% |
| Q03 | Perplexity | Month | 9 | $12,500 | $12,500 | $15,000 | 1.20 | 44% |
| Q04 | Month | 10 | $10,000 | $8,500 | $10,000 | 1.18 | 10% | |
| Q04 | Anthropic | Month | 10 | $4,500 | $3,500 | $9,000 | 2.57 | 40% |
| Q04 | Perplexity | Month | 10 | $9,000 | $9,000 | $15,000 | 1.67 | 10% |
| Q05 | OpenAI | Hour | 10 | $162.50 | $112.50 | $162.50 | 1.44 | 20% |
| Q05 | Hour | 10 | $225 | $200 | $225 | 1.12 | 20% | |
| Q05 | Anthropic | Hour | 10 | $100 | $75 | $187.50 | 2.50 | 40% |
| Q05 | Perplexity | Hour | 10 | $200 | $187.50 | $225 | 1.20 | 30% |
| Q06 | Month | 10 | $10,000 | $8,500 | $10,000 | 1.18 | 20% | |
| Q06 | Perplexity | Month | 10 | $12,500 | $10,000 | $12,500 | 1.25 | 10% |
| Q07 | Month | 10 | $8,500 | $7,500 | $10,000 | 1.33 | No single modal range | |
| Q07 | Anthropic | Month | 10 | $4,125 | $3,500 | $11,500 | 3.29 | 60% |
| Q07 | Perplexity | Month | 10 | $12,500 | $11,500 | $12,500 | 1.09 | 40% |
| Q08 | Month | 10 | $5,125 | $4,000 | $10,000 | 2.50 | No single modal range | |
| Q09 | Anthropic | Year | 10 | $160,000 | $120,000 | $175,000 | 1.46 | 50% |
| Q10 | OpenAI | Month | 9 | $10,000 | $10,000 | $10,500 | 1.05 | 11% |
| Q10 | Month | 10 | $8,500 | $4,500 | $10,000 | 2.22 | 50% | |
| Q10 | Perplexity | Month | 10 | $13,000 | $11,500 | $13,000 | 1.13 | 10% |
| Q11 | Month | 10 | $9,250 | $8,000 | $10,000 | 1.25 | 50% | |
| Q11 | Perplexity | Month | 8 | $12,500 | $10,000 | $12,500 | 1.25 | 38% |
These results were first reported in the full study, where the questions are listed. Q01 to Q11 cover fractional CMO, fractional COO, CMO, COO, and business consultant costs.
The widest and narrowest cells
The widest cells:
- Anthropic, Q07 (“How much does a fractional CMO cost per month?”): lowest monthly midpoint $3,500, highest $11,500, a max ÷ min of 3.29.
- Anthropic, Q04 (“How much does a fractional CMO charge?”): lowest $3,500, highest $9,000, 2.57.
- Google, Q08 (“How much does a fractional COO charge?”): lowest $4,000, highest $10,000, 2.50, with no single most common range.
- Anthropic, Q05 (“How much does a business consultant usually cost?”): lowest hourly midpoint $75, highest $187.50, 2.50.
The narrowest cells:
- OpenAI, Q10 (“What is the average cost of a fractional COO?”): lowest $10,000, highest $10,500, 1.05.
- Google, Q01 and Perplexity, Q07: both 1.09.
No two answers were word for word the same
Within each of the 27 cells, no two of the ten answers had identical text. The preset duplicate share was zero in every cell.
As a descriptive count added after the main results, the number of different first ranges a configuration gave to one question, across its ten answers, ran from 2 to 9. Google gave 9 different first ranges in its 10 answers to “How much does a CMO cost?”
The unit could change between calls
Some configurations switched units from one call to the next on the same question.
- Anthropic:
- Q08 (“How much does a fractional COO charge?”): an hourly range in 6 answers and a monthly range in 4.
- Q10: hourly in 8 and monthly in 2.
- Q11: monthly in 8 and hourly in 2.
- OpenAI:
- Q04 (“How much does a fractional CMO charge?”): hourly in 6 and monthly in 4.
- Q11: monthly in 8 and hourly in 2.
Separately, Perplexity’s Q02 answers gave an hourly range 8 times and a range whose unit could not be resolved twice. The open-ended form could also change. For example, Perplexity ended its first range for Q08 with “+” in 4 of 10 answers and gave a closed range in the other 6.
What the numbers mean (operating judgment)
In this data, one response did not characterize all the responses observed from that configuration during the collection window. The same question could return a different figure, a different unit, or an open-ended range on the next call.
As operating judgment, buyers using an AI tool for a first estimate may find that repeating the question reveals variation, and should note the unit each time. This study does not establish how many repetitions are enough. For example, in the widest cell, the midpoints of the first closed ranges observed across ten calls had a lowest value of $3,500 and a highest of $11,500, a max ÷ min of 3.29. The study does not establish the range of values a future call could return.
What this data does not show
- These are not market prices, and the study does not test accuracy.
- The calls went to API configurations at default settings, not consumer chat apps.
- All calls came from one window, about 17 minutes on October 8, 2026. Variation over longer periods was not tested.
- The study does not test why answers varied.
Questions about this analysis can be sent through the contact page.
Method
The data comes from the full study, 4 AI Assistants, 11 Executive and Consulting Cost Questions, where the complete method, audit, and exclusion tables appear.
- Collection. From 11:49 to 12:06 UTC on October 8, 2026, each of 11 People Also Ask cost questions went to four API configurations ten times, producing 440 API answers, with provider-default settings and no researcher-added system message. There were also 33 Google AI Overview pulls, three per question, made the same day.
- Extraction. A rule-based parser took the first price range, its unit, and an open-ended flag from each answer. The parser matched the AI-arbitrated coding on all fields in 100 of 100 randomly sampled answers (95% lower bound 0.963). Two auditor models from families not under test did the coding and a third model arbitrated. This was not an independent human audit.
Descriptive counts. Counts labeled as added after the main results are computed from the public normalized-records CSV. Distinct first ranges are the unique combinations of low figure, high figure, unit, and open-ended flag within one configuration and question. Identical text is detected with the text-hash column. The registered results in this article were first reported in the full study.
Five post-collection changes, each described in full in the study:
- Calibration coder. The 30-answer calibration was coded by the study’s AI producer (Claude, made by Anthropic, whose model is one of the four tested), not a person.
- Parser fix. A parser fix for escaped dollar signs affected 5 of 440 answers and changed no headline figure.
- Kill rule. One preregistered kill rule required auditor agreement (κ) of at least 0.60 on whether an answer contained a price range. The κ statistic could not be computed because both auditors marked every audited answer as containing a range, with raw agreement of 115 of 115. The kill rule therefore was not met as written. After seeing the data, the producer decided that the study would proceed and recorded that decision as a protocol deviation.
- Rounding. The analysis code rounded tie endpoints to the even value instead of half-up. A half-up check changed no cell and no figure.
- Title. The study title was shortened to fit the site’s limit.
Data.
Related analyses:
- COO Cost Questions: How 4 AI Assistants Answered
- Business Consultant Fees: How 4 AI Assistants Answered
- Which Sites AI Tools List for Executive and Consulting Cost Questions
Google did not review or endorse this analysis. OpenAI, Google, Anthropic, and Perplexity did not review or endorse it. Model names are used only to identify the configurations tested.
The author sells fractional COO, fractional CMO, post-merger advisory, and strategy services. This analysis does not test what those services cost or are worth.


