HIGH SALIENCE / RESEARCH / ACCOUNTING TEARDOWN

TEARDOWN 05 · PUBLISHED SEPTEMBER 17, 2026 · DATASET INCLUDED

100 Questions, Fifth Category: Accounting, Where the Incumbent Survives Every Vertical Prompt

Fifth category, same method: 100 accounting software queries, ChatGPT with web search on, every answer coded under the rulebook the first four teardowns used, Google's top ten as a control. In CRM and project management a vertical word replaced the field. In accounting the field has a floor, and the floor is QuickBooks: the vertical prompts put a specialist on top of it and leave it there.

01Method

Same query structure, same coding, one new category.

Queries: 100 accounting software queries built from the same templates as the earlier teardowns: 30 category, 30 comparison, 20 alternative and 20 recommendation. The full list is in the dataset.

Surface: ChatGPT with web search forced on, logged out, United States, English, via the DataForSEO scraper, one run per query, collected September 17, 2026.

Control: Google's top ten organic results for the same queries in the same window, via the DataForSEO SERP API at depth 20 and truncated to the first ten organic results.

Coding: every brand named was coded as recommended (selected for a stated case), listed (in a table or list without being chosen), passing, or anchor (the brand being replaced in an "alternatives" query), against a 55-brand dictionary fixed before collection, under codebook v1.5. Cited URLs were deduplicated to their domain and classed as first-party or third-party. Five answers drawn with a fixed seed were hand-checked: 25 coded mentions, 24 agreed with the reader. The one disagreement was a hedged pick ("particularly attractive if you prefer Xero's interface") the rulebook codes as listed; it is left in and reported.

Collection note: the ChatGPT answers for this category were collected through the DataForSEO connector's live scraper rather than the standard queue used for the earlier categories, one query at a time, same surface and settings; seven queries needed a second or third attempt after dropped connections and all 100 completed. One answer ("which accounting software is best value for money") came back written for a UK business, with pound pricing and FreeAgent, despite the United States location setting; it is coded as returned and noted in Limitations. Three aliases in the brand dictionary ("square", "stripe", "paypal" on their own) were narrowed after collection because they matched point-of-sale and payment mentions rather than accounting products; the dictionary shipped with the dataset is the narrowed one.

One run per query, so this is a teardown, not the benchmark. Frequencies describe this window only. Absence means not observed in this sample, never zero visibility.

02Findings

One incumbent, and it survives every vertical prompt.

Bar chart of 8 accounting software brands showing how many of 100 ChatGPT answers each appears in versus how many recommend it. QuickBooks 75 and 58, Xero 57 and 43, Zoho Books 46 and 33, FreshBooks 38 and 28, Wave 38 and 25, Sage 33 and 20, Sage Intacct 26 and 14, NetSuite 18 and 9.
BrandAppears inRecommended inRecommended share of appearances
QuickBooks755877%
Xero574375%
Zoho Books463372%
FreshBooks382874%
Wave382566%
Sage332061%
Sage Intacct261454%
NetSuite18950%

Out of 100 answers, codebook v1.5. "Recommended" means the answer selected the brand for a stated case. "Appears" adds brands listed in a table or bullet list without being chosen. The brand being replaced in an "alternatives" query is excluded from both columns.

QuickBooks leads, Xero follows, and both get chosen when named

QuickBooks is named in 75 answers and recommended in 58. Xero is named in 57 and recommended in 43, Zoho Books 46 and 33, FreshBooks 38 and 28, Wave 38 and 25. The recommendation share of appearances is above 70 percent for the top four, the same shape as CRM: a field of big brands each handed a condition. Below them the share drops. Sage is named in 33 answers and chosen in 20; Sage Intacct in 26 and 14; NetSuite in 18 and 9. The mid-market and enterprise products are named for completeness and chosen only when the buyer describes a mid-market or enterprise company.

When the buyer asks outright, it is a two-horse race. Nineteen of 20 recommendation-intent queries produced a dictionary pick, and QuickBooks was among the picks in 15, Xero in 14. No other brand was picked more than four times. The one query with no pick was the real estate brokerage, which was told to run Lone Wolf Back Office, Brokermint or Loft47 for commissions with QuickBooks underneath as the ledger.

The vertical prompts stack a specialist on the incumbent instead of replacing it

This is the finding that separates accounting from every category before it. Ask for a CRM for a law firm and the general CRMs vanish. Ask for accounting software for a law firm and the answer is "Clio plus QuickBooks Online", "LeanLaw plus QuickBooks Online", or CosmoLex if you want one system. Restaurants get "QuickBooks plus MarginEdge" or Restaurant365. Ecommerce stores get "A2X plus QuickBooks or Xero". A real estate brokerage gets "Brokermint plus QuickBooks Online". Small contractors get "QuickBooks Online plus construction add-ons" until they pass a revenue line the answer sets at roughly three to five million, after which the field does swap, to Sage 100 Contractor, Foundation and Sage Intacct Construction. Fifteen of the 100 answers contain an explicit "QuickBooks plus something" stack. Across the ten vertical and industry prompts we looked at, QuickBooks is recommended in eight and still named in the other two.

The mechanism is simple: the model treats the general ledger as infrastructure and the vertical tool as an application on it. A vertical vendor in this category is not competing with QuickBooks in the answer. It is being recommended next to it, and its pitch in the model's words is "we sync to QuickBooks".

Forks, with one internal winner

In 27 of 30 head-to-head queries ChatGPT recommended every brand named in the query. Xero versus Sage produced a single winner, Xero, with Sage described in the table but never selected for a case. QuickBooks Online versus QuickBooks Desktop produced a single winner, Online, because the answer reports that Desktop is no longer sold to new customers outside the Enterprise edition. Sage Intacct versus QuickBooks produced no pick from either: the answer described a size threshold and asked which side of it the buyer was on.

Half the recommendations go to brands that rank

Of the 268 recommended brand mentions, 136 were for brands whose own website does not rank in Google's top ten organic results for that query, 51 percent. That is the lowest overlap gap of the five categories: 65 in CRM, 70 in email marketing, 76 in help desk, 83 in project management. Google's top ten for these queries is 74 percent third-party pages, so the difference is not that Google favors vendors here. It is that intuit.com and xero.com hold their own category and comparison queries, and they are the brands the model picks.

The publishers are back

Across 100 answers there were 302 citation events to 106 domains, one per cited domain per answer. Forty-one percent went to vendors' own sites, led by intuit.com (36) and xero.com (26). But 27 percent went to review platforms and recognized publishers, the highest share of any category so far, and two of them dominate: NerdWallet with 25 citations and Fit Small Business with 17, more than zoho.com, waveapps.com or sage.com. An ERP comparison site, erpresearch.com, follows with 12. The ten most-cited domains cover 55 percent of citation events, the most concentrated census of the five, and 99 of the 302 events involved a domain that also sat in Google's top ten for the query, 33 percent, also the highest so far. Accounting is the category where the model's reading list looks most like a Google results page. Reddit and Wikipedia have zero citations, for the fifth teardown running.

035 categories, side by side

Same rulebook, every category so far.

Measure (codebook v1.5)Project management, Sept 8CRM, Sept 17Email marketing, Sept 17Help desk, Sept 17Accounting, Sept 17
Recommendations per category answer (average)4.83.33.33.93.1
Answers with no recommended dictionary brand5 of 10013 of 1007 of 1008 of 1009 of 100
Most-recommended brand: appears / recommendedAsana 76 / 68HubSpot 77 / 62Mailchimp 61 / 40Zendesk 74 / 54QuickBooks 75 / 58
Head-to-head queries recommending every named brand29 of 3026 of 3030 of 3028 of 3027 of 30
Recommendation-intent queries with a dictionary pick18 of 2016 of 2017 of 2018 of 2019 of 20
Recommended mentions where the brand does not rank in Google's top 10305 of 368 (83%)187 of 287 (65%)216 of 309 (70%)255 of 336 (76%)136 of 268 (51%)
Google top 10 that is third-party pages746 of 1,000 (75%)710 of 1,000 (71%)720 of 1,000 (72%)714 of 1,000 (71%)743 of 1,000 (74%)
Citation events to vendors' own sites216 of 362 (60%)136 of 320 (42%)193 of 367 (53%)152 of 345 (44%)125 of 302 (41%)
Answers citing only vendor pages50 of 10040 of 10037 of 10027 of 10034 of 100
Share of citations in the 10 most-cited domains54%40%42%45%55%
Citation events whose domain is in Google's top 1083 of 362 (23%)83 of 320 (26%)97 of 367 (26%)80 of 345 (23%)99 of 302 (33%)
Reddit and Wikipedia citations00000

All columns are coded under codebook v1.5; the project management column is the September 8 dataset recoded under it. Each teardown is one run per query in its own window, so differences between columns mix category with collection date.

04What this changes

Four things, in order.

If you are a vertical accounting tool, sell the stack. The model recommends you as a layer on QuickBooks, in those words. The page that wins is the one that says how you sit on the ledger, not the one that says you replace it.

If you are a general ledger, the size line is the fight. Every answer that moves off QuickBooks does it at a revenue or complexity threshold the model states out loud. Owning the sentence that defines that threshold is worth more than another feature page.

Two publishers carry this category. NerdWallet and Fit Small Business are cited more than most vendors. In the earlier categories the third-party half of the census was a long tail of unknown sites; here it is two names, and they are reachable.

Measure the ranking overlap per category. Fifty-one percent here, 83 in project management. The claim that AI recommendation and Google ranking are unrelated is true in some categories and false in this one.

05Limitations

What this teardown cannot tell you.

One run per query means answer variance is unmeasured; Benchmark 01 runs each query three times across three surfaces. The teardowns are collected on different days, so differences between categories mix the category with the date. The brand dictionary covers the general accounting software market and deliberately excludes vertical tools, so their appearances are described in prose and not counted. The rule-based coder has one known failure: a brand named as a contrast on the same line as a selecting verb is coded as recommended. Sentiment coding is rule-based and not reported. No vendor-tool cross-check was read for this category. Google's control counts a brand as ranking only when its own domain is in the top ten; a listicle that features the brand does not count. One answer was returned for a UK market despite the United States setting and is coded as returned; its picks (FreeAgent, QuickBooks, Xero) are in the data. Adjacent tools that appear in answers as payroll or expense add-ons (Gusto, Bill.com, Ramp, Brex) are in the dictionary because they were listed before collection; they are named in answers and never recommended as accounting software, and they do not affect the figures above.

06Dataset

Check it, don't believe it.

Every number above can be recomputed from these files. CC BY 4.0: use them, cite the page.

  • queries.csv: the 100 queries with intent labels.
  • brands.csv: the 55-brand dictionary with aliases and canonical domains.
  • mentions.csv: 434 coded brand mentions with position, type and whether the brand's domain was in Google's top ten.
  • citations.csv: 302 citation events with domain class and Google overlap.
  • observations.csv: one row per query with brand, recommendation, citation and control counts.
  • codebook.md: the rulebook, with dated amendments through v1.5.

The earlier teardowns: Teardown 01, project management, Teardown 02, crm, Teardown 03, email marketing, Teardown 04, help desk. This teardown is also being published on the High Salience Substack.

Five categories, and the vertical rule just broke.

Three categories where one word erases the field, one where it does not, and one where the field gets a specialist bolted on top. The only way to know which rule your category follows is to measure it. A Category Salience Brief runs this exact method on it, with a query set you approve first.