HIGH SALIENCE / RESEARCH / CRM TEARDOWN

TEARDOWN 02 · PUBLISHED SEPTEMBER 17, 2026 · DATASET INCLUDED

100 Questions, Second Category: In CRM, ChatGPT Picks Four Brands at Once

We ran the project management teardown again on CRM software: 100 queries, ChatGPT with web search on, every answer coded under the same rulebook, Google's top ten as a control. The model does not crown a CRM. It hands the buyer a fork with four tines, each with a condition attached, and the brand that ranks in Google is chosen more often here than it was in project management.

01Method

Same query structure, same coding, one new category.

Queries: 100 CRM software queries built from the same templates as the earlier teardowns: 30 category, 30 comparison, 20 alternative and 20 recommendation. The full list is in the dataset.

Surface: ChatGPT with web search forced on, logged out, United States, English, via the DataForSEO scraper, one run per query, collected September 17, 2026.

Control: Google's top ten organic results for the same queries in the same window, via the DataForSEO SERP API at depth 20 and truncated to the first ten organic results.

Coding: every brand named was coded as recommended (selected for a stated case), listed (in a table or list without being chosen), passing, or anchor (the brand being replaced in an "alternatives" query), against a 52-brand dictionary fixed before collection, under codebook v1.5. Cited URLs were deduplicated to their domain and classed as first-party or third-party. Five answers drawn with a fixed seed were hand-checked: 22 coded mentions, 21 agreed with the reader. The one disagreement is a brand named as a contrast on the same line as a selecting verb, left in and reported.

Codebook amendments v1.4 and v1.5: this teardown's hand-check found the rulebook missing hedged selecting verbs ("I'd lean HubSpot", "worth considering if"), matching brand names inside link addresses, and counting the brand being replaced in an "alternatives" query as a recommendation. The email marketing teardown's hand-check then found it missing conditional assignments ("small team: Pipedrive"). Both were fixed in dated amendments and every dataset was recoded. The figures on this page are v1.5. Under the original v1.1 cues HubSpot was recommended in 33 answers, under v1.4 in 39, under v1.5 in 62; the order of the top four brands is the same under all three.

One run per query, so this is a teardown, not the benchmark. Frequencies describe this window only. Absence means not observed in this sample, never zero visibility.

02Findings

The big four get chosen most of the time they are named.

Bar chart of 10 CRM software brands showing how many of 100 ChatGPT answers each appears in versus how many recommend it. HubSpot 77 and 62, Pipedrive 56 and 44, Zoho CRM 56 and 40, Salesforce 50 and 38, Freshsales 27 and 13, monday CRM 16 and 6, Microsoft Dynamics 365 15 and 11, Copper 13 and 6, Close 12 and 8, Attio 10 and 8.
BrandAppears inRecommended inRecommended share of appearances
HubSpot776281%
Pipedrive564479%
Zoho CRM564071%
Salesforce503876%
Freshsales271348%
monday CRM16638%
Microsoft Dynamics 365151173%
Copper13646%
Close12867%
Attio10880%

Out of 100 answers, codebook v1.5. "Recommended" means the answer selected the brand for a stated case. "Appears" adds brands listed in a table or bullet list without being chosen. The brand being replaced in an "alternatives" query is excluded from both columns.

Four picks, each with a condition

HubSpot is named in 77 answers and recommended in 62. Pipedrive, Zoho CRM and Salesforce are each recommended in roughly three quarters of the answers that name them. Thirty of the 100 answers recommend three or more of those four brands at once, and 61 recommend at least two. The typical CRM answer is not a verdict. It is a table, then a list of the form "want sales, marketing and service together: HubSpot; want a straightforward pipeline: Pipedrive; on a budget: Zoho CRM; expecting complex enterprise sales: Salesforce." Under the rulebook each of those lines is a recommendation for a stated case, and that is how the buyer reads them too.

Below the big four the pattern breaks. Freshsales is named in 27 answers and chosen in 13; monday CRM in 16 and chosen in 6. Those are the furniture brands in this category: present in the comparison table, rarely handed a condition.

Every Salesforce pick comes with the word enterprise

Salesforce is named in half the answers and recommended in 38. All 38 of those recommendations sit on a line that also says enterprise, complex, large, customization, scale or grow. The model does not recommend Salesforce; it recommends Salesforce once you are big. HubSpot's condition is just as stable: in 53 of its 62 recommendations, the line that names it pairs marketing with sales or service.

HubSpot is the default when the buyer asks for one

Sixteen of the 20 recommendation-intent queries produced a pick from the dictionary. HubSpot was among the picks in 15 of them, Pipedrive in 13, Salesforce in 12, Zoho CRM in 5. The four that produced no dictionary pick were a real estate brokerage, a law firm, an insurance agency and "best value for money", which answered with a price table and a request for team size.

Comparison queries fork here too

In 26 of 30 head-to-head queries ChatGPT recommended every brand named in the query. Project management, under the same rules, did it in 29 of 30. One three-way (Salesforce, HubSpot, Zoho CRM) dropped Salesforce and kept the other two. Three produced no pick from the query brands: Nutshell versus Pipedrive and SugarCRM versus Salesforce described the differences and asked for context, and Zendesk Sell versus Pipedrive reported that Zendesk is retiring Sell and reframed the comparison around what to migrate to.

Eight prompts swapped the field

Ask for a CRM for real estate agents and the answer names Follow Up Boss, Lofty, BoldTrail, Real Geeks, Wise Agent and Sierra Interactive. Not one of the 52 dictionary brands appears; the same happened for a real estate brokerage. Insurance agents get AgencyBloc, HawkSoft, EZLynx, AgencyZoom and Applied Epic, with HubSpot as the only general brand in the room. Financial advisors get Wealthbox, Redtail and Practifi. Nonprofits get Bloomerang, Neon CRM, Little Green Light and DonorPerfect. A law firm gets Clio Grow, Lawmatics, Filevine and Smokeball. In each case the general CRM market vanished on one word in the prompt, exactly as construction and legal did in project management.

In CRM, ranking and recommendation overlap more

Of the 287 recommended brand mentions, 187 were for brands whose own website does not rank in Google's top ten organic results for that query. That's 65 percent. In project management, under the same rules, it was 83 percent. Google's top ten for these CRM queries is still 71 percent listicles, review sites and comparison pages, but HubSpot, Zoho and Pipedrive rank their own category and comparison pages far more often than Asana or ClickUp do, and those are the brands the model chooses. That is consistent with the model reading what ranks; it does not prove it, and one category cannot settle it. Benchmark 01 is built to.

The citations are a long tail of sites you have never heard of

Across 100 answers there were 320 citation events to 138 domains, one per cited domain per answer. Forty-two percent went to vendors' own sites, against 60 percent in project management, and 40 of the 100 answers cited nothing but vendor pages. HubSpot is both the most-recommended brand and the most-cited domain, at 37 citations; in project management the most-cited site (ClickUp) was not the most-recommended brand (Asana). The rest is spread thin: the ten most-cited domains cover 40 percent of citation events, against 54 percent in project management, and the third-party half of the census is dominated not by G2 or Capterra but by sites such as crmnewspaper.com, itechguides.com, agiled.app and layer3labs.io, each cited four to six times. TechRadar is the most-cited independent publisher at eight. Reddit and Wikipedia have zero citations in this sample, as they did in the first teardown.

Eighty-three of the 320 citation events involved a domain that also sat in Google's top ten for that query, 26 percent. Three quarters of the evidence behind the answer is not in the results a rank tracker shows you.

032 categories, side by side

Same rulebook, every category so far.

Measure (codebook v1.5)Project management, Sept 8CRM, Sept 17
Recommendations per category answer (average)4.83.3
Answers with no recommended dictionary brand5 of 10013 of 100
Most-recommended brand: appears / recommendedAsana 76 / 68HubSpot 77 / 62
Head-to-head queries recommending every named brand29 of 3026 of 30
Recommendation-intent queries with a dictionary pick18 of 2016 of 20
Recommended mentions where the brand does not rank in Google's top 10305 of 368 (83%)187 of 287 (65%)
Google top 10 that is third-party pages746 of 1,000 (75%)710 of 1,000 (71%)
Citation events to vendors' own sites216 of 362 (60%)136 of 320 (42%)
Answers citing only vendor pages50 of 10040 of 100
Share of citations in the 10 most-cited domains54%40%
Citation events whose domain is in Google's top 1083 of 362 (23%)83 of 320 (26%)
Reddit and Wikipedia citations00

All columns are coded under codebook v1.5; the project management column is the September 8 dataset recoded under it. Each teardown is one run per query in its own window, so differences between columns mix category with collection date.

04What this changes

Four things, in order.

Own a condition, not a ranking. In this category the model assigns each big brand a case and repeats it. HubSpot's is "sales, marketing and service together"; Salesforce's is "once you are enterprise". If the model has not assigned you a condition, you are furniture, however often you appear.

Write the fork yourself. On comparison queries the model recommends both sides 26 times in 30. It is going to say when the other brand wins. Better it says it in your words.

Find the one-word prompts that erase you. Real estate, insurance, financial advisors, nonprofits, law. If those buyers matter, the vertical field is where you have to exist, and the general category pages will not put you there.

Treat the ranking-recommendation overlap as a per-category measurement. Sixty-five percent in one category and eighty-three in another is not a rule of thumb. It is two data points, and the third one is yours.

05Limitations

What this teardown cannot tell you.

One run per query means answer variance is unmeasured; Benchmark 01 runs each query three times across three surfaces. The teardowns are collected on different days, so differences between categories mix the category with the date. The brand dictionary covers the general CRM software market and deliberately excludes vertical tools, so their appearances are described in prose and not counted. The rule-based coder has one known failure: a brand named as a contrast on the same line as a selecting verb is coded as recommended. Sentiment coding is rule-based and not reported. No vendor-tool cross-check was read for this category. Google's control counts a brand as ranking only when its own domain is in the top ten; a listicle that features the brand does not count. The brand dictionary was written before collection and does not score the vertical tools named above.

06Dataset

Check it, don't believe it.

Every number above can be recomputed from these files. CC BY 4.0: use them, cite the page.

  • queries.csv: the 100 queries with intent labels.
  • brands.csv: the 52-brand dictionary with aliases and canonical domains.
  • mentions.csv: 438 coded brand mentions with position, type and whether the brand's domain was in Google's top ten.
  • citations.csv: 320 citation events with domain class and Google overlap.
  • observations.csv: one row per query with brand, recommendation, citation and control counts.
  • project-management-mentions-v1.5.csv: the September 8 project management mentions recoded under v1.5, for the side-by-side table.
  • codebook.md: the rulebook, with dated amendments through v1.5.

The earlier teardowns: Teardown 01, project management. This teardown is also being published on the High Salience Substack.

Two categories in, and they already disagree.

The third one should be yours. A Category Salience Brief runs this exact method on your category, against your competitors, with a query set you approve first.