# AI writing tools teardown: coding notes, codebook v2.0 (collected September 26, 2026)

The published codes are in coded_v2/ and come from AI readers, not people. Protocol: gold2/READER_PROTOCOL_v2.md (the six codes: picked, recommended, listed, passing, anchor, not_a_mention). "Recommended" in the page tables and in coded_v2/stats.json is the broad level: picked counts as recommended. The column "picked" in coded_v2/mentions.csv marks the answer's own verdict.

## Collection

All 100 ChatGPT tasks and 100 Google tasks were posted and fetched through the DataForSEO standard queue on the evening of September 26, 2026, US Mountain time (the raw files carry UTC timestamps, early September 27). Zero failed tasks, no retries. Raw responses are saved as returned.

## Coding method and agreement

- Every answer was coded by two independent AI readers (Claude agents; files gold2/out/A58-A60 and B58-B60 for this teardown), working blind from the published protocol on the answer text with link addresses removed. Where they disagreed, a third AI reader decided (gold2/out/J03 and J04; case files gold2/adj/21_ai-writing_*.md).
- Reader vs reader agreement: 97.3% on the six codes, 100.0% at the broad level (recommended or not). Codes decided by the third reader: 11, in 7 answers. (gold2/CODED_V2_SUMMARY.md)
- The rules-based coder used for the earlier draft of this teardown (coded/, codebook v1.5) agreed with the final codes on 62.8% of brand codes at the broad level; the revised v1.6 rules, which supplied the candidate brand list and positions for the readers, agreed on 91.1%. Neither is published.
- Positions, bold and table flags and all citations come from the v1.6 first pass; the reader codes decide each mention's type and which rows exist. Citations are identical to coded/citations.csv.

## Calls settled by the third reader

- q035 "chatgpt vs claude for writing": one reader picked both, the other coded both recommended. Settled: both picked, as head-to-head cases (ChatGPT as the versatile working partner, Claude for prose-first style, with a suggestion to test both).
- q041 "sudowrite vs novelai": one reader picked both, the other coded both recommended. Settled: both picked, on the answer's closing distinction (Sudowrite for a conventional novelist, NovelAI for a sandbox with models, lore and images).
- q058 and q060: same split (one reader picked both brands, the other coded both recommended), settled as picked both.
- q055 "chatgpt or claude for writing a novel": ChatGPT picked by both readers ("I'd lean ChatGPT overall"). Claude settled as picked on "I'd test both on your voice".
- q058 "claude or jasper for brand voice": both picked ("solo writer: Claude"; "marketing team: Jasper").
- q060 "jenni ai or paperpal for academic writing": both picked (Jenni for drafting a thesis from research, Paperpal for improving a draft).
- q082 "which AI writing assistant should I use for a 50 person company": Writer settled as recommended, a fit row in a table; the answer's own recommendation picks Grammarly and Microsoft Copilot.
- q097 "what AI writing tool is safest for confidential company documents": ChatGPT settled as recommended, not picked ("worth comparing" for an independent platform, hedged, not a verdict). Microsoft Copilot is the one pick.

## Judgment calls that matter for this category

- ChatGPT not offered as a writing tool is coded passing (9 answers): as an AI search surface whose visibility Writesonic, Surfer or Frase tracks (q032, q043, q044, q047, q050, q053); as a comparison anchor, "not a better ChatGPT" (q031); OpenAI models used inside Sudowrite (q041); only in an offer to compare tools later (q086). Gemini is passing on the same surface lines (q032, q043, q047, q050, q053), as a model inside Novelcrafter (q088, q098) and in a usage statistic (q008). Claude is passing as an Anthropic model inside Novelcrafter, Sudowrite or Copy.ai (q017, q041, q047) and in offers to compare later (q016, q081, q086). Writer is passing in q081 for the same reason (offer to compare later).
- ChatGPT is anchor in the two ChatGPT-alternative questions (q074, q079). Both pick Claude first: q074 "I'd try Claude first" (plus Perplexity for research); q079 "I'd narrow the choice to Claude vs. Writesonic vs. Jasper" (all three picked), with Copy.ai and Rytr picked from its practical shortlist.
- q083 "what AI writing tool do you recommend for marketing teams": Jasper picked ("my recommendation would be Jasper"); Writer recommended ("worth serious consideration for larger organizations").
- q037 "claude vs gemini for writing": Claude picked ("the one I'd test first"); Gemini recommended ("particularly attractive" for research-heavy writing), not picked. The only one-sided head-to-head.
- q047 (jasper vs copy.ai vs writesonic) and q048 (chatgpt vs claude vs gemini): each brand gets a "Consider X if" or "particularly good at" fit and the answer declines to name a winner, so no pick. q063 (writesonic alternatives) and q066 (sudowrite alternatives) are the other two answers without a pick.
- q091 "which AI tool should I use to write SEO blog posts": "What I'd use" is a two-step workflow (ChatGPT or Claude to draft, Surfer or Frase to optimize); all four picked, plus Jasper from the shortlist.
- q008 "most popular AI writing tool": ChatGPT picked as the direct answer; seven other tools listed in an "other major tools" list, no fit given.

## Brand mentions the readers added (outside alias matching)

Both readers flagged the same eight mentions (16 flags in all, gold2/CODED_V2_SUMMARY.md): Writer in q029, q061, q081, q082, q083, q085; Hemingway Editor in q071 (bare "Hemingway" as the app); Notion AI in q018. Final codes: Writer picked in q029 and q061, recommended in q082, q083, q085, passing in q081; Hemingway Editor recommended in q071; Notion AI recommended in q018 (already a first-pass row). Writer's final count is 7 named, 7 recommended, 4 picked.

## Dictionary notes that still apply

- brands.csv is identical to brands_original.csv: the dictionary is as fixed before collection. A case-sensitivity change to the ChatGPT aliases was tried after collection and reverted once the coder removed link addresses before matching (DICTIONARY_CHANGES.md; the tried version is kept as brands_casesensitive_chatgpt_v15.csv). The readers never saw link addresses, so the "utm_source=chatgpt.com" tag in every cited link cannot credit ChatGPT.
- The ChatGPT row also carries "OpenAI" and GPT model names.
- Writer (writer.com) matches only "writer.com" and "palmyra", because bare "writer" is an ordinary word; the readers added the mentions listed above. Hemingway Editor is not matched on bare "Hemingway" (fiction answers could name the novelist); the readers added q071.
- Case-sensitive aliases (capitalized only): Claude, Gemini, Copilot, Perplexity, Grok, Jasper, Surfer, Koala, Byword, Junia, Lex, Hypotenuse, Lavender, Jenni (14). Notion is case-sensitive in the shared coder.
- Folded rows, as drafted: Semrush ContentShake counts any Semrush mention; Notion AI counts bare "Notion"; Microsoft Copilot counts bare "Copilot"; HubSpot counts any HubSpot mention. Semrush ContentShake and HubSpot have no rows in coded_v2.
- The first-pass Writesonic matches inside third-party link slugs (q007, q075) are not in coded_v2: the readers coded text without addresses and neither answer names Writesonic.
- out_of_dictionary.csv is empty (the scraper returned no brand entities for this run).
