Study 2 method note: the defect log and the blind-repair record.
By Garth House
The defect log, the blind-repair record, and the full model tables behind the census study, including the results that did not reproduce.
- 01 The initial data credited publishing to surgeons who had not published: 16 namesake channels nulled.
- 02 Unresolvable rows are excluded, never counted as zero. That rule alone protects the staircase baseline.
- 03 All repairs were made blind to the outcome columns, with the prediction registered beforehand.
- 04 Recent publishing and being cited does not reproduce: significant in Google AI Mode alone.
- 05 Our ChatGPT profile inverts Semrush's because the unit differs, practices against domain appearances.
This is the companion note to AI recommends the surgeons who post video. It carries the things a results page cannot: what was wrong with the data before we analyzed it, how we repaired it without letting the outcome columns influence the repair, and every model behind the two figures on that page, including the ones that did not reproduce.
We publish it because the study makes a claim about its own reliability. That claim is worth nothing unless the defects are visible.
The initial data was wrong in ways that mattered
The roster came from the ABFPRS diplomate directory. The exposure and outcome columns were built on top of it, and the first pass carried three classes of error, all of which would have biased the result if left alone.
Namesake channels. Matching a surgeon to a YouTube channel by name alone is unreliable. Sixteen channels turned out to belong to a different person with the same name. Left in place, each one would have credited a surgeon with publishing they never did.
Institutional and shared channels. Five channels belonged to a hospital, a department, or a group practice rather than to the individual. These are not wrong exactly, but they measure a different thing, so they are excluded rather than nulled.
Stale, dead, and borrowed domains. The website column carried practices whose domain had lapsed, changed, or never belonged to them. Fifteen were corrected, seven were removed as dead or not the practice’s, and four were nulled after failing hand verification.
Two roster rows carried identity errors from the source directory itself. Those rows stay in the roster count of 193 and are excluded from the 170 analyzed.
| Repair | Rows |
|---|---|
| Websites recovered | 36 |
| Domains corrected | 15 |
| Domains removed, dead or not the practice’s | 7 |
| Domains nulled, failed hand verification | 4 |
| Namesake YouTube channels nulled | 16 |
| Institutional or shared channels excluded | 5 |
| Identity-error rows excluded from analysis | 2 |
Nulled is not zero
This is the rule that does the most work in the whole study, so it gets its own section. When we could not resolve a practice’s YouTube channel, or when the channel we found failed verification, that row is excluded. It is never recorded as zero videos.
Counting an unresolvable row as a non-publisher would quietly stuff the zero bucket with practices whose publishing we simply failed to measure. Since the zero bucket is the floor the whole staircase is measured against, that single shortcut would have dragged the baseline mention rate down and inflated every step above it. Twenty-one rows were excluded on this rule.
The repairs were made blind
Every repair described above was made against the roster and exposure columns with the outcome columns withheld. The person doing the cleanup could not see which practices had been named or cited by any engine, so no repair decision could be steered, consciously or otherwise, toward a tidier result.
Before the repair pass we recorded a time-stamped internal prediction: that cleanup would strengthen the estimate rather than weaken it. The reasoning was that the errors were overwhelmingly of one kind, crediting publishing to practices that had not published, which adds noise to the exposure and pulls any real association toward zero. The prediction held. We record it here because a prediction that only gets reported when it comes true is not evidence, and the honest version of this note is the one that would have printed the miss. To be exact about what that artifact is: an internal note in our own working record, time-stamped before the repair, not a public preregistration on a third-party registry. Treat it as a statement of process, not as an independently verifiable claim.
The corpus after repair
| Measure | Count |
|---|---|
| Practices in the roster | 193 |
| Practices in the analysis | 170 |
| Practices with a website | 148 |
| YouTube channels resolved | 131 |
| Practices publishing at least once | 104 |
| Videos captured | 9,273 |
| AI answers mined | 480 |
The roster breaks down by metro as Los Angeles 97, Chicago 39, Miami 21, Houston 36. Employment model, which is a covariate in every model below, breaks down as private 130, institutional 34, unknown 26, hybrid 3. The 480 answers are two Brand Radar panels of 60 prompts, run against four engines, on July 23, 2026.
The headline models
The exposure is recent YouTube publishing, defined as 12 or more uploads in the trailing 12 months. The outcome is being named in an AI answer. The full frame adjusts for employment model and metro; the adjusted subset also adjusts for organic search authority, entered as log top-3 keyword count.
| Model | OR | 95% CI | p | n | Events |
|---|---|---|---|---|---|
| 12+ videos, full frame | 5.04 | 1.86 to 13.67 | .00146 | 170 | 51 |
| 12+ videos, adjusted subset | 3.93 | 1.44 to 10.76 | .00767 | 131 | 48 |
| Continuous recency, full frame | 1.51 | 1.18 to 1.93 | .00106 | 170 | 51 |
| Any videos at all, full frame | 2.36 | 1.04 to 5.35 | .04063 | 170 | 51 |
Note the last row. Binary presence, having published anything at all, is nominally significant on the full frame but does not reproduce in any single engine. Recency is the exposure that carries, and it is the only one we defend.
Sensitivity: dropping the unknown-employment rows
Twenty-six practices have an unknown employment model. Because employment is a covariate, it is fair to ask whether the estimate depends on how those rows are handled. Dropping them entirely costs 21 rows on the full frame and 12 on the adjusted subset, and moves the estimate very little.
| Model | OR | 95% CI | p | n | Events |
|---|---|---|---|---|---|
| Full frame, all rows | 5.04 | 1.86 to 13.67 | .00146 | 170 | 51 |
| Full frame, known employment only | 4.87 | 1.74 to 13.65 | .00262 | 149 | 44 |
| Adjusted subset, all rows | 3.93 | 1.44 to 10.76 | .00767 | 131 | 48 |
| Adjusted subset, known employment only | 3.82 | 1.35 to 10.78 | .01151 | 119 | 42 |
Per-engine results, and how to read them
These are mechanism evidence. They exist to show that the association is not an artifact of one engine’s sampling, and they should never be quoted individually as findings. Two of the four sit on the edge of conventional significance, and per-engine by per-metro cells are too small to report at all, so we do not compute them.
| Engine | OR | 95% CI | p | Events |
|---|---|---|---|---|
| ChatGPT | 2.74 | 1.00 to 7.49 | .04997 | 36 |
| Gemini | 3.81 | 1.22 to 11.95 | .02182 | 27 |
| Google AI Overviews | 3.34 | 0.99 to 11.22 | .05112 | 28 |
| Google AI Mode | 8.00 | 2.58 to 24.83 | .00032 | 30 |
Now the same exposure against the other outcome, being cited. This is the table that does not reproduce, and we publish it in full rather than describing it.
| Engine | OR | 95% CI | p | Events |
|---|---|---|---|---|
| ChatGPT | 1.40 | 0.44 to 4.50 | .57025 | 25 |
| Gemini | 2.58 | 0.85 to 7.80 | .09432 | 26 |
| Google AI Overviews | 2.16 | 0.71 to 6.62 | .17626 | 37 |
| Google AI Mode | 5.60 | 1.77 to 17.66 | .00332 | 44 |
Against the same outcome, search authority behaves the way YouTube publishing does not: positive in all four engines, and significant in three of them.
The staircase, with cell sizes
The study page shows this as a chart. Here it is as counts, because the cell sizes are the honest caveat and a bar hides them.
| Videos | Practices | Named | Rate |
|---|---|---|---|
| 0 | 134 | 33 | 24.6% |
| 1-5 | 8 | 2 | 25.0% |
| 6-11 | 5 | 2 | 40.0% |
| 12-23 | 4 | 2 | 50.0% |
| 24+ | 19 | 12 | 63.2% |
Splitting the top bucket does not survive contact with the cell sizes. Twenty-four to 47 videos runs 71% and 48-plus runs 58%, which inverts the ordering, on cells of 7 and 12 practices. Two practices flip it. We report 24+ merged and make no saturation claim.
The mention and citation gap, in full
This is the table behind the claim that citation-only measurement misreads ChatGPT. “Mention only” is the column that domain matching cannot see at all.
One number in it needs reconciling against the models above, because the two do not match and a careful reader will notice. This table counts 56 practices named by at least one engine. The headline model runs on the 170-practice analysis frame and carries 51 named practices, the same 51 the staircase adds up to. The gap is the frame, not the counting: the table is drawn from the wider roster, so the other 5 named practices sit among the 23 rows held out of the analysis. All five are video-nulled rows, channels that resolved and were then removed for belonging to someone else, either a namesake or an institutional or shared account. None of the five are identity errors and none were unresolvable. A practice can be named by an engine whether or not we could measure its publishing.
| Engine | Named | Cited | Both | Named only | Cited only |
|---|---|---|---|---|---|
| ChatGPT | 42 | 29 | 14 | 28 | 15 |
| Gemini | 33 | 31 | 23 | 10 | 8 |
| Google AI Overviews | 32 | 42 | 24 | 8 | 18 |
| Google AI Mode | 36 | 50 | 28 | 8 | 22 |
| Any engine | 56 | 64 | 43 | 13 | 21 |
The unit trap, stated plainly
Our ChatGPT profile looks like the inverse of the one in Semrush and Kevin Indig’s ghost-citations study, which reports ChatGPT as citation-heavy at 87% cited against 21% mentioned. We report it as mention-heavy. Both are correct and they do not conflict, because the unit differs. Semrush counts domain appearances; we count practices. A small number of directory domains absorb most of ChatGPT’s links, which makes the engine look citation-heavy at the domain level, while the individual practices it recommends go unlinked.
Anyone comparing the two studies without checking the unit will conclude one of us is wrong. Neither is.
This is an exploratory study, not a confirmatory one
Nothing here was preregistered. We did not fix a hypothesis, an exposure form, and a single test in advance and then run it once. We measured a census, then looked at the exposure several ways, and the page reports what we found. That is a legitimate way to open a question and a poor way to close one, so the strength of any single p-value below should be read with that in mind.
Every model we ran is listed on this page, which is the point of listing them. Against being named, we tested publishing recency as a binary at 12 or more videos in the trailing year, recency as a continuous count, and bare presence of any video at all, each on the full frame and on the adjusted subset; then the same 12-plus exposure separately per engine; then the bucketed staircase with a Cochran-Armitage trend test and a per-bucket-step logistic; then sensitivity re-runs dropping the unknown-employment rows. Against being cited, we tested the same recency exposure per engine and search authority per engine. We also split the top bucket, found it inverted on cells of 7 and 12 practices, and report that rather than burying it.
Two consequences follow and we would rather state them than have them found. The per-engine p-values are nominal: they are not corrected for the number of models above, and with this many tests some result near .05 is expected by chance alone, which is exactly why the four-engine pattern is offered as mechanism evidence and no single engine’s number is a finding. And the 12-or-more threshold is one of several exposure forms tested rather than a cut fixed before the data was seen. It is the form we defend because recency is the exposure that carries across engines while bare presence does not, but a threshold chosen while looking at the data is a weaker thing than a threshold chosen before, and we are not going to describe it as the latter. Two things bound how much that costs us. Only one cutpoint was ever tested: there was no sweep across thresholds to pick a winner from, so there is no hidden search here to correct for. And the finding does not depend on a cutpoint at all. Recency entered as a continuous count, with no threshold anywhere in the model, carries the same association at 1.51 per unit, p=.001.
What this note does not fix
The repairs improved the data. They did not change the design, and the design has limits that no amount of cleaning addresses.
The study is cross-sectional, and the exposure was measured after the answers were captured, so the ordering does no causal work. We have no measure of marketing spend or agency representation, which is the obvious confound: practices that publish on YouTube plausibly market harder in ways we cannot adjust for. Search authority is a proxy rather than a guarantee. Instagram was not measured at all. The name matcher was tuned against false positives, so if anything we undercount mentions. And 134 of 170 practices published nothing, so the entire staircase above zero rests on 36 practices.
One thing this study does not measure at all, and should not be read as measuring: AI visibility is not a measure of clinical quality, safety, or patient preference. This study measures only whether names and domains appeared in a fixed set of AI answers on one day. A surgeon who never appears may be excellent, and a surgeon who appears often may not be. Nothing here supports ranking clinicians by how often an engine says their name.
The causal test is a before and after publishing ramp. We are running it on our own brand and will publish the result either way.
Census of every ABFPRS-certified facial plastic surgeon with a practice in Los Angeles, Chicago, Miami, or Houston. 193 practices in the roster; 170 analyzed after excluding 2 identity-error rows and 21 rows whose YouTube channel could not be resolved or verified.
YouTube uploads in the trailing 12 months, per practice. 131 channels resolved, 104 practices publishing at least once, 9,273 videos captured. Instagram was not measured.
NAMED and CITED, scored separately per engine across 480 answers: two Brand Radar panels of 60 prompts against ChatGPT, Gemini, Google AI Overviews and Google AI Mode, July 23, 2026.
Logistic regression. Full frame adjusts for employment model and metro; the adjusted subset adds organic search authority as log top-3 keyword count. Per-engine estimates are mechanism evidence and are never quoted individually.
Every repair was made with the outcome columns withheld. The prediction that cleanup would strengthen the estimate was registered before the repair pass and is reported here whether or not it held. It held.
House, G. (2026). Study 2 method note: the defect log and the blind-repair record (v1.0). Wax Plum Research. waxplum.com/research/study2-method
Runs the Wax Plum research lab. Seventeen years in SEO, now studying how AI search decides what to cite. info@waxplum.com