# Disclosure Effect v1.2 — audited results summary

**Revised 14 September 2026.** Replaces the 13 September summary. The original data and analysis outputs remain preserved; corrections and additional analyses below are explicitly identified.

## Design and completion

120 authored questions = three indications × five decision stages × eight questions. Each question was tested on three API configurations, in two text conditions, with three repetitions: **2,160 planned observation records**.

**2,149 captured responses; 11 terminal GPT-5 timeouts; 357 eligible question–system pairs.** Sonar and Gemini each captured 720 responses; GPT-5 captured 709. Three GPT-5 pairs had insufficient repetitions. These are repeated observations of authored questions, not 2,160 independent questions or patients.

Both conditions were stateless. “ANON” means the base query; “DISCLOSED” adds a fixed context prefix. Some base queries already contain personal information. Prefixes ranged from 13 to 77 words and sometimes contained multiple sentences.

## Main results

| Analysis | Audited result | Interpretation |
|---|---|---|
| H1, prespecified stage contrast: S1/S5 versus S3/S4 | 31.5% versus 30.8%; GEE OR 0.98, 95% CI 0.53–1.81, p = 0.95 | Not supported; not an equivalence finding |
| Main descriptive Jaccard endpoint | 105/357 = 29.4%; original Wilson 95% CI 24.9–34.3% | Reproduced; “material” denotes the operational threshold, not clinical importance |
| Threshold-free description | Between-condition similarity lower in 280/357 = 78.4%; mean within 0.548, between 0.318 | Reproduced; the original pooled Wilcoxon p ≈ 6.7 × 10⁻⁴⁴ did not account for cross-system prompt clustering |
| System differences | Sonar 58/120 = 48.3%; Gemini 41/120 = 34.2%; GPT-5 6/117 = 5.1% | Strong differences within these tested configurations |
| H2, source classes | No support for the hypothesized broad shift toward higher-authority classes | Does not establish unchanged composition or equal quality |
| Indication comparison, exploratory | Oncology 39.0%, T2D 27.5%, obesity/OSA 21.8% | Context length and detail differ by indication; no isolated oncology effect |
| Post-hoc S4/S5 versus S1–S3 | 36.4% versus 24.8%; OR 1.91, 95% CI 1.09–3.37, p = 0.024 | Hypothesis-generating; pattern largely driven by Sonar |

The Jaccard threshold was the system-specific 5th percentile of pooled **separate within-condition means**. Thresholds were approximately Sonar 0.206, Gemini 0.203, GPT-5 0.111.

## Stage counts

| Stage | Eligible pairs | Crossed Jaccard threshold | Rate | Wilson 95% CI |
|---|---:|---:|---:|---|
| S1 Clinical relevance | 72 | 19 | 26.4% | 17.6–37.6% |
| S2 Evidence and options | 71 | 16 | 22.5% | 14.4–33.5% |
| S3 Eligibility | 71 | 18 | 25.4% | 16.7–36.5% |
| S4 On-treatment risk / adverse events | 72 | 26 | 36.1% | 26.0–47.6% |
| S5 Action and access | 71 | 26 | 36.6% | 26.4–48.2% |

The original stage figure plots this Jaccard endpoint and remains numerically consistent. Its bars represent question–system pairs, with approximately 24 questions per system/stage, not independent patients.

## Additional audit analyses — all post-hoc

| Check | Result |
|---|---|
| Resample prompts and recalculate the Jaccard thresholds, 4,000 bootstrap draws | Overall 95% interval 23.3–35.5% |
| System comparison with prompt-clustered GEE | Joint p = 1.36 × 10⁻¹⁰ |
| Matched Cochran Q on 117 complete prompts | p = 3.53 × 10⁻¹³ |
| Mean within-minus-between Jaccard difference | Sonar 0.448; Gemini 0.182; GPT-5 0.054 |
| Correct literal www-prefix handling | Main binary Jaccard result unchanged at 29.4% |
| Strict registrable-domain sensitivity | 28.3% overall |
| Exclude responses with no citations; recalibrate | 101/329 = 30.7% |
| Omit unresolved Gemini redirects | 29.7% overall |
| Correct unequal-list RBO formula and align threshold calculation | 23.2%, replacing 40.1% from the original computation (39.8% was a manuscript transcription error) |
| Add prefix length to exploratory oncology model | Oncology versus obesity/OSA OR 0.89, 95% CI 0.25–3.11, p = 0.85; this does not isolate a causal length effect |

The mean added prefix was 49.2 words in oncology, 20.9 in obesity/OSA and 21.9 in T2D. The original OR 2.66 is an odds ratio; 39.0% versus 21.8% is approximately 1.79 times the rate, not 2.7 times.

## Corrections to methods and interpretation

The runner allowed up to three attempts per observation, with 90 seconds per attempt, and retained only the terminal outcome. “No retries” was incorrect. Intermediate failures and the exact request count cannot be reconstructed. Refusal-like answers were sometimes recorded as captured answers. There were 119 captured responses without citations: 31 Gemini and 88 GPT-5. Empty–empty citation sets received Jaccard 1 in the original code.

The domain normalizer used character stripping rather than removal of a literal www-prefix. Correction changed 211 citation domain entries and 129 class assignments, without changing the main divergence flags. The original taxonomy also used custom domain aggregation, including separate NIH subdomains; it was not a strict registrable-domain measure.

Source-class statements mixed classifier versions. In the original classifier, guideline/society share changed by about −0.41 percentage points. The −1.52-point estimate belonged to the post-hoc classifier. Academic/hospital share also had an interval excluding zero. The appropriate H2 conclusion is lack of support for the hypothesized broad directional shift.

## Timing and provenance

Protocol, prompts and runner: 09:31:15 UTC, commit 31f64ef. First completed record: 09:31:34; the first request began approximately 09:31:30.641. Classifier and analysis implementation: 09:40:10, commit 7840332, after collection began. The contemporaneous commit says rules were constructed from pooled, condition-blind domains. That is an author-recorded procedure, not independently verified blinding.

Deviations were recorded during collection; the GEE substitution is documented in the history/manuscript but was not a second separately labeled deviation in the recovered protocol. Results: commit 29a18f6. A public OSF registration and public data/code deposit have not been verified.

The resolved upload agrees with the recorded SHA-256, and its parsed records agree with the recovered repository despite line-ending differences. Supplied pair/citation tables match the repository. The unresolved observation archive’s checksum differs from the commit’s stated hash and remains to be reconciled.

## Claim boundary

These are domain patterns in citation and grounding fields, collected from Sonar, Gemini 2.5 Flash with Google Search, and GPT-5 with web search through developer APIs. The study did not evaluate clinical accuracy, actual patients, consumer app sessions, equine health, the effectiveness of communications interventions, or the full Surface Divergence Framework.

For SFG, this supports including realistic context and multiple systems in a diagnostic. Client-specific positioning and source decisions require client-specific evidence and judgment.

See MANUSCRIPT_DRAFT_v1.2.md and SFG-Disclosure-v1.2-Independent-Audit.md for the revised account and reproducible audit evidence.
