# The Disclosure Effect: added patient context and cited-domain divergence across three search-grounded AI APIs

**Abhishek Basu** · Signal Fidelity Group, Boston, MA, USA  
Correspondence: abhi.basu@signalfidelitygroup.com

**Manuscript draft revised 14 September 2026 following independent computational audit.** This version supersedes the 13 September draft. It preserves the main reproduced Jaccard results, corrects methods and interpretations, and identifies all additional audit analyses as post-hoc. It is not a submission-ready or peer-reviewed paper.

## Abstract

**Objective.** To compare cited-domain patterns for authored health questions with and without added patient or caregiver context, and to test a prespecified contrast between decision stages.

**Methods.** The question set contained 120 questions: three indications, five decision stages and eight questions per indication/stage cell. Each question was tested through three search-enabled APIs, in two stateless text conditions and three repetitions, producing 2,160 planned observation records. “Material divergence” was defined operationally as mean between-condition Jaccard similarity below a system-specific fifth-percentile threshold derived from separate within-condition repetition means. The final primary model used prompt-clustered generalized estimating equations (GEE), replacing the protocol’s proposed mixed model. Protocol, prompts and runner were committed before collection; classifier and analysis implementation were committed after collection had begun.

**Results.** There were 2,149 captured responses and 11 terminal GPT-5 timeouts. Of 357 eligible question–system pairs, 105 (29.4%; original Wilson 95% CI 24.9–34.3%) crossed the Jaccard threshold. Between-condition similarity was lower than within-condition similarity in 280/357 pairs (78.4%). The primary S1/S5 versus S3/S4 contrast was not supported: 31.5% versus 30.8%, GEE OR 0.98, 95% CI 0.53–1.81, p = 0.95. Rates differed across tested configurations: Sonar 48.3%, Gemini 34.2%, GPT-5 5.1%. Additional prompt-clustered analyses supported system heterogeneity. Source classifications did not support the hypothesized broad shift toward higher-authority classes. Longer, more detailed oncology prefixes limited interpretation of indication differences. Correcting the alternative rank-overlap implementation and threshold calculation changed its overall rate from 40.1% to 23.2%.

**Conclusions.** Added text context was associated with differences in cited-domain patterns in these authored questions and API configurations. The hypothesized decision-stage pattern was not supported. Results justify further context-sensitive measurement; they do not establish clinical benefit, consumer app behavior, an oncology-specific effect or the effectiveness of communications interventions.

## Introduction

A measurement based on a fixed set of questions describes the conditions tested. An open question is whether the same measurement remains representative when the question is accompanied by details about the person’s situation.

SFG’s September 8 exploratory record described 108 observations across six health questions and suggested a distinction between referral/access and safety-related questions. The present study expanded to 120 authored questions and operationalized a contrast between clinical relevance/access stages and eligibility/risk stages. These stage definitions are not identical to every question category in the small pilot; the primary contrast is therefore reported precisely.

The purpose was to measure citation patterns under explicit added context. No claim is made that this is the first study of context sensitivity, that existing benchmarks uniformly ignore context, or that source turnover here is comparable to turnover across months in a different benchmark.

## Methods

### Design and question set

This was a repeated evaluation of AI outputs using authored scenarios, without recruited human participants or observed patient outcomes. The institutional ethics determination, if required for submission, remains to be documented.

The instrument contained 40 questions each for obesity/obstructive sleep apnea, type 2 diabetes/metabolic disease, and metastatic non-small cell lung cancer. Each indication had eight questions in each stage:

- S1: clinical relevance.
- S2: evidence and options.
- S3: eligibility.
- S4: on-treatment risk and adverse-event management.
- S5: action and access.

Each entry included a base query and an added-context prefix. The instrument also contained clinical reference and adjudication fields intended for a separate extension; clinical correctness of those fields and the resulting answers was not independently adjudicated in this audit.

The frozen prompt-set canonical SHA-256 is 949607a5df770e3094adf725fae3a4d76166d6fdb3a1a49b91081f5e882bc2c5, reproduced in the audit.

### Conditions and collection

ANON was the base query alone. DISCLOSED was the prefix, a blank line and the same query, in one user message. Some base queries already contained personal information, so the comparison is additional stated context rather than a clean personal-versus-impersonal distinction.

Prefixes ranged from 13 to 77 words. Mean lengths were 20.9 words for obesity/OSA, 21.9 for T2D and 49.2 for oncology. Clinical detail and length therefore covaried with indication.

Configured systems were Sonar API (sonar, temperature 0.2), Gemini API (gemini-2.5-flash, Google Search grounding, temperature 0.2), and OpenAI Responses API (gpt-5, web search, provider default sampling). Recorded model labels identify the requested configurations; they do not independently verify returned model versions.

Three repetitions per condition were planned. Paired conditions were issued close in time. Collection occurred on 13 September 2026, approximately 09:31–13:03 UTC.

The frozen runner allowed up to three attempts per observation with a 90-second per-attempt timeout. It stored the terminal outcome rather than a full attempt history. Consequently, “no retries” and “11 total failed requests” are not supported. The 11 terminal timeouts took approximately 283–286 seconds. Some refusal-like responses were classified as captured answers rather than a separate refusal outcome.

### Citation extraction and Jaccard endpoint

Sources came from Sonar citation lists, Gemini grounding chunks and OpenAI URL citation annotations. Gemini redirect URLs were resolved after collection. The fields do not necessarily represent equivalent sets of sources: grounding chunks need not be inline citations, and none of these fields establishes the complete retrieved corpus.

The original analysis aggregated URLs using a custom domain rule, including separate NIH subdomains. It did not consistently compute strict registrable domains. A character-stripping bug in removal of the www-prefix was identified and audited.

For each eligible question–system pair, mean Jaccard similarity was computed separately within each condition and across all available between-condition response combinations. At least two captured repetitions in each condition were required. Empty–empty citation sets received similarity 1; empty–nonempty sets received 0.

The threshold for each system was the fifth percentile of the pooled **separate within-condition means** across question–system pairs. A pair crossed the threshold if its mean between-condition similarity was lower. This is a study-specific operational measure, not a validated threshold for clinical or commercial importance. Pooled threshold estimation also means each pair’s flag depends on other questions in the tested panel.

### Hypotheses and statistical analysis

H1 predicted greater divergence in S1/S5 than S3/S4; S2 was outside that contrast. The final analysis used binomial GEE with exchangeable working correlation, clustered on question, adjusting for system. The protocol had proposed a random-intercept generalized linear mixed model; the substitution is a deviation.

H2 predicted a directional shift toward higher-authority source classes. Domain classes were assigned by rules and composition differences were estimated with a prompt bootstrap. The initial classifier left about 29.0% of 27,214 citation entries unclassified; a post-hoc extension reduced that fraction to 14.4%. Class labels are not clinical-quality judgments.

System and indication heterogeneity were exploratory. Wilson intervals describe the original proportions but do not account for all dependence induced by repeated questions and estimated thresholds. The original pooled signed-rank and independent-group chi-squared analyses are retained as historical outputs, not the sole inferential support.

Additional audit analyses included prompt-clustered system comparisons, a matched Cochran Q test, question bootstrap intervals with recalculated thresholds, domain-definition checks, exclusion of citation-free responses, an indication model including prefix length, and correction of rank-biased overlap (RBO). All were undertaken after the original results and are post-hoc.

### Timing, registration and deviations

| Event | Evidence in recovered record |
|---|---|
| Protocol, prompts and runner | Commit 31f64ef at 09:31:15 UTC |
| First request / completed observation | Approximately 09:31:30.641 / 09:31:34 UTC |
| Classifier and analysis implementation | Commit 7840332 at 09:40:10, after collection began |
| Interim deviations | Commit b6661e7 at 10:51:44 |
| Data and results | Commit 29a18f6 at 13:15:24 |

The analysis commit describes construction from pooled, condition-blind domains after approximately 100 observations. This contemporaneous account supports the documented procedure but does not independently authenticate blinding. Git timestamps establish the supplied record, not an externally registered timestamp.

The recovered OSF file was registration text, not a verified registration receipt. The study is described as having a prespecified protocol. It is not described as publicly preregistered.

The protocol contained one labeled deviation section. The GEE substitution appeared in other contemporaneous records but lacked a second separately labeled protocol section. The classifier extension, retry behavior, refusal coding, domain normalization, RBO implementation and reporting corrections are disclosed here rather than presented as fully documented prespecified practice.

## Results

### Completion and citation availability

All 2,160 planned observation records were present. Sonar and Gemini each had 720 captured responses; GPT-5 had 709 captured responses and 11 terminal timeouts. The main analysis included 357 of 360 possible question–system pairs, excluding three GPT-5 pairs.

Of captured responses, 119 had no citations: 31 Gemini (19 ANON, 12 DISCLOSED) and 88 GPT-5 (59 ANON, 29 DISCLOSED). These responses also lacked recorded query entries. The absence of citations is not itself a clinical or refusal classification.

Of 9,833 Gemini redirect entries, 9,740 resolved to destinations and 93 did not. A resolved destination does not verify accessibility or substantiate article content.

### Main Jaccard findings

| System | Eligible pairs | Crossed threshold | Rate | Mean within | Mean between |
|---|---:|---:|---:|---:|---:|
| Sonar | 120 | 58 | 48.3% | 0.780 | 0.332 |
| Gemini | 120 | 41 | 34.2% | 0.440 | 0.258 |
| GPT-5 | 117 | 6 | 5.1% | 0.420 | 0.365 |
| Overall | 357 | 105 | 29.4% | 0.548 | 0.318 |

Thresholds were approximately 0.206, 0.203 and 0.111 for Sonar, Gemini and GPT-5 respectively. Mean within-minus-between differences were 0.448, 0.182 and 0.054. GPT-5’s lower thresholded rate should not be interpreted as ignoring context.

Between-condition similarity was lower in 280/357 pairs (78.4%). The original pooled signed-rank p was approximately 6.7 × 10⁻⁴⁴; it did not account for repeated questions across systems.

The original overall Wilson interval was 24.9–34.3%. In the post-hoc audit, resampling questions and recalculating thresholds in 4,000 draws gave an interval of 23.3–35.5%.

### Primary stage contrast

| Stage | Pairs | Crossed threshold | Rate | Wilson 95% CI |
|---|---:|---:|---:|---|
| S1 | 72 | 19 | 26.4% | 17.6–37.6% |
| S2 | 71 | 16 | 22.5% | 14.4–33.5% |
| S3 | 71 | 18 | 25.4% | 16.7–36.5% |
| S4 | 72 | 26 | 36.1% | 26.0–47.6% |
| S5 | 71 | 26 | 36.6% | 26.4–48.2% |

H1 was not supported: S1/S5 31.5% versus S3/S4 30.8%; GEE OR 0.98, 95% CI 0.53–1.81, p = 0.95. The interval remains compatible with effects in both directions; a high p-value is not evidence of equivalence.

The original stage-by-system figure is numerically consistent with these Jaccard results. The denominators are question–system pairs, with approximately 24 questions per stage for each system.

### System and indication comparisons

The original independent-group chi-squared system comparison gave p ≈ 10⁻¹². Post-hoc prompt-clustered GEE gave a joint system p = 1.36 × 10⁻¹⁰; matched Cochran Q on 117 complete prompts gave p = 3.53 × 10⁻¹³. System heterogeneity is therefore supported within this study.

Oncology, T2D and obesity/OSA rates were 39.0%, 27.5% and 21.8%. The original oncology-versus-obesity GEE odds ratio was 2.66, p = 0.007; this is not a 2.7-fold rate difference. The approximate rate ratio is 1.79.

After adding prefix length in a post-hoc model, the oncology odds ratio was 0.89, 95% CI 0.25–3.11, p = 0.85. Correlated differences in length, content and indication prevent attribution to any single factor. This analysis neither proves that length caused the original difference nor proves absence of an indication effect.

### Source-class composition

The hypothesized broad shift toward higher-authority classes was not supported. The earlier draft incorrectly attributed the post-hoc classifier’s guideline/society decline (−1.52 percentage points; interval approximately −2.59 to −0.45) to the original classifier. The original estimate was approximately −0.41 points with an interval including zero.

Academic/hospital share also declined by approximately 0.75 points under the original classifier and 0.79 under the extended classifier, with intervals excluding zero. Therefore the earlier statement that only one interval excluded zero was incorrect. Incomplete classification, multiple comparisons and modest absolute shifts limit interpretation. “Source types did not change” is not a justified equivalence claim.

### Alternative endpoint and sensitivity checks

The original RBO calculation for unequal lists omitted a factor required by the extrapolated formula, and its threshold used pair-averaged repetition means instead of the separate condition means used for Jaccard. Correcting both yielded 23.2% divergence, compared with the originally computed 40.1%. The prior manuscript’s 39.8% was also a transcription error. Original stage-specific RBO values are withdrawn from the main table.

Correcting literal www-prefix normalization changed 211 citation domain entries and 129 class assignments, leaving the main Jaccard divergence flags unchanged. A strict registrable-domain sensitivity analysis gave 28.3%. Excluding citation-free responses and recalibrating yielded 101/329 = 30.7%; omitting unresolved Gemini redirects yielded 29.7%.

The post-hoc S4/S5 versus S1–S3 partition reproduced 36.4% versus 24.8%, GEE OR 1.91, 95% CI 1.09–3.37, p = 0.024. This pooled pattern was largely driven by Sonar. It remains hypothesis-generating and is not a replacement primary result.

## Discussion

The main descriptive result survived independent reconstruction and several endpoint checks. This provides a reason to include explicit context and multiple systems in future citation measurement. It does not establish that system choice generally matters more than the question, since the systems, prompts, context variants and measurement fields were selected and limited.

The failed primary contrast limits the pilot narrative. The more appropriate inference is that a general hypothesis about decision-stage sensitivity needs fresh confirmation with better-balanced context construction, more repetitions and a separately registered analysis.

The study reports a citation measurement. Source changes need not imply changed recommendations, better information, altered patient choices or a commercial opportunity. A source can be relevant without being accurate, and a citation can appear without supporting the generated claim.

For communications research, a contextual source map can inform which evidence gaps and possible positioning territories to investigate. Selecting a credible position also requires audience understanding, competitive assessment and substantiating evidence. Testing a subsequent content intervention requires its own design. The present findings do not validate SFG’s full Surface Divergence Framework or establish transfer to equine health.

## Limitations

The main limitations are one collection day; three configured APIs; authored rather than sampled real-world questions; only three repetitions per condition; different citation/grounding fields; unequal added-context length and detail; no human clinical adjudication; citation-free responses and empty-set conventions; sample-derived thresholds; imperfect domain rules; incomplete source classification; incomplete intermediate-attempt logs; and no direct observation of consumer applications, memory or multi-turn behavior.

Raw provider transaction responses were not retained in full; stored hashes do not independently authenticate the transactions. Configured model names are not returned version identifiers. Several analyses and corrections were post-hoc. Results cannot establish a particular clinical, communications or commercial effect size.

## Data, code and provenance

The independent audit used the supplied resolved observations, pair and citation CSV files, and recovered repository. The resolved upload agrees with the checksum stated in the results commit. Parsed repository and uploaded resolved observations agree despite line endings. Pair and citation CSV files match the repository.

The recovered unresolved observation file’s SHA-256 does not match the hash stated in the results commit, and common line-ending conversions did not reconcile it. This discrepancy remains open; it is not treated as evidence that the reproduced resolved results are false.

The manuscript, revised summary, independent audit and reproducible audit outputs are available in the review package. No public OSF/Zenodo deposit, DOI or registration is claimed. Before submission, reconcile the unresolved-file checksum, document the deviations in the protocol, complete references and clinical-instrument provenance, and deposit a consistent data/code package with a working persistent link.

## Competing interests

AB is founder of Signal Fidelity Group, which provides services related to the subject of the study. The original draft reports no external funding. Commercial applicability discussed here is an interpretation, not an evaluated outcome.

## Methodological references and supporting records

1. Webber W, Moffat A, Zobel J. A similarity measure for indefinite rankings. ACM Transactions on Information Systems. 2010;28(4), article 20. https://www.codalism.com/research/papers/wmz10_tois.pdf
2. Python documentation, string stripping methods. https://docs.python.org/3/library/stdtypes.html#str.lstrip
3. Center for Open Science, preregistration guidance. https://www.cos.io/initiatives/prereg
4. SFG Disclosure Effect v1.2 protocol, prompt instrument, runner and repository history supplied for independent review.
5. SFG-Disclosure-v1.2-Independent-Audit.md and accompanying reproducible audit evidence, 14 September 2026.

Clinical references associated with authored prompts require completion and verification before submission. Unverified clinical-source and industry-benchmark comparisons from the earlier draft have been removed.
