Skip to content
Signal Fidelity GroupBack to SFG

Answer Fidelity · Technical study record

The Disclosure Effect v1.2: study design, results and limitations

Signal Fidelity Group · Data collected September 13, 2026
Public technical summary prepared September 21, 2026 from the revised September 14 study records

This is a technical summary of a descriptive research preview. It is not a peer-reviewed publication, a new data collection or a new computational analysis. The accompanying revised results and manuscript are preserved as research records rather than represented as journal publications.

Return to the article · Revised results summary · Revised manuscript draft

Design

The instrument contained 120 authored questions: 40 on obesity and obstructive sleep apnea, 40 on type 2 diabetes and metabolic disease, and 40 on metastatic non-small cell lung cancer. Within each indication were eight questions at each of five decision stages: clinical relevance; evidence and options; eligibility; on-treatment risk and adverse events; and action and access.

Each question was tested on three configured APIs in two text conditions, with three planned repetitions. Both conditions were stateless. The base condition supplied the question; the added-context condition supplied a fixed patient or caregiver prefix, a blank line, and the same question in one user message. Some base questions already contained personal information. Prefixes ranged from 13 to 77 words, with different lengths and clinical detail across indications.

Configured models were Perplexity Sonar (sonar, temperature 0.2), Gemini 2.5 Flash (gemini-2.5-flash, Google Search grounding, temperature 0.2), and GPT-5 (gpt-5, OpenAI Responses API web search, provider-default sampling). These are requested configuration labels, not independently verified returned model-version identifiers.

The collection occurred on September 13, 2026, approximately 09:31–13:03 UTC. Conditions were issued close in time, but prompt order and collection timing were not randomized. Context length, information content and indication were not independently manipulated.

Completion and eligibility

Measure Count
Authored base questions 120
Planned observation records 2,160
Captured responses 2,149
Terminal timeouts, all GPT-5 11
Possible question–system pairs 360
Eligible pairs 357
Captured responses without citations 119

Eligibility required at least two captured repetitions in each condition. Three GPT-5 pairs did not meet that requirement. Sonar and Gemini each supplied 720 captured responses; GPT-5 supplied 709. Citation-free responses comprised 31 Gemini and 88 GPT-5 responses.

The runner allowed up to three attempts per observation, with a 90-second timeout per attempt, and retained the terminal result rather than all intermediate attempts. Thus 2,160 is not the total number of API requests, and 11 is not a complete attempt-level failure count. Some refusal-like answers were counted as captured responses.

Source identity and outcome

The source fields were Sonar citation lists, Gemini grounding chunks and OpenAI URL-citation annotations. Grounding chunks need not equal inline citations. None of these fields establishes the full retrieved corpus, what a consumer necessarily saw, or whether a source supports a generated claim.

The original domain aggregation used custom rules, including separate NIH subdomains. It was not a universal registrable-domain parser. A domain-level measure also misses article substitutions within the same domain.

Jaccard similarity equals the number of shared domains divided by the number of distinct domains in the union. Empty–empty sets received a score of 1; empty–nonempty sets received 0. The endpoint therefore includes changes in whether citations were provided as well as changes in domain identity.

For each eligible question–system pair, the analysis calculated mean similarity separately within each condition and across all available between-condition combinations. For each system, the threshold was the fifth percentile of the pooled separate within-condition means. A pair crossed the threshold when its mean between-condition similarity was lower than that cutoff. Approximate thresholds were Sonar 0.206, Gemini 0.203 and GPT-5 0.111.

The threshold rule was prespecified, but its numerical values were estimated from the test panel. The endpoint is an operational measure of source divergence, not a validated clinical or commercial effect threshold and not an individual significance test. Since thresholds were estimated from pooled panel values, one pair’s classification depends in part on the other questions in that panel.

Main descriptive findings

Configuration Eligible pairs Crossed threshold Rate Mean within-condition Jaccard Mean between-condition Jaccard
Perplexity Sonar 120 58 48.3% 0.780 0.332
Gemini 2.5 Flash with Google Search grounding 120 41 34.2% 0.440 0.258
GPT-5 with web search 117 6 5.1% 0.420 0.365
Overall 357 105 29.4% 0.548 0.318

Between-condition similarity was lower than within-condition similarity in 280 of 357 pairs (78.4%). This is a threshold-free descriptive result, not a count of clinically or commercially important changes. GPT-5 had a smaller measured source response, not an absence of response to context. Differences concern complete API configurations, sampling, search behavior, citation reporting and data-derived thresholds; they do not rank model quality.

The original Wilson 95% interval for the 29.4% rate was 24.9%–34.3%. It does not account for all dependence from repeated questions and estimated thresholds. A post-hoc audit bootstrap with 4,000 draws, resampling questions and recalculating thresholds, gave 23.3%–35.5%. Neither interval turns this authored panel into a representative sample of patient questions.

Prespecified hypotheses

The primary stage hypothesis contrasted S1/S5 (clinical relevance and access) with S3/S4 (eligibility and on-treatment risk). The rates were 31.5% versus 30.8%; a question-clustered GEE model gave an odds ratio of 0.98, 95% CI 0.53–1.81, p = 0.95. The hypothesis was not supported; this is not proof of equivalence.

The hypothesized broad shift toward higher-authority source categories was also not supported. The classification was rule-based and incomplete, and category labels were not clinical-quality assessments. Absence of support for a broad directional shift is not proof that source composition or quality was unchanged.

A later S4/S5 versus S1–S3 regrouping yielded 36.4% versus 24.8%, but was post-hoc, largely driven by Sonar and hypothesis-generating. It is not confirmation that later-stage or high-intent questions are universally more context-sensitive. Indication comparisons were also confounded by differences in prefix length and detail; they do not isolate an oncology-specific effect.

Corrections and sensitivity checks

These checks were conducted after the original analysis. They are not retroactively described as prespecified results.

Check Result
Correct literal www. prefix handling Primary Jaccard threshold flags unchanged: 105/357, 29.4%
Strict registrable-domain definition 28.3%
Exclude citation-free responses, retain at least two responses per condition, recalibrate 101/329, 30.7%
Omit unresolved Gemini redirects, recalibrate 29.7%
Use the 2.5th rather than fifth percentile 19.3%
Use the tenth rather than fifth percentile 38.7%
Correct alternative rank-biased overlap implementation and align threshold construction 23.2%, replacing 40.1% from the original computation

The manuscript’s earlier 39.8% alternative-endpoint figure was a transcription error. The revised record reports the corrected rank-biased overlap result; the public headline uses Jaccard. Different thresholds and endpoints answer different operational questions. They are not interchangeable estimates of a population effect.

The domain-normalization correction changed 211 citation-domain entries and 129 class assignments but left the main Jaccard classifications unchanged. Of 9,833 Gemini redirect entries, 93 did not resolve. A resolved URL identifies a destination; it does not demonstrate readable access or substantiate article content.

Additional post-hoc system comparisons using prompt-clustered GEE and matched Cochran’s Q supported heterogeneity across the tested configurations. The original pooled signed-rank and independent-group chi-square calculations did not fully account for questions repeated across systems and should not be used as the sole inferential support.

Timing, analysis changes and provenance

The recovered record places the protocol, prompt set and runner at commit 31f64ef, September 13, 2026, 09:31:15 UTC, before the first request at approximately 09:31:30.641. The classifier and analysis implementation were committed at 09:40:10, after collection had started. The final analysis used GEE rather than the mixed model proposed in the protocol; this is an analysis deviation.

The records describe development of classifier rules from pooled, condition-blind domains. This is a contemporaneous author-recorded procedure, not independently authenticated blinding. Local Git timestamps do not establish external preregistration. No public registration receipt or DOI is claimed.

The computational review reproduced the main Jaccard figures from the supplied resolved records, pair table and citation table. It did not authenticate provider transactions or establish clinical ground truth. Raw provider responses were not retained in full, and intermediate request logs were incomplete.

The resolved observation data and supplied pair/citation tables were consistent with the recovered repository. A checksum stated for the earlier unresolved observation archive did not match its recovered file and remained unreconciled in the revised record. This limits the provenance account; it does not, by itself, negate the calculations reproduced from the resolved records.

The protocol-deviation record, unresolved checksum and clinical-instrument provenance require further reconciliation before representing the record as submission-ready or fully deposited. The documents linked here are supporting research drafts, not a registered study or a complete public raw-data/code archive.

Interpretation boundary

The supported finding is that added patient or caregiver information was associated with changed source-domain patterns in this authored panel and these configured APIs. It motivates context-sensitive measurement.

The study did not isolate the contribution of each contextual variable; compare validated context packs against personas; evaluate consumer-app personalization, account memory, multi-turn behavior or agent task completion; adjudicate medical correctness; or measure patient choices, conversion, sales or communications effectiveness. It does not validate every component of SFG’s applied method or demonstrate superiority over a competitor.

SFG’s research-led context-pack method is a separate applied methodology: investigate consequential variables, use that evidence to construct decision-specific contexts, and test the resulting scenarios. This study contributes evidence for investigating context, not a completed variable-by-variable validation of that method.

Source documents and disclosure

This technical summary is based on RESULTS_SUMMARY_v1.2, revised September 14, 2026; MANUSCRIPT_DRAFT_v1.2, revised September 14, 2026; and the accompanying SFG-Disclosure-v1.2-Independent-Audit computational review. The downloadable revised results and manuscript are preserved as supplied. “Computational review” does not imply journal peer review or endorsement by an external institution.

Abhi Basu is founder of Signal Fidelity Group, which develops related products and services. Commercial implications are SFG’s interpretations, not outcomes established by this study. This record is not medical advice.

Read the revised results summary · Read the revised manuscript draft · Return to the article