Answer Fidelity · Original research preview
The Disclosure Effect.
Your audience brings context.
Your AI visibility map should too.
Research preview. Not peer reviewed. This study measured AI-reported source patterns, not medical accuracy or patient outcomes.
Before acting on an AI visibility score, ask what decision it actually represents.
A fixed set of prompts can show what appeared under the conditions tested. It cannot, on its own, establish whether those conditions represent the situations your audience brings. SFG’s Disclosure Effect study examines one reason that distinction matters: what changes when the same health question is accompanied by additional patient or caregiver information.1
Same question. Added context. A different source map.
We tested 120 authored health questions across three search-enabled AI configurations, with and without added patient or caregiver context. Each condition was scheduled to run three times, so repeated answers to unchanged prompts supplied a reference for measuring variation.1
In plain language: the reported source sets were different enough to cross a cutoff defined from repeated-prompt variation.1
The questions covered obesity and obstructive sleep apnea, type 2 diabetes and metabolic disease, and metastatic non-small cell lung cancer. The collection produced 2,149 captured responses from 2,160 planned observation records, with 11 terminal timeouts.1
These were controlled tests through developer APIs, not observations of patients using consumer chatbots. The study measured domains reported in citation and grounding records. It did not determine whether an answer was medically correct or whether a patient made a different decision.1
Three configurations. Different responses.
| API configuration | Crossed threshold | Count |
|---|---|---|
| Perplexity Sonar | 48.3% | 58 of 120 |
| Gemini 2.5 Flash with Google Search grounding | 34.2% | 41 of 120 |
| GPT-5 with web search | 5.1% | 6 of 117 |
Each configuration was assessed against its own threshold. These results describe the settings tested on September 13, 2026—not a ranking of medical accuracy, intelligence or overall model quality. GPT-5’s lower threshold-crossing rate does not establish that it ignored context.1
For measurement, the implication is to test the combination of question, context and AI configuration rather than assume that a result from one combination transfers to another.
The words can stay the same. The task can change.
Consider the question: “What are my options?”
One person adds: “I’m preparing for my first appointment and want to understand what to ask.” Another adds: “I’ve already discussed treatment with my clinician; I’m trying to understand an insurance barrier.”
The words of the question are identical. The information needed to help is not. One situation calls for preparation; the other calls for understanding an access problem.
Illustrative scenarios, not study prompts or observed patient conversations.
That is the difference between describing a person and specifying the decision they face. A visibility map built around the broad question should not automatically stand in for both situations.
Adding context is easy. Knowing which context matters takes research.
A profile can describe who is asking. A research-led context pack is built around evidence about which details change what the question is asking—and what the AI system returns.
SFG’s context-pack method starts with variable discovery, not persona invention. Research identifies which contextual variables matter for a defined question or decision. Those findings inform what goes into a pack, which combinations to test, and which differences to measure.
The standard is evidence for including a detail—not simply making the profile more elaborate. The method separates tested variables from working assumptions and asks whether the observed differences hold across relevant questions and AI configurations.
This study supports testing context as part of measurement. Because its added-context passages combined several details, it does not isolate the contribution of each variable. Establishing those contributions requires separate tests; it is not something the 29.4% figure establishes.1
Research-led packs should also be checked against customer interviews, observed questions and other real-user evidence. A model of a decision situation becomes more useful when its assumptions can be challenged—not when its description becomes more vivid.
A useful baseline is not a complete strategy.
A standard prompt panel gives a team a defined baseline. The next step is to ask whether the findings hold when circumstances relevant to the decision change.
The issue is not simply the number of turns in a conversation. Both conditions in this study used single-turn, stateless requests; the difference was the additional information supplied.1
For communications teams, our recommendation is to preserve the baseline and test the decision contexts that could change the interpretation. Does an apparent source priority remain relevant? Does a competitor’s framing still fit? Is important evidence missing from the answer—or simply irrelevant to that particular situation?
Before optimizing for an answer, establish which decision the answer is serving.
As AI moves from answers to action, the assignment gets bigger.
Context is already part of how some consumer AI systems search. OpenAI documents that, when memory is enabled, ChatGPT may use relevant saved memories to rewrite a search query.2
In its May 2026 Search announcements, Google described context carrying forward through follow-up conversations, expanded connections to personal information, and broader agentic booking capabilities that use a person’s criteria to find options and help them take a next step.3
Our strategic conclusion is that measurement should examine the task and its constraints—not only the wording of the first question. An assistant asked to explain a topic and one asked to find a workable next step are not performing the same assignment.
The Disclosure Effect tested one component: additional context supplied in prompt text. Conversation history, account memory, connected apps, tool use and task completion require their own evaluation.1
Map the decision. Find an opening your evidence can earn.
For SFG, the commercial application begins with a consequential audience decision. What is the person trying to understand or do? Which circumstances could change the relevance of your evidence? Where is there an unmet need that your organization can credibly address?
Research-led context packs make those questions testable. Comparing the responses can help identify candidate evidence gaps, competing narratives and source priorities. The next step is strategic judgment: assess those findings against audience needs, the organization’s objectives and the proof it can supply.
A cited website is not automatically an outreach priority. A gap in an answer is not automatically a market opportunity. The communications job is to determine where a useful contribution can be made—and how to make it credible.
The intended output is a decision: which issue to address, which message to lead with, what evidence to supply, and which voices or channels deserve attention. Client-specific opportunities and outcomes require client-specific evidence; they are not findings of this study.1
From insight to communications.
This research informs SFG’s approach to audience intelligence—and the development of CommsRx, our agentic platform for pharmaceutical and biotech communications. CommsRx connects strategic counsel, message adaptation and multilingual multimedia, with your team directing the work. The study informs the approach; it does not validate the product’s performance.
Your next communications decision
Bring the decision. Let’s investigate what changes it.
Tell us which audience you need to understand and which business or communications question you are working through. Explore SFG’s research-led approach to context mapping, evidence priorities and credible competitive openings.
Discuss a decision with SFG ↗
Request research updates ↗
These links open a prefilled email to Abhi Basu. Nothing is submitted by this page; you choose whether to send the message. Please do not include patient-identifying information or individual medical records.
Study design and limitations
Design and unit of analysis. The study used 120 authored health questions, three search-enabled API configurations, two text conditions and three planned repetitions: 2,160 observation records. It captured 2,149 responses and recorded 11 terminal GPT-5 timeouts. The analysis included 357 of 360 possible question–system pairs; three GPT-5 pairs lacked at least two captured repetitions in each condition. The observations are not independent patients or 2,160 different questions.1
What “added context” means. Both conditions were stateless, single-turn API requests. One contained the base question; the other added patient or caregiver text before the same question. Some base questions already included personal information. Added passages ranged from 13 to 77 words and varied in clinical detail. This was not a test of login state, memory, multi-turn conversation or consumer-app personalization, and it did not isolate individual contextual variables.1
What was measured. The analysis compared domain sets extracted from Sonar citation lists, Gemini grounding records and OpenAI URL-citation annotations. These fields are not identical measures of the sources visible to a consumer, and they do not reveal everything a system retrieved or relied on. Domain grouping used a custom rule; it also cannot detect every change of article within the same website.1
How the threshold works. Jaccard similarity is the number of shared domains divided by the number of distinct domains across two sets. For each eligible question–system pair, the analysis averaged similarity separately within each condition and across conditions. A pair crossed the threshold when its mean between-condition similarity was below the system-specific fifth percentile of pooled, separate within-condition means. The rule was prespecified; the numerical thresholds were estimated from the test panel. Crossing the threshold is not a per-question significance test or a validated measure of clinical or commercial importance.1
Empty citation sets and sensitivity. The primary analysis treated two empty source sets as identical and an empty versus a nonempty set as completely different. There were 119 captured responses without citations. In a post-hoc check that excluded those responses, required at least two remaining repetitions in each condition and recalculated the thresholds, 101 of 329 comparisons crossed the threshold: 30.7%. A strict registrable-domain check gave 28.3%. These are alternative analyses, not replacements for the 29.4% primary result.1
The broader descriptive pattern. Between-condition source similarity was lower than within-condition similarity in 280 of 357 comparisons, or 78.4%. Mean Jaccard similarity was 0.548 within conditions and 0.318 between them. These are descriptive comparisons; 78.4% is not a second estimate of threshold-crossing or clinically meaningful change. A post-hoc bootstrap that resampled questions and recalculated thresholds produced a 95% interval of 23.3%–35.5% around the threshold-crossing rate. It describes uncertainty within the authored test panel, not the prevalence of an effect among real patients.1
What the study did not establish. The prespecified hypothesis that certain decision stages would show more divergence was not supported. Neither was the hypothesized broad shift toward higher-authority source categories. The study did not evaluate clinical correctness, patient behavior, conversion, sales, the effectiveness of a communications intervention, or whether a context-pack method outperforms another measurement approach. It does not establish that detailed prompts are higher intent, that all competitors ignore context, or that the results generalize to all users or AI products.1
Research record. The study covers one collection day and only three repetitions per condition. This page uses the revised September 14, 2026 study records. Technical corrections, analysis deviations and data-provenance limits are described in the supporting study record. No peer review or public preregistration is claimed.1
Research and commercial disclosure
Abhi Basu is the founder of Signal Fidelity Group, which develops products and services related to the subject of this research, including CommsRx. The commercial applications described here are SFG’s interpretation and approach, not outcomes demonstrated by the study. This page concerns communications research and is not medical advice.
Sources and study record
1. Signal Fidelity Group. The Disclosure Effect v1.2. Revised results summary and manuscript, September 14, 2026, with accompanying computational review. Research preview; not peer reviewed. Read the study record, revised results summary and revised manuscript draft.
2. OpenAI. Searching the web with ChatGPT. Help Center, accessed September 21, 2026. See “Information shared with search providers” for the documented use of saved memories in query rewriting. Read the source.
3. Google. A new era for AI Search. May 19, 2026. Provider announcement covering conversational context, personal information connections and agentic Search capabilities; it does not establish identical availability for every user. Read the source.