
Sydney-based AI search monitoring company Somantra has released a study showing that small changes to insurance-shopping queries can substantially change how ChatGPT describes brands, even when the same brands continue to appear in the answers.
The study analyzed 4,445 ChatGPT responses to Australian insurance-shopping queries containing brand comparison tables. Researchers identified 135 matched query pairs in which a single decision-related word, such as “cheapest,” “safest” or “most trusted,” was added to an otherwise similar query and the same brand appeared in both responses.
Somantra said the analysis was designed to examine how sensitive an AI system’s description of a brand is to small changes in the wording of a consumer’s request. The company refers to the method as perturbation testing.
Across the 135 observations, the median word shift was 0.94 on a scale from 0 to 1. Somantra said this meant that a typical modified query changed almost all of the visible wording used to describe a brand. The median semantic distance was 0.53, indicating that the differences extended beyond individual word substitutions to changes in the underlying meaning of the descriptions.
The study found that price-related terms produced some of the largest changes. Queries modified with “cheapest” or “most affordable” recorded an average word-shift score of 0.97. The word “best” produced the largest shift in reasoning among the tested modifiers, with a median semantic distance of 0.60.
“Most trusted” generated the largest positive sentiment movement in the study. Across 16 observations using that modifier, Somantra reported a +0.50 mean net sentiment shift.
Other modifiers also produced measurable changes. Somantra reported 20 observations for “most popular,” 19 for “easiest,” 17 for “best,” 17 for “cheapest,” 16 for “most trusted,” 13 for “most affordable,” 13 for “most reliable,” 13 for “safest,” six for “premium” and one for “worst.”
The company found differences across four insurance categories. Home insurance accounted for 46 observations, followed by roadside assistance with 34, motorcycle insurance with 31 and travel insurance with 24.
Home insurance showed the largest observed movement. Across 46 clean observations, Somantra reported a median word shift of 0.98. Roadside assistance, motorcycle insurance and travel insurance also showed measurable changes, although their reported movement was smaller.
One example in the study involved AAMI and QBE in a home-insurance query. Without an additional decision word, ChatGPT described AAMI as a large mainstream insurer offering home and contents cover. When “safest” was added to the query, the resulting description used a more detailed value and coverage rationale linked to Queensland performance and a named industry award.
QBE also appeared in the matched response but was framed differently, with the answer emphasizing comprehensive and flood coverage and referring to a named guide. The example illustrates the distinction Somantra is examining: the brand can remain present while the reasons given for considering it change.
The 25 brands represented in the observations were 1Cover, AAMI, AANT, AIG, Allianz, Apia, Budget Direct, Cover More, GIO, InsureandGo, National Motorcycle Insurance, NRMA, NRMA Insurance, QBE, RAA, RAC, RACQ, RACT, RACV, Shannons, Suncorp, Tick Travel Insurance, Travel Insurance Direct, World Nomads and Youi.
Somantra cautioned that some brand labels should not automatically be treated as a single entity in analysis. In particular, NRMA and NRMA Insurance were retained as separate source labels in the research dataset.
The company has made research materials available through a public GitHub repository, including the response data, analysis scripts, modifier information and files used to calculate word, semantic and sentiment changes.
The word-shift measurement is based on Jaccard distance between tokens in the two brand descriptions. For semantic distance, the researchers used OpenAI’s text-embedding-3-small model to compare the descriptions using embeddings. Sentiment movement was calculated using the Hugging Face cardiffnlp/twitter-roberta-base-sentiment-latest model, with the reported net sentiment change based on movements in positive and negative probabilities.
Those measures capture different aspects of the response changes. A high word-shift score indicates that the wording changed substantially, while semantic distance attempts to determine whether the underlying rationale changed. The sentiment figure is a model-derived measurement and does not represent a survey of consumer trust or purchasing intent.
Somantra’s public methodology also places limits on how the results should be interpreted. The 135 matched observations were reconstructed from historical query logs rather than produced as a randomized experiment in which every base query was systematically run with every modifier. The company describes the results as evidence of observed sensitivity rather than proof that an individual word caused a particular change.
The researchers also note that the current corpus does not provide the repeated identical-query controls needed to establish how much response variation ChatGPT can produce without a modifier being changed. That creates a limitation when attributing the observed differences specifically to the added decision word.
A stronger follow-up study would repeatedly run identical base queries and modified versions of those queries so researchers could establish a baseline level of normal model variation and compare it with the changes associated with individual modifiers. Somantra identifies establishing that “noise floor” as a direction for future research.
The distinction is significant because the study shows an association between query changes and different brand descriptions, but it does not establish that adding a specific word will always produce the same result. It also does not show that the changes lead to different purchasing decisions among consumers.
Somantra published the research materials earlier in August, while the company issued a new release on August 17, 2026 announcing the findings. In that release, Somantra founder Arun Prasad said brands have traditionally focused on whether they are mentioned in AI search, while the company’s research examines whether the surrounding description helps them get chosen.
The study adds another dimension to measuring brand visibility in AI-generated answers. Instead of counting only whether a company appears in a response, Somantra’s approach examines the language and reasoning used to explain the company’s inclusion when the user’s decision criteria change.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.

