
Ask ChatGPT where to stay in a town, then ask it again tomorrow. You will get two different lists. Every property manager who has tried this has the same reaction: if the answer changes every time, how do you plan for it?
We stopped guessing and measured it. Same 24 questions, every Monday for six weeks, on Perplexity, Gemini, Google AI Overviews and ChatGPT, five times each, in two vacation rental markets where one of our clients operates. About 480 answers a week. This is one client and six weeks, so treat the exact numbers loosely. The series runs through October. The shape is already clear.
What We Measured, and How
The questions are the ones a guest actually types. Four are broad (“where should I stay in Hampton Beach NH”). Fourteen carry one or two requirements (“best beach houses for families in Hampton Beach NH”). Six stack three or more (“pet-friendly rental with a hot tub and parking near the beach”). Same wording every week, so the only thing that can change is the answer.
Each Monday we ask all 24 questions five times on each engine, record every site the engine names, and compare the list to the previous Monday’s list for the same question on the same engine. The score is simple: of the sites named in both weeks combined, what share appeared in both? A score of 1.0 means the identical list. A score of 0 means nothing carried over.
We also track five competing property managers in the same two markets, so a change in our client’s numbers can be read against the field.
How Much of an AI Answer Changes Week to Week
Here is the median overlap per engine across six weeks, with the middle half of questions in parentheses.
- Perplexity: about 7 in 10 sites carry over (0.62 to 0.80). Positions hold.
- Gemini: about half (0.37 to 0.60). Half the sources re-roll each week.
- Google AI Overviews: about half (0.33 to 0.56). Same story, plus Google served no AI answer at all on 20 to 29% of questions, depending on the week.
- ChatGPT: almost none. On two runs of the current model the median overlap was zero. Ask it twice and you are mostly looking at two different lists.
So the slot machine is real. On three of the four engines, somewhere between a third and all of the sources named for a question get replaced within a week. That matches what we found earlier inside a single day, running the same ChatGPT question twice: about half the sources changed between runs.
Why the Bottom of the Answer Re-rolls
The churn is not the engine being careless. It is how the engine works. When a guest types one question, the engine turns it into several hidden searches, runs them, and builds the answer from whatever comes back. Gemini runs about five hidden searches per question in our data, and in about half of broad questions it adds a requirement the guest never typed, like “family” or “with a pool.” Each week those hidden searches land on slightly different pages, so the long tail of the answer shifts.
The sites that survive the shuffle are the ones that match the question no matter how it gets rephrased behind the scenes. A page titled “Plum Island vacation rentals” that actually covers the beach, the drive, the sleeping capacity and the season comes back for “where to stay in Plum Island,” “Plum Island beach houses for families,” and “Plum Island rentals near the refuge.” A page that only answers one phrasing is in the tail, and the tail re-rolls.
That also explains the engine differences. Perplexity leans on a stable set of sources and re-uses them. Gemini and Google’s AI answers pull from a wider net each time. ChatGPT, on its current model, effectively starts over on every run. They all build the answer the same way. They differ in how much of the source list they keep.
The Part That Does Not Move: Anchors
The re-rolling happens at the bottom of the answer. The top holds.
Our client was named for the same 19 of the 24 questions in all six runs. Not 19 on average. The same 19, every week. Airbnb showed up for 22 questions every week, Vrbo for 21. Week to week, the client kept 19 to 21 questions, lost between zero and three, and gained between zero and two. None of the five competitors we tracked came close, and none of them moved in a way that tracked the client’s numbers, which is what you would expect when the client’s numbers did not move either.
We call a site like that an anchor: named for the same question every week regardless of what else changes around it. Every AI answer has two layers. A few anchors, and a tail that gets reshuffled every time. The reshuffling is what people notice, because it is what changes. The anchors are what decide the booking, and they are not random.
The overall presence number tells the same story from the other side. Across Perplexity, Gemini and AI Overviews together, the client appeared in 44 to 49% of answers every single week. Newburyport and Plum Island questions ran 55 to 64%. Hampton Beach questions ran 29 to 38%. Six weeks, nothing published, and the line is flat. That is the baseline any change has to clear before it counts.
Which Pages Held the Spot
When an engine cites the client, we record the page. Three pages carry the site.
- The Plum Island location page, 192 citations over six weeks, first position on Google AI Overviews.
- The Newburyport location page, 159 citations.
- The Hampton Beach location page, 109 citations.
That is roughly 460 citations from three pages, each one answering “where should I stay” for one town in more depth than anyone else in that town does. The homepage is sixth on the list. And not a word on those three pages changed during the six weeks. The client has five new pages drafted for a pilot, and none of them had gone live by the end of the window, which is exactly why this series is a clean “before.”
This lines up with what we found across nearly 7,000 citations of property manager sites in twenty markets: seven in ten point to an interior page, not the homepage. Our guide to AI search for vacation rentals covers which page types win which questions.
What This Means for a Property Manager
Stop reading single runs. If someone shows you one screenshot of an AI answer as proof you are in or out, they have shown you a coin flip. On ChatGPT, one run is an anecdote. Ask five times, on more than one engine, and look at what repeats. That is also how we measure everything we publish.
The goal is a spot that holds, not a one-time win. A page that answers the question a guest is asking, in your town, with real detail, earns a place the engine keeps coming back to. That is the whole mechanism we have seen so far. Backlinks, schema and homepage redesigns did not separate the cited operators from the skipped ones in our larger study. Local pages did.
Judge your visibility weekly, on a fixed set of questions. Pick the ten questions your guests ask about your market, ask them every week, and count how many you hold. One week up or down is noise. Three weeks in a row on the same question set is a position.
Engines differ, so measure more than one. Perplexity is the steadiest engine for this client and the one most likely to cite an independent operator. ChatGPT re-rolls almost everything and, for this client, only started naming the site at all after a model change in late August (one answer in 120, then five in 120). We would not build a plan on either extreme alone.
What We Do Not Know Yet
This is one property manager in two coastal markets, measured for six of a planned twelve weeks. The percentages will move. The pattern, a stable top and a churning tail, has held on every run so far and matches the day-to-day churn we measured separately. We will publish the full twelve weeks, including what happens to the page table when the client’s new pages go live, because new pages entering the list is the one signal here that has no baseline noise at all.
Frequently Asked Questions
Do AI search results change every time you ask?
Partly. Across six weekly runs of the same 24 vacation rental questions, Perplexity kept about seven in ten of the sites it named from one week to the next, Gemini and Google AI Overviews kept about half, and ChatGPT kept almost none. The sites at the top of the answer tended to repeat every week; the rest of the list was reshuffled.
Which AI engine gives the most consistent answers?
In our six-week series Perplexity was the most consistent, keeping about seven in ten cited sites week over week on the same question. Gemini and Google AI Overviews kept about half. ChatGPT was the least consistent, with almost no overlap between consecutive weeks on the current model.
How do you measure AI search visibility reliably?
Ask a fixed set of questions, several times each, on more than one engine, on a regular schedule, and count how many questions name you in every run. A single answer is a coin flip on some engines. A site that appears for the same question week after week is holding a position; we call that an anchor.
Why does one property manager keep showing up in AI answers when others do not?
In our data the sites that hold a spot week after week have a page that answers the specific question in depth, usually a location or guest-type page rather than the homepage. Our client held 19 of 24 questions for six straight weeks on the strength of three location pages that were not changed during the test.
Does being cited by AI once mean you will stay cited?
Not on its own. Between a third and all of the sites named for a question can be replaced within a week depending on the engine. The sites that stay are the ones with the best page for that question in that market. Being cited once is a signal; being cited six weeks running is a position.
Ask an AI where to stay in your market tonight, then ask it again next Monday. Whoever is there both times is your real competition, and the page that put them there is the one to study. The Sightline Report runs that check across your market and tells you which page to build first.

