The question every single-site study should be asked
Everything published in this research programme comes from one corpus about one place. That is the standard weakness of this kind of work, and the standard response is to acknowledge it in a footnote and carry on.
We had a second corpus available β 4,291 rated reviews about the Colosseum, collected separately, months apart, sharing no reviews with the Vatican set β so instead of acknowledging the limit we tested it. We re-ran fifteen of our measurements on the other monument and compared.
The point of this article is not to say more about either building. It is to separate the findings that describe tourism from the findings that describe a specific address.
The result that stopped us
We had already published that longer Vatican reviews rate lower β 4.68 stars in the shortest tenth against 3.39 in the longest, a drop of 1.29. Running the identical split on the Colosseum returned 4.70 down to 3.41. A drop of 1.29. The same number, to the second decimal, from a dataset collected by a different effort about a different monument in a different part of Rome.
Review length vs rating: the same curve in two different buildings
Vatican, shortest tenth4.68 β
Colosseum, shortest tenth4.70 β
Vatican, longest tenth3.39 β
Colosseum, longest tenth3.41 β
Bars start at 3.35. Vatican drop: β1.29. Colosseum drop: β1.29. Neither corpus knew about the other.
7,714 and 4,291 rated on-topic items, split into ten equal groups by character count. Source: Vatican Tour Research Corpus 2026 and the Colosseum research corpus.
That is not a finding about the Vatican and it never was. It is a finding about how people write reviews: satisfaction produces a sentence, dissatisfaction produces paragraphs. Anyone reading any review site should carry it β the most detailed accounts are systematically the least positive ones, everywhere we have looked.
Fourteen of fifteen held
Every measurement is expressed as a distance from its own corpus average, which matters because the two baselines differ: 4.08 at the Vatican, 4.40 at the Colosseum. Comparing raw ratings across the two would be meaningless. Comparing distances is not.
| What the review mentions | Vatican | Colosseum | Verdict |
| Guide quality | +0.28 | +0.18 | Replicates closely |
| Audio guide | −0.10 | −0.24 | Replicates closely |
| Visiting with children | −0.31 | −0.19 | Replicates closely |
| Duration and pacing | +0.22 | +0.08 | Replicates closely |
| Toilets | −0.62 | −0.42 | Replicates closely |
| Engaging with the history | +0.62 | +0.40 | Replicates |
| Cancellations | −2.86 | −3.14 | Replicates |
| Crowding | −0.60 | −0.31 | Replicates |
| Price | −0.82 | −1.14 | Replicates |
| Meeting points | −0.51 | −0.95 | Replicates |
| Accessibility | −0.80 | −0.36 | Replicates |
| Group size | −0.71 | −0.49 | Replicates |
| Weather | −0.49 | −0.11 | Replicates |
| Scams and resellers | −1.71 | −2.45 | Same direction, larger gap |
| Photography | −0.05 | +0.31 | Inverts |
Each figure is the difference from that corpus's own average β 4.08 at the Vatican, 4.40 at the Colosseum β so the two columns are comparable even though the baselines are not. Source: Vatican Tour Research Corpus 2026 and the Colosseum research corpus.
The five marked as replicating closely land within 0.15 of each other. Guide quality helps by roughly the same amount at both sites; audio guides underperform at both; visiting with children costs something at both; toilets are a mild drag at both. None of that is obvious in advance, and none of it could be claimed from one corpus.
The one that inverted, and why it is the most useful row
Photography sits at β0.05 at the Vatican and +0.31 at the Colosseum.
Same word, opposite sign, and the explanation is not statistical β it is architectural. Photography is banned in the
Sistine Chapel, so Vatican visitors mention it when they are being told to put the phone away. At the Colosseum photography is most of what people came to do, so they mention it when they are delighted.
That single row is the argument for doing this exercise at all. Without the comparison, a reader could reasonably take any of our tag findings as being about visiting monuments in general. This one proves that some of them are not, and that the tag is measuring the visitorβs relationship to a rule rather than a subject.
It also sets a standard for the rest of the programme: a finding that has not been tested somewhere else should be read as being about this building until proven otherwise.
What we could not test
Three of our strongest Vatican findings have no verdict here, and we would rather say so than quietly omit them. Staff, wayfinding and the herding complaint all fall below usable sample sizes on the Colosseum side once the necessary exclusion is applied β the Colosseumβs TripAdvisor venue sample was collected with a deliberate 1β3 star filter and had to be removed, and it took most of the monument-specific commentary with it.
So the claim that staff are the lowest-rated element of a Vatican visit remains a single-corpus finding. It may well generalise. We simply have not shown that it does.
The verdict
Two corpora are not a sample of monuments, and this article does not pretend otherwise. What it establishes is narrower and still worth having: the review-length effect is not about the Vatican, the guide and context effects hold in a second place, and at least one of our findings is genuinely local. For a reader, the practical takeaway is the one from the length curve β when you are researching any attraction, the long reviews are the angry ones, and weighting them as more informative because they are more detailed will systematically mislead you. That holds in both buildings we have measured, and we would expect it to hold in yours.
Common questions
β Do longer reviews really mean worse experiences?
In both corpora we have measured, yes, and by almost exactly the same amount. Vatican reviews fall from 4.68 stars in the shortest tenth to 3.39 in the longest; Colosseum reviews fall from 4.70 to 3.41. Both are a drop of 1.29 stars. Satisfied visitors write a sentence and dissatisfied ones write paragraphs, which means the most detailed reviews on any site are systematically the least positive.
β Which findings hold at more than one monument?
Fourteen of the fifteen we tested replicate in direction. Guide quality, audio guides, visiting with children, duration and toilets land within 0.15 of each other at both sites. Engaging with the history, cancellations, crowding, price, meeting points, accessibility, group size and weather all point the same way with larger gaps. Only photography inverts.
β Why does photography rate differently at the two sites?
Because it means different things in each building. Photography is banned inside the Sistine Chapel, so Vatican reviews mention it in the context of being stopped, giving β0.05. At the Colosseum it is a main activity, so reviews mention it in the context of enjoying themselves, giving +0.31. The measurement is picking up a visitorβs relationship to a local rule, not an attitude to photography.
β Does this prove your Vatican findings are correct?
No, and it is worth being precise. Replication shows a result is not an artefact of one dataset. It does not establish causation, and it does not validate findings we could not test β staff, wayfinding and the herding complaint remain single-corpus results because the Colosseum samples were too small after a necessary exclusion. Two corpora also are not a sample of monuments.
The Colosseum team have published their own read of the same comparison β written from the position of the corpus that did the testing rather than the one being tested β as
Testing the Vatican Findings Against the Colosseum.
Author and Method
Research and analysis by the Intercoper Curator Team for VaticanTourGuides. Reviewed by Mario Dalo, founder of Intercoper.
Corpora: the Vatican Tour Research Corpus (22,771 items, 7,714 rated on-topic, corpus average 4.08) and the Colosseum research corpus (4,291 rated on-topic after exclusion, corpus average 4.40). The two were assembled by separate collection runs months apart and share no reviews.
A mandatory exclusion on the Colosseum side. That corpus contains 1,928 TripAdvisor reviews attached to the Colosseum venue that were collected with a deliberate 1β3 star filter to surface pain points; they contain zero 4 and 5 star ratings. Every Colosseum figure in this article excludes them. Including them would depress the Colosseum baseline from 4.40 to 3.79 and make the length gradient appear twice as steep as it is β which is precisely the false result we obtained on a first pass before checking the rating distribution by venue.
Method: each figure is the difference between the average rating of reviews carrying a given enrichment tag and the average of that corpus as a whole. Differences rather than raw ratings are compared, since the two baselines differ by 0.32. Categories overlap within each corpus. Tag assignment was automated per-item enrichment applying a consistent rule at scale.
What replication does and does not establish. It shows a result is not an artefact of one datasetβs assembly. It does not establish causation, it does not extend to findings we could not test, and two corpora are not a sample of monuments. Sample sizes range from 49 to 3,163 and every one is stated in the underlying articles.