Hotel Guest Intelligence From 67,962 Reviews: The Hark Case Study
Nearly 68,000 guest reviews across 16 luxury hotels, read sentence by sentence and turned into a private portal, a one-page brief and a monthly owner report per hotel, with the guests' own words behind every number.
- 67,962
- reviews read
- 285,344
- labelled observations
- 84
- aspects under 16 topics
- 16
- hotels live on one score

A star rating compresses everything a guest said into one number. For a set of luxury hotels in one city we built Hark, the product that reads the sentences instead: every public review of a hotel and of the hotels it competes with, labelled against a hospitality ontology and served to the general manager as a private portal, a one-page brief and a monthly owner report.
The problem
A general manager knows the last bad review. Nobody knows whether check-in complaints are trending up quarter over quarter, whether breakfast sentiment beats the hotel across the road, or which specific aspect of the stay is dragging the score down before it reaches the rate. Chain tools benchmark against the world; reading reviews by hand does not scale past last week's worst one; and a rating averages away exactly the detail that is fixable.
What we built
- An engine that reads everything. Every public review from Booking.com, TripAdvisor, Google and Agoda, deduplicated into one archive, split into sentences, and labelled by a frontier language model against a 16-topic, 84-aspect ontology that grew from the reviews themselves across 26 versions. Each label carries a sentiment, a severity, an intensity, a confidence and the sentence as evidence: 285,344 observations from 67,962 reviews, 4.2 per review.
- A score a hotelier can read. Observations are confidence-weighted so one furious paragraph does not outweigh ten mild mentions, then rescaled to a guest score out of 10. Six departments roll the topics up. Every hotel in a competitive set is processed identically, so a gap against a competitor is an apples-to-apples gap.
- A portal for the daily open. Today: the four KPIs with their change, the platform ratings, alerts that fire only with volume behind them and carry the guest's sentence, department health, and what guests praise and complain about. Trends: 24 months of guest score, what improved and what declined this quarter, and a topic detail with the guests' own words. Compare: the named set ranked, a department scorecard with the leader per department, where you win and where you trail. Actions: recommended actions in three panels, an estimated revenue range for the current score drift, and an intervention log.
- A loop, not a report. A GM logs a change; Hark compares sentiment on that topic before and after and reports Improving, Worsening or Unchanged in points, or "Measuring" until enough guests have written.
- Two documents that write themselves. A one-page Guest Voice Brief, the first thing a GM sees, with the top five strengths and pain points and two guest quotes under each. A monthly owner report that opens with the period in one generated paragraph.
- Isolation by construction. One pre-computed bundle per hotel; competitors inside it as aggregates only, enforced at export; one private 192-bit link per hotel, a 404 for anything else, no database in the request path. Refreshed weekly, with every screen carrying the date its data runs through.


How it works
Labelling is parallel and checkpointed in 50-review chunks, so a stopped run resumes where it stopped and never re-bills finished work. We benchmarked frontier and open models before committing: the configuration we shipped runs a cloud model with reasoning disabled and kept roughly 95% of full-reasoning quality at about a quarter of the cost; local models on a consumer GPU reached 46% and 62% of that quality and were rejected. A 529-review backlog labels in about 25 minutes for about a dollar. Each of the four feeds is watched for freshness per hotel, so a feed that goes quiet is noticed while the others keep flowing.
What it proves
Unstructured feedback becomes an auditable instrument: every chart drills to the sentences behind it, every alert carries its evidence, every money figure names its source and hides itself when the data is thin. That is our AI practice in one line: not a chatbot, but structure at scale that a general manager can verify on a phone between two meetings.
Hark is available as a product.
Every screen shown is Hark's demo hotel: the real product over an invented property, an invented competitive set and invented guests.
Frequently asked questions
- Is it fair to compare a hotel against different hotels?
- The comparison is relative, not absolute: identical sources, identical processing. A hotel with more reviews does not score higher; volume affects confidence, not the score.
- Can competitors see a hotel's data?
- Other hotels see only aggregate scores derived from public reviews, exactly what the hotel sees about them. The private link, the intervention log and the detail views are private, and the server never loads another hotel's data to answer a request.
- What about fake or malicious reviews?
- Only platform-verified reviews are ingested, every observation is weighted by confidence, and nothing is flagged without volume behind it. A single review, genuine or not, cannot move the numbers meaningfully.
- Where does the revenue number come from?
- From published hospitality research that ties each point on a 10-point review scale to a band of 4 to 8 percent of annual room revenue, applied to the hotel's own rooms and rate. It is shown as a range and labelled a directional estimate, never a forecast or a measured outcome.