Skip to main content
Back to Community Pulse
Methodology · v1.1

How Community Pulse is built

We want you to be able to trust what's on these pages — and to know exactly where to push back when you shouldn't. This page is long on purpose. Skim what you want, skip what you don't.

1. What data we use

Two layers — read this first. The live market-temperature score on each community comes from DLD market signals (transaction velocity, asking-price compression, momentum and stability). It refreshes with our data pipeline. The Reddit corpus described below is the historical qualitative layer (2021–2026): it powers paraphrased pull-quotes and aspect themes, shown as dated context — not a live score or resident survey.

The qualitative layer is built from four subreddits where people who live in, or are thinking about moving to, Dubai actually talk about neighborhoods: r/dubai, r/UAE, r/expats, and r/digitalnomad — a snapshot spanning 2021–2026.

The initial archive was extended with an authorised Arctic Shift export. The combined archive contains 23,330 area-item records sent to the AI scorer, with the latest collected item dated July 21, 2026. Deleted or removed content, off-topic chatter, non-Dubai locations, advertiser-heavy subreddits and property listings are excluded before publication.

After scoring and applying a confidence threshold (more on that below), 5,720 items cover 80 Dubai areas. Each area page shows you its own sample size and its trend over time.

A note on volume. The Reddit conversation about Dubai real estate has roughly tripled between 2024 and 2026. That means newer years have more data, which can make an area look like it's suddenly trending when it's really just being discussed more. We mitigate this by reporting mean sentiment (not raw counts) and suppressing any quarter with fewer than 5 on-topic mentions.

2. How we score

Every post and comment goes through Anthropic's Claude Haiku 4.5. We give the model the item text plus the area it was tagged with, and ask it to return a structured JSON object:

  • on_topic — is this actually about this area, or noise?
  • sentiment — -1 (very negative) to +1 (very positive)
  • aspects — up to four tags from a fixed vocabulary (price, traffic, community, noise, schools, etc.)
  • pros / cons — up to three short paraphrased claims each
  • persona_signal — resident, considering a move, investor, tourist, professional commenter, or unknown
  • confidence — 0 to 1, how sure the model is

We only use items where on_topic = true AND confidence ≥ 0.5. The rest are discarded — typically about 30–40% of the raw sample. Area pages show both numbers so you can see the drop-off.

Sentiment for each area is the weighted mean of individual-item sentiment, weighted by confidence × log(1 + Reddit upvotes). High-confidence, well-upvoted comments carry more weight than low-confidence throwaway ones — but zero-score items still count.

The 3-year sentiment trend is a linear-regression slope over the last 4 available quarters of rolling sentiment, expressed per year.

3. What this does NOT tell you

Community Pulse is useful because it's the lived-experience signal that listings sites never show. But it has important limitations you should keep in mind.

The July 2026 increment is incomplete

Source query caps can omit high-volume discussions, and comment collection only covers posts returned by those searches. Treat July 21 as the latest collected item date, not a claim that every discussion through that date is present.

  • Run at 2026-07-21T20:29:52Z before requested end 2026-07-22T00:00:00Z.
  • Arctic Shift API returns max 100 results per search query; high-volume communities may have partial coverage.
  • Comment collection limited to posts returned in search results; not all comments for matched posts may be included.

This is Reddit, not Dubai

Our sample is English-speaking and expat-heavy. Long-term Emirati residents, Arabic-speaking tenants, labour-camp residents, and most of the city's working population are under-represented or absent. Treat this as a slice, not a consensus.

Sample sizes vary — a lot

Dubai Marina has ~160 analysed (on-topic) mentions. Al Barari has ~40. Newer communities may have single-digit coverage. We suppress any quarterly data point with fewer than 5 items, and leaderboards require at least 30 on-topic mentions, but small-n noise is still possible in aspect breakdowns. Trust the direction more than the decimal.

AI scoring is imperfect

Claude gets most things right and often picks up sarcasm and implicit sentiment. It also makes mistakes. We paraphrase every claim rather than quoting the user verbatim, and every pull quote links to the original Reddit thread so you can judge for yourself. If a claim looks wrong, the evidence is one click away.

Generic area names over-match

Some Dubai areas have names that are also common English phrases — "City Walk", "Downtown", "The Valley", "Meadows", "Town Square". The raw scraper matches on the phrase, which can grab discussions that aren't actually about that community. Claude's on-topic filter catches most of these, but expect slightly higher noise on these areas.

Astroturfing exists — we don't filter it

Real estate brokers are active on Reddit. Some of what looks like resident enthusiasm or resident complaint is professionally posted. v1 of this system has no bot filter. We label comments as "professional commenter" when the AI can tell, and the persona mix for each area shows you how much of the discussion looks industry-driven.

4. How often this updates

Collection and scoring run periodically rather than live. Each area carries its own latest-source date, while the hub discloses the newest item in the combined archive and whether collection was complete.

Re-scoring the same items is wasteful, so the pipeline is incremental: new items get scored, existing items keep their original score unless we intentionally rerun them.

Methodology version: v1.1. If the scoring prompt, confidence threshold or weighting changes, the version increments and we'll note it here.

Flood exposure rating

We flag communities reported flooded in the April 2024 event (Dubai's heaviest rainfall in ~75 years) with one of two tiers: Severe (prolonged standing water, evacuations) or Reported (waterlogging). Severity is a conservative editorial read of public reporting, with sources cited.

We deliberately do not derive flood risk from elevation. Dubai is flat and flooding is driven by storm-drainage capacity, not height — an elevation proxy mislabels low-lying coastal communities (Dubai Marina, Palm, JBR) that drain straight to the sea and did not flood, while the inland master-communities that did flood sit at higher elevation. So elevation is shown only as context, never as a risk score.

This is indicative only — not an engineering, drainage, planning or insurance assessment. There is no official public per-community flood dataset for Dubai, so this reflects public reporting of a single event. A community not listed was simply not reported flooded — that is not evidence of safety. Corrections welcome — email us.

5. Report a problem

If a pull quote misrepresents the thread it links to, or an area page draws a conclusion that the underlying posts don't support, we want to know. It sharpens the prompt and re-scores the offending items.

Email hello@dubuy.ai with:

  • the area page URL
  • the specific claim or quote
  • what you think it should say instead (or a pointer to the Reddit thread that contradicts it)

As of this version

Items scored

23,330

High-confidence

5,720

On-topic items

Areas covered

80

Dubai communities

Methodology

v1.1

Current version