Scraping a page once is easy. The hard part comes on the second run, and every run after that: working out what actually changed since last time, and telling the right people without flooding them with noise. In my experience, most scraping systems don't fail at extraction. They fail at scraped data change detection. Every run looks like a wall of "new" records, alerts fire for whitespace edits, and nobody trusts the notifications anymore.
If you monitor prices, listings, permits, compliance records or competitor catalogues, the value is in the change, not the snapshot. A client doesn't pay to be told a property has 14 open violations. They pay to hear, within minutes, that it has a new one.
In this article I'll walk through how I design change detection for scraped data: stable record identity, normalisation, content hashing, field-level diffs, and turning those diffs into alerts people actually read. The examples use Laravel and MySQL, but the pattern works in any stack.
Why comparing raw scraped pages doesn't work#
The naive approach is to store the raw HTML or JSON from each run and compare it with the previous one. It breaks almost immediately:
- Pages change without the data changing. Session tokens, timestamps, ad slots, "last updated" footers and randomised CSS class names make every fetch different.
- Order isn't stable. Many sources return the same rows in a different order each time, so a line-by-line diff reports everything as changed.
- Formatting drifts.
"1,200","1200"and"1200.00"are the same value."Open "and"OPEN"usually are too. - Pagination shifts. One new record at the top pushes every other record to a different page, and a page-level diff sees churn across the whole dataset.
The fix is to stop comparing documents and start comparing records. Each scraped item becomes a row with a stable identity and a normalised set of fields, and change detection happens per record.
Step 1: Give every scraped record a stable identity#
Change detection is only as good as your answer to "is this the same thing I saw yesterday?" Before you store anything, decide on a natural key for each record type:
| Source type | Good natural key | Avoid |
|---|---|---|
| Government registers | Official record or case number | Row position, page number |
| Product catalogues | SKU or the store's product ID | Product title |
| Property data | Parcel or building identifier plus record type | Street address text |
| Job or listing boards | Listing ID from the URL | Scrape timestamp |
When the source has no ID, build a composite key from the fields that define the entity and never change, then hash it. Keep the key's inputs in one place in code, because changing them later makes every record look new.
function recordKey(string $source, array $raw): string
{
// Only fields that identify the entity, never fields that describe it.
return hash('sha256', implode('|', [
$source,
strtoupper(trim($raw['agency'])),
preg_replace('/\s+/', '', $raw['case_number']),
]));
}If you can't find a stable key at all, treat that as a design problem to solve with the client before building alerts. Any alert built on a guessed identity will produce false "new" and "removed" events.
Step 2: Normalise before you compare#
Normalisation turns the messy values a page shows into the canonical values your system stores. Do it once, at ingestion, and store only the normalised form for comparison.
Typical rules:
- Trim and collapse whitespace, and standardise case for enum-like fields such as statuses.
- Parse numbers and money into integers (store cents, not
"$1,200.00"). - Parse dates into ISO 8601 in a single time zone.
- Map source vocabulary to your own:
"CLOSED","Closed - Resolved"and"RESOLVED"might all becomeresolved. - Sort any list-valued field, so order changes don't register as edits.
- Drop fields that are volatile by nature: view counts, "fetched at" stamps, tracking parameters in URLs.
Keep the raw payload too, in a separate column or object store. When a client asks why an alert fired, you want to show exactly what the source said, not just your interpretation of it.
Step 3: Hash the fields you care about#
Once records are normalised, a content hash makes the "did anything change?" check cheap. Hash only the fields that matter to the business, in a fixed order:
function contentHash(array $normalised, array $trackedFields): string
{
$subset = [];
foreach ($trackedFields as $field) {
$subset[$field] = $normalised[$field] ?? null;
}
return hash('sha256', json_encode($subset, JSON_UNESCAPED_UNICODE));
}The tracked field list is a product decision, not just a technical one. For a compliance product, a change in status or penalty_amount matters. A change in the inspector's display name probably doesn't. Getting this list right is the single biggest lever on alert quality.
The storage model I use is simple: a current-state table keyed by the record key, and an append-only change log.
CREATE TABLE scraped_records (
record_key CHAR(64) PRIMARY KEY,
source VARCHAR(64) NOT NULL,
content_hash CHAR(64) NOT NULL,
data JSON NOT NULL,
first_seen_at DATETIME NOT NULL,
last_seen_at DATETIME NOT NULL,
missing_runs INT NOT NULL DEFAULT 0
);
CREATE TABLE record_changes (
id BIGINT AUTO_INCREMENT PRIMARY KEY,
record_key CHAR(64) NOT NULL,
change_type ENUM('created', 'updated', 'removed') NOT NULL,
diff JSON NULL,
detected_at DATETIME NOT NULL,
INDEX (record_key, detected_at)
);On each run, the comparison is a lookup by key and a hash comparison. If the hash matches, you only bump last_seen_at, which is cheap enough to do for hundreds of thousands of records. MySQL's INSERT ... ON DUPLICATE KEY UPDATE handles the "seen again, unchanged" path in bulk.
Step 4: Compute field-level diffs, not just "changed"#
A hash tells you that something changed. An alert needs to say what changed. When the hash differs, compute a field-level diff against the stored data and write it to the change log:
function fieldDiff(array $old, array $new, array $trackedFields): array
{
$diff = [];
foreach ($trackedFields as $field) {
$before = $old[$field] ?? null;
$after = $new[$field] ?? null;
if ($before !== $after) {
$diff[$field] = ['from' => $before, 'to' => $after];
}
}
return $diff;
}That diff is what makes an alert useful: "Status changed from open to resolved" is actionable, "Record 48213 updated" is not. It also gives you an audit trail for free, which matters when a customer disputes whether a change happened. I covered the wider pattern in my post on tamper-evident audit logging.
Handling removed records without false alarms#
Deletions are where most change detection systems embarrass themselves. A record missing from one run doesn't mean it was removed. The source might have timed out on page 37, a filter might have changed, or the site might have been briefly down.
A few rules keep removals honest:
- Only mark removals after a complete run. If any page in the run failed, skip removal detection for that run entirely.
- Require several consecutive misses. Increment
missing_runseach time a record isn't seen, and emitremovedonly after two or three misses in a row. Reset the counter when the record reappears. - Sanity-check the run size. If a run returns 40% fewer records than the last successful one, treat it as a failed run, not a mass deletion. Alert the engineering team, not the customer.
The same instinct applies to "created" events. After you add a new source or change a key, the first run will mark everything as new. Run it silently as a baseline and start alerting from the second run.
Turning diffs into alerts people trust#
With a clean change log, alerting becomes a separate, much simpler problem. I treat each change row as an event and fan it out to subscribers asynchronously:
- The scraper writes change rows in the same transaction as the record update.
- A queued job picks up new change rows and matches them against each user's watch rules (a specific building, a record type, a threshold).
- Matching changes are delivered through the channels the user chose: in-app, email, real-time push or webhook.
Because delivery is queued, a slow email provider never slows down ingestion. Keep jobs small and retry-safe, as I described in Laravel queue architecture for production, and give each delivery an idempotency key built from the change ID and subscriber ID, so retries never send the same alert twice.
Two product details make a large difference to trust:
- Batch low-urgency changes into digests. Not every change deserves an instant notification. Let users choose instant alerts for critical fields and a daily summary for the rest.
- Show the before and after. Every alert should include the field-level diff and a link to the record, so the user can verify it in seconds.
Example from real delivery: Violerts#
On Violerts, a NYC compliance monitoring SaaS, the product's whole promise was change detection. The platform consolidates fragmented NYC municipal property data from multiple agencies into a single compliance view, and property owners and managers use it to find out quickly when something changes on a building they care about.
I led the modernisation of the Laravel backend and React frontend, including multi-agency data ingestion, asynchronous scraping and real-time alerts. The architecture follows the pattern in this article. Ingestion from public sources such as NYC Open Data and scraped agency pages runs as queued work on Redis, supervised with Laravel Horizon. Each source's records are normalised into a common shape before comparison, and alerts are delivered separately from ingestion, with real-time updates pushed to the browser over websockets.
The lesson I took from that project is that alert quality is decided upstream: when an alert is wrong, the cause is usually identity or normalisation, not the notification code. I wrote more about the ingestion side in building a resilient scraping system for government data.
When you don't need all of this#
Not every scraping project needs a full change detection pipeline. Skip most of it if:
- You only need the latest snapshot. A price-comparison page that shows current values can simply overwrite rows.
- The source already publishes changes. Many APIs and open data portals expose an updated-at field, a change feed or conditional requests via ETags. Use those first. Scraping and diffing is the fallback, not the default.
- The dataset is small and nobody acts on changes in real time. A nightly CSV export compared in a spreadsheet may be all the business needs.
Build the full pipeline when changes drive decisions, money or compliance, and when users will judge the product by whether its alerts are right.
Frequently asked questions#
How often should a scraper run for change detection?#
Match the schedule to how quickly the source changes and how quickly users need to act. Compliance and pricing data often justify runs every 15 to 60 minutes. Most other data is fine daily. Always respect the source's rate limits and terms.
Should I store every version of a record?#
Store the current normalised state plus a change log of field-level diffs. That lets you rebuild history without keeping a full copy of every record from every run. Keep raw payloads for a limited window for debugging.
What's the best way to detect changes in scraped data?#
Give each record a stable key, normalise its fields, hash the fields you care about, and compare hashes between runs. Compute a field-level diff only when the hash changes. This is fast, cheap to store and easy to explain to users.
How do I avoid duplicate alerts?#
Make alert delivery idempotent. Use a unique key per change and subscriber, and check it before sending. Combined with retry-safe queued jobs, this prevents duplicates even when jobs are retried.
Key takeaways#
- Compare records, not pages. Stable identity is the foundation of change detection.
- Normalise values at ingestion, and hash only the fields that matter to the business.
- Store field-level diffs in an append-only change log, and build alerts from that log.
- Treat removals with suspicion: confirm complete runs and require consecutive misses.
- Deliver alerts asynchronously and idempotently, with before-and-after values users can verify.
If you're planning a monitoring product or your current scraper's alerts have lost your users' trust, my web scraping and data extraction service covers exactly this, or you can book a call to talk it through.
Share this technical insight with your network
Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.
Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS
Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.
Related Technical Articles
View all articles →
How I Built a Resilient Web Scraping System for Complex Government Data Sources
A real-world look at how I designed a resilient web scraping system for Violerts to collect, normalize, and manage data from complex government sources.

Laravel Queue Architecture: Jobs That Survive Production
dispatch() is one line. Making that job safe to retry, cheap to scale, and impossible to double-run is the actual engineering problem.

API Idempotency Keys: Stopping Duplicate Writes in Laravel
A retried POST request isn't a safe assumption — it's a duplicate charge waiting to happen. Here's how idempotency keys make retries actually safe.

Building a Reliable Webhook Delivery System in Laravel
A webhook that fires once and hopes isn't an integration. It's a support ticket waiting to happen the first time a customer's server is down.
Have a complex technical project in mind?
Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

