Skip to content
Data & Web Scraping•11 min read•Published September 28, 2026

Scraped Data Change Detection: From Snapshots to Reliable Alerts

How I design change detection for scraped data: stable record identity, normalisation, content hashing, field-level diffs and alerts users trust, with Laravel and MySQL examples from a real NYC compliance monitoring platform.

Aqib Javaid
Aqib Javaid
Senior Full-Stack Engineer
Diagram of scraped data change detection: snapshot, normalise, hash and diff, then alert

Scraping a page once is easy. The hard part comes on the second run, and every run after that: working out what actually changed since last time, and telling the right people without flooding them with noise. In my experience, most scraping systems don't fail at extraction. They fail at scraped data change detection. Every run looks like a wall of "new" records, alerts fire for whitespace edits, and nobody trusts the notifications anymore.

If you monitor prices, listings, permits, compliance records or competitor catalogues, the value is in the change, not the snapshot. A client doesn't pay to be told a property has 14 open violations. They pay to hear, within minutes, that it has a new one.

In this article I'll walk through how I design change detection for scraped data: stable record identity, normalisation, content hashing, field-level diffs, and turning those diffs into alerts people actually read. The examples use Laravel and MySQL, but the pattern works in any stack.

Why comparing raw scraped pages doesn't work#

The naive approach is to store the raw HTML or JSON from each run and compare it with the previous one. It breaks almost immediately:

  • Pages change without the data changing. Session tokens, timestamps, ad slots, "last updated" footers and randomised CSS class names make every fetch different.
  • Order isn't stable. Many sources return the same rows in a different order each time, so a line-by-line diff reports everything as changed.
  • Formatting drifts. "1,200", "1200" and "1200.00" are the same value. "Open " and "OPEN" usually are too.
  • Pagination shifts. One new record at the top pushes every other record to a different page, and a page-level diff sees churn across the whole dataset.

The fix is to stop comparing documents and start comparing records. Each scraped item becomes a row with a stable identity and a normalised set of fields, and change detection happens per record.

Step 1: Give every scraped record a stable identity#

Change detection is only as good as your answer to "is this the same thing I saw yesterday?" Before you store anything, decide on a natural key for each record type:

Source typeGood natural keyAvoid
Government registersOfficial record or case numberRow position, page number
Product cataloguesSKU or the store's product IDProduct title
Property dataParcel or building identifier plus record typeStreet address text
Job or listing boardsListing ID from the URLScrape timestamp

When the source has no ID, build a composite key from the fields that define the entity and never change, then hash it. Keep the key's inputs in one place in code, because changing them later makes every record look new.

php
function recordKey(string $source, array $raw): string
{
    // Only fields that identify the entity, never fields that describe it.
    return hash('sha256', implode('|', [
        $source,
        strtoupper(trim($raw['agency'])),
        preg_replace('/\s+/', '', $raw['case_number']),
    ]));
}

If you can't find a stable key at all, treat that as a design problem to solve with the client before building alerts. Any alert built on a guessed identity will produce false "new" and "removed" events.

Step 2: Normalise before you compare#

Normalisation turns the messy values a page shows into the canonical values your system stores. Do it once, at ingestion, and store only the normalised form for comparison.

Typical rules:

  1. Trim and collapse whitespace, and standardise case for enum-like fields such as statuses.
  2. Parse numbers and money into integers (store cents, not "$1,200.00").
  3. Parse dates into ISO 8601 in a single time zone.
  4. Map source vocabulary to your own: "CLOSED", "Closed - Resolved" and "RESOLVED" might all become resolved.
  5. Sort any list-valued field, so order changes don't register as edits.
  6. Drop fields that are volatile by nature: view counts, "fetched at" stamps, tracking parameters in URLs.

Keep the raw payload too, in a separate column or object store. When a client asks why an alert fired, you want to show exactly what the source said, not just your interpretation of it.

Step 3: Hash the fields you care about#

Once records are normalised, a content hash makes the "did anything change?" check cheap. Hash only the fields that matter to the business, in a fixed order:

php
function contentHash(array $normalised, array $trackedFields): string
{
    $subset = [];
    foreach ($trackedFields as $field) {
        $subset[$field] = $normalised[$field] ?? null;
    }

    return hash('sha256', json_encode($subset, JSON_UNESCAPED_UNICODE));
}

The tracked field list is a product decision, not just a technical one. For a compliance product, a change in status or penalty_amount matters. A change in the inspector's display name probably doesn't. Getting this list right is the single biggest lever on alert quality.

The storage model I use is simple: a current-state table keyed by the record key, and an append-only change log.

sql
CREATE TABLE scraped_records (
    record_key     CHAR(64) PRIMARY KEY,
    source         VARCHAR(64) NOT NULL,
    content_hash   CHAR(64) NOT NULL,
    data           JSON NOT NULL,
    first_seen_at  DATETIME NOT NULL,
    last_seen_at   DATETIME NOT NULL,
    missing_runs   INT NOT NULL DEFAULT 0
);

CREATE TABLE record_changes (
    id           BIGINT AUTO_INCREMENT PRIMARY KEY,
    record_key   CHAR(64) NOT NULL,
    change_type  ENUM('created', 'updated', 'removed') NOT NULL,
    diff         JSON NULL,
    detected_at  DATETIME NOT NULL,
    INDEX (record_key, detected_at)
);

On each run, the comparison is a lookup by key and a hash comparison. If the hash matches, you only bump last_seen_at, which is cheap enough to do for hundreds of thousands of records. MySQL's INSERT ... ON DUPLICATE KEY UPDATE handles the "seen again, unchanged" path in bulk.

Step 4: Compute field-level diffs, not just "changed"#

A hash tells you that something changed. An alert needs to say what changed. When the hash differs, compute a field-level diff against the stored data and write it to the change log:

php
function fieldDiff(array $old, array $new, array $trackedFields): array
{
    $diff = [];
    foreach ($trackedFields as $field) {
        $before = $old[$field] ?? null;
        $after  = $new[$field] ?? null;
        if ($before !== $after) {
            $diff[$field] = ['from' => $before, 'to' => $after];
        }
    }

    return $diff;
}

That diff is what makes an alert useful: "Status changed from open to resolved" is actionable, "Record 48213 updated" is not. It also gives you an audit trail for free, which matters when a customer disputes whether a change happened. I covered the wider pattern in my post on tamper-evident audit logging.

Handling removed records without false alarms#

Deletions are where most change detection systems embarrass themselves. A record missing from one run doesn't mean it was removed. The source might have timed out on page 37, a filter might have changed, or the site might have been briefly down.

A few rules keep removals honest:

  • Only mark removals after a complete run. If any page in the run failed, skip removal detection for that run entirely.
  • Require several consecutive misses. Increment missing_runs each time a record isn't seen, and emit removed only after two or three misses in a row. Reset the counter when the record reappears.
  • Sanity-check the run size. If a run returns 40% fewer records than the last successful one, treat it as a failed run, not a mass deletion. Alert the engineering team, not the customer.

The same instinct applies to "created" events. After you add a new source or change a key, the first run will mark everything as new. Run it silently as a baseline and start alerting from the second run.

Turning diffs into alerts people trust#

With a clean change log, alerting becomes a separate, much simpler problem. I treat each change row as an event and fan it out to subscribers asynchronously:

  1. The scraper writes change rows in the same transaction as the record update.
  2. A queued job picks up new change rows and matches them against each user's watch rules (a specific building, a record type, a threshold).
  3. Matching changes are delivered through the channels the user chose: in-app, email, real-time push or webhook.

Because delivery is queued, a slow email provider never slows down ingestion. Keep jobs small and retry-safe, as I described in Laravel queue architecture for production, and give each delivery an idempotency key built from the change ID and subscriber ID, so retries never send the same alert twice.

Two product details make a large difference to trust:

  • Batch low-urgency changes into digests. Not every change deserves an instant notification. Let users choose instant alerts for critical fields and a daily summary for the rest.
  • Show the before and after. Every alert should include the field-level diff and a link to the record, so the user can verify it in seconds.

Example from real delivery: Violerts#

On Violerts, a NYC compliance monitoring SaaS, the product's whole promise was change detection. The platform consolidates fragmented NYC municipal property data from multiple agencies into a single compliance view, and property owners and managers use it to find out quickly when something changes on a building they care about.

I led the modernisation of the Laravel backend and React frontend, including multi-agency data ingestion, asynchronous scraping and real-time alerts. The architecture follows the pattern in this article. Ingestion from public sources such as NYC Open Data and scraped agency pages runs as queued work on Redis, supervised with Laravel Horizon. Each source's records are normalised into a common shape before comparison, and alerts are delivered separately from ingestion, with real-time updates pushed to the browser over websockets.

The lesson I took from that project is that alert quality is decided upstream: when an alert is wrong, the cause is usually identity or normalisation, not the notification code. I wrote more about the ingestion side in building a resilient scraping system for government data.

When you don't need all of this#

Not every scraping project needs a full change detection pipeline. Skip most of it if:

  • You only need the latest snapshot. A price-comparison page that shows current values can simply overwrite rows.
  • The source already publishes changes. Many APIs and open data portals expose an updated-at field, a change feed or conditional requests via ETags. Use those first. Scraping and diffing is the fallback, not the default.
  • The dataset is small and nobody acts on changes in real time. A nightly CSV export compared in a spreadsheet may be all the business needs.

Build the full pipeline when changes drive decisions, money or compliance, and when users will judge the product by whether its alerts are right.

Frequently asked questions#

How often should a scraper run for change detection?#

Match the schedule to how quickly the source changes and how quickly users need to act. Compliance and pricing data often justify runs every 15 to 60 minutes. Most other data is fine daily. Always respect the source's rate limits and terms.

Should I store every version of a record?#

Store the current normalised state plus a change log of field-level diffs. That lets you rebuild history without keeping a full copy of every record from every run. Keep raw payloads for a limited window for debugging.

What's the best way to detect changes in scraped data?#

Give each record a stable key, normalise its fields, hash the fields you care about, and compare hashes between runs. Compute a field-level diff only when the hash changes. This is fast, cheap to store and easy to explain to users.

How do I avoid duplicate alerts?#

Make alert delivery idempotent. Use a unique key per change and subscriber, and check it before sending. Combined with retry-safe queued jobs, this prevents duplicates even when jobs are retried.

Key takeaways#

  • Compare records, not pages. Stable identity is the foundation of change detection.
  • Normalise values at ingestion, and hash only the fields that matter to the business.
  • Store field-level diffs in an append-only change log, and build alerts from that log.
  • Treat removals with suspicion: confirm complete runs and require consecutive misses.
  • Deliver alerts asynchronously and idempotently, with before-and-after values users can verify.

If you're planning a monitoring product or your current scraper's alerts have lost your users' trust, my web scraping and data extraction service covers exactly this, or you can book a call to talk it through.

Share this technical insight with your network

Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.

📁 Production Case Study

Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS

Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.

Related Technical Articles

View all articles →
✦ Let's Build Together

Have a complex technical project in mind?

Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

Need a web or software development partner?

Tell me what you’re building, what’s getting in the way, and where you need help. Whether you need a custom web application, SaaS platform, API integration, or full-stack development, I’ll give you a clear answer on scope, cost, and timeline usually within one business day.

AqibJavaid

Senior Full-Stack Engineer building backend systems, cloud infrastructure and product platforms for teams that need them to stay up.

Available for new projects

Get in touch

© 2026 Aqib Javaid. All rights reserved.

Built and maintained by Aqib Javaid