Skip to content
Backend & Architecture•12 min read•Published September 1, 2026

How I Built a Resilient Web Scraping System for Complex Government Data Sources

A real-world look at how I designed a resilient web scraping system for Violerts to collect, normalize, and manage data from complex government sources.

Aqib Javaid
Aqib Javaid
Senior Full-Stack Engineer
Resilient data pipeline connecting multiple public and government data sources

How I Built a Resilient Web Scraping System for Complex Government Data Sources#

When people hear that a system collects information from public or government websites, they often assume the difficult part is simply writing a scraper.

It isn't.

While working on Violerts, I had to solve a much bigger engineering problem: reliably collecting data from multiple public and government sources that all behaved differently.

The platform needed to work with sources that had inconsistent HTML structures, different search workflows, JavaScript-driven pages, rate limits, anti-automation protections, and changing website structures.

Some of the sources included New York City government websites and public information systems, including property and building-related portals such as BISWeb, along with other source-specific government and public data systems.

A scraper that worked perfectly one day could fail the next.

The real challenge was not collecting data once.

The challenge was building a production system that could continue collecting reliable data as external sources changed.

The Core Problem: Public Data Does Not Mean Easy Data Access#

Violerts needed to collect and process information from multiple public-facing and government data sources.

The problem was that there was no single integration pattern.

Every source had its own:

  • Website structure
  • Data format
  • Search flow
  • Request limits
  • Availability issues
  • Session behavior
  • Protection mechanisms
  • Update frequency

For example, one government website might expose information through relatively simple HTML pages, while another required navigating a multi-step search workflow before the relevant information could be retrieved.

A basic request such as:

php
$response = Http::get($url);

$html = $response->body();

can work for simple sources.

But when building a production data collection system, I could not assume every source would respond consistently to the same approach.

The architecture needed to answer questions such as:

  • What happens when a government website temporarily goes offline?
  • What happens when an HTML structure changes?
  • What happens when a request fails halfway through processing?
  • How do I avoid duplicate records?
  • How do I process multiple data sources without blocking the application?
  • How do I manage source-specific sessions and request behavior?
  • How do I prevent one failing source from stopping the entire pipeline?
  • How do I detect that a source has changed before the data quality is affected?

That changed how I approached the entire system.


Working With NYC Government Data Sources#

One of the interesting challenges in Violerts was working with New York City public and government information systems.

These systems were not designed as a single unified API ecosystem.

Different portals had different:

  • Navigation patterns
  • Search parameters
  • Record structures
  • HTML layouts
  • Session requirements
  • Update schedules

For example, property or building-related information could require source-specific logic to locate, extract, and interpret the relevant records.

That meant I could not build a generic scraper and expect it to work everywhere.

Instead, each source needed to be treated as its own integration.

The system had to understand:

text
Source A
    v
Search / Request Strategy A
    v
Source-Specific Parser A
    v
Normalized Record

Source B
    v
Search / Request Strategy B
    v
Source-Specific Parser B
    v
Normalized Record

This approach was much more maintainable than trying to create one large scraper containing logic for every website.


Why I Didn't Build One Giant Scraper#

One of the first architectural decisions was avoiding a single scraper responsible for every source.

That approach becomes difficult to maintain very quickly.

Instead, I treated each source as an independent integration.

Conceptually, the system looked like this:

text
Scheduler
    v
Queue Jobs
    v
Source-Specific Collectors
    v
Request / Session Management
    v
Validation Layer
    v
Data Normalization
    v
Deduplication
    v
Database
    v
Application

Each source-specific collector had one responsibility:

Retrieve and process data from one particular source and return it in a predictable format.

For example:

php
interface DataSource
{
    public function collect(): Collection;
}

Each source could then have its own implementation:

php
class GovernmentSourceCollector implements DataSource
{
    public function collect(): Collection
    {
        // Source-specific collection logic

        return collect();
    }
}

This separation made the system significantly easier to maintain.

If one source changed, I could fix that integration without affecting the others.


Managing Requests and Proxy Infrastructure#

Another important part of building a reliable collection system was request infrastructure.

When a system communicates with multiple external websites at scale, sending all traffic through a single network path can create reliability problems.

For appropriate and authorized use cases, I designed the collection workflow so that request routing could be managed independently from the scraper itself.

Conceptually:

text
Collector
    v
Request Manager
    v
Proxy / Network Routing
    v
External Source

This separation was important because the scraper should focus on:

  • What data needs to be collected
  • How the source should be parsed
  • How the data should be normalized

While the request layer handles operational concerns such as:

  • Connection reliability
  • Request timeouts
  • Network failures
  • Controlled request distribution
  • Session continuity where required
  • Source-specific request configuration

The key architectural decision was making this configurable rather than embedding network logic directly into every scraper.

A simplified structure could look like:

php
class SourceRequestService
{
    public function get(string $url)
    {
        return Http::timeout(30)
            ->retry(3, 5000)
            ->get($url);
    }
}

The actual production request strategy can then be configured independently based on the source and its supported access requirements.

This separation also made troubleshooting much easier.

When something failed, I could determine whether the problem was:

  • The external source
  • The network connection
  • The request configuration
  • The parser
  • The data itself

When Simple HTTP Requests Were Not Enough#

Some sources were straightforward.

Others were built around modern browser behavior, JavaScript rendering, interactive sessions, or multi-step workflows.

In those situations, a simple HTTP request was not always enough to reproduce a legitimate user workflow.

The important lesson was:

Don't force every source through the same collection method.

The strategy needed to match the source.

That could mean:

  • Official APIs where available
  • Structured public datasets
  • Standard HTTP requests for simple pages
  • Browser automation for legitimate interactive workflows
  • Source-specific session handling
  • Alternative supported integration methods

The collection layer was designed so that the application did not care how a source was retrieved.

It only cared about receiving a predictable, normalized result.


Handling Anti-Automation and CAPTCHA Challenges#

Some of the external sources introduced additional challenges such as anti-automation systems, JavaScript checks, or CAPTCHA-gated workflows.

This meant that the collection architecture could not simply assume every request would receive the expected page.

A production pipeline needs to identify different response states:

text
Successful Response
    v
Validate Data

Temporary Failure
    v
Retry Strategy

Unexpected Response
    v
Flag for Investigation

Access Restriction
    v
Review Supported Access Method

The important engineering challenge was not simply making a request.

It was understanding whether the request actually returned usable data.

For example, receiving an HTTP 200 response does not necessarily mean the collection succeeded.

The response might contain:

  • An error page
  • A challenge page
  • An empty result
  • An unexpected HTML structure
  • A maintenance page

That is why response validation became part of the collection pipeline.


Normalizing Completely Different Sources#

One of the biggest problems was that different sources did not describe information in the same way.

One might return:

text
Name
Date
Location
Reference Number

Another might return:

text
Organization
Created Date
Region
Case ID

The application should not need to understand every source.

So I created a normalization layer.

The responsibility of the collector was to retrieve source data.

The responsibility of the normalizer was to transform it into a consistent internal structure.

Conceptually:

php
$normalizedData = [
    'title' => $sourceData['name'],
    'reference' => $sourceData['reference_number'],
    'published_at' => $sourceData['date'],
    'location' => $sourceData['region'],
    'source' => 'government_source',
];

This meant the rest of Violerts could work with a predictable data model regardless of where the information originated.

That separation turned out to be one of the most important architectural decisions in the project.


Queues Were Essential#

Data collection should not block a user request.

Imagine one government source taking 30 seconds to respond while another source processes hundreds of records.

You do not want users waiting for that process.

Instead, collection tasks were moved into background jobs.

php
CollectSourceData::dispatch($source);

Laravel queues allowed the system to:

  • Process multiple sources independently
  • Retry temporary failures
  • Prevent long-running requests
  • Schedule recurring collection
  • Isolate source failures
  • Scale workers independently

A failure in one source should not stop the entire pipeline.

That principle shaped the queue architecture.

text
BISWeb Source       -> Success
Source B            -> Retry
Source C            -> Failed
Source D            -> Success

The system could continue processing while individual problems were isolated and investigated.


Designing for Failure Instead of Assuming Success#

One of the biggest lessons from Violerts was that external systems fail.

Government and public websites can be:

  • Slow
  • Temporarily unavailable
  • Under maintenance
  • Inconsistent
  • Changed without notice

So the pipeline needed to expect failure.

For temporary issues, retry strategies can help:

php
public $tries = 3;

public function backoff(): array
{
    return [60, 300, 900];
}

The goal was not to endlessly retry every failure.

The goal was to distinguish between:

  • Temporary failures
  • Network failures
  • Invalid responses
  • Permanent failures
  • Source structure changes
  • Unexpected access restrictions

Those categories require different responses.

A temporary network failure might justify a retry.

A changed HTML structure requires maintenance.

An unexpected response requires investigation.

That distinction makes a major difference when the system depends on external websites.


Deduplication Became a Data Problem#

Collecting data is only useful if the same record does not appear repeatedly.

Different collection runs could encounter:

  • Previously processed records
  • Slightly modified records
  • Duplicate listings
  • Updated source data

So the pipeline needed a way to determine whether something was genuinely new.

A simplified approach might use a unique source reference:

php
Record::updateOrCreate(
    [
        'source' => $source,
        'reference' => $reference,
    ],
    [
        'title' => $title,
        'published_at' => $publishedAt,
    ]
);

However, real-world deduplication becomes more complicated when a source does not provide a stable identifier.

In those situations, the system may need to compare multiple attributes or generate a consistent internal fingerprint.

The important part was making deduplication part of the pipeline architecture—not something added later.


The Biggest Maintenance Problem: Government Websites Change#

The first version of a collector is rarely the final version.

External websites change:

  • HTML structures
  • CSS selectors
  • Search flows
  • URLs
  • JavaScript behavior
  • Data formats

A scraper can be working perfectly and still fail when the source changes.

That is why observability matters.

I wanted the system to make failures visible.

Useful signals include:

  • Number of records collected
  • Number of failures
  • Processing duration
  • Source response errors
  • Unexpected response types
  • Sudden drops in collected records
  • Validation failures

For example, if a source normally returns hundreds of records and suddenly returns zero, that is not necessarily success.

Technically, the request may have completed successfully.

From a business perspective, something is probably wrong.

That difference between technical success and data success is critical in production data pipelines.


What I Learned Building the Violerts Data Pipeline#

The biggest lesson was that the difficult part of web scraping is rarely the first successful request.

The difficult part is everything that happens after that.

A production data collection system needs to survive:

  • Source changes
  • Temporary outages
  • Slow responses
  • Network problems
  • Invalid data
  • Duplicate records
  • Background processing failures
  • Changing external workflows

The architecture matters more than a clever scraper.

For Violerts, the solution was to build a system around separation of concerns:

text
Source Collection
        v
Request Management
        v
Validation
        v
Normalization
        v
Deduplication
        v
Background Processing
        v
Monitoring

Each layer had a clear responsibility.

That made the system easier to debug, maintain, and extend as new government and public data sources were introduced.


Final Thoughts#

Building Violerts taught me that web scraping and external data collection are fundamentally reliability and architecture problems.

The data might be publicly available.

The first scraper might take an afternoon.

But building a system that can reliably work across complex government websites, changing source structures, network challenges, and multiple data formats requires a completely different level of engineering.

The most important things I focused on were:

  • Treating every external source as an independent integration
  • Building source-specific collectors
  • Managing requests separately from parsing logic
  • Using queues for long-running collection tasks
  • Normalizing data early
  • Preventing duplicate records
  • Designing explicit failure and retry strategies
  • Monitoring data quality, not just HTTP responses
  • Building for changing external systems

The result was not just a collection script.

It became a more resilient data pipeline capable of processing fragmented external information and turning it into structured data that Violerts could actually use.

And that was the real engineering challenge: not scraping a website once, but building a system that can keep working when everything outside your application keeps changing.

Share this technical insight with your network

Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.

📁 Production Case Study

Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS

Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.

Related Technical Articles

View all articles →
✦ Let's Build Together

Have a complex technical project in mind?

Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

Need a web or software development partner?

Tell me what you’re building, what’s getting in the way, and where you need help. Whether you need a custom web application, SaaS platform, API integration, or full-stack development, I’ll give you a clear answer on scope, cost, and timeline usually within one business day.

AqibJavaid

Senior Full-Stack Engineer building backend systems, cloud infrastructure and product platforms for teams that need them to stay up.

Available for new projects

Get in touch

© 2026 Aqib Javaid. All rights reserved.

Built and maintained by Aqib Javaid