How I Built a Resilient Web Scraping System for Complex Government Data Sources#
When people hear that a system collects information from public or government websites, they often assume the difficult part is simply writing a scraper.
It isn't.
While working on Violerts, I had to solve a much bigger engineering problem: reliably collecting data from multiple public and government sources that all behaved differently.
The platform needed to work with sources that had inconsistent HTML structures, different search workflows, JavaScript-driven pages, rate limits, anti-automation protections, and changing website structures.
Some of the sources included New York City government websites and public information systems, including property and building-related portals such as BISWeb, along with other source-specific government and public data systems.
A scraper that worked perfectly one day could fail the next.
The real challenge was not collecting data once.
The challenge was building a production system that could continue collecting reliable data as external sources changed.
The Core Problem: Public Data Does Not Mean Easy Data Access#
Violerts needed to collect and process information from multiple public-facing and government data sources.
The problem was that there was no single integration pattern.
Every source had its own:
- Website structure
- Data format
- Search flow
- Request limits
- Availability issues
- Session behavior
- Protection mechanisms
- Update frequency
For example, one government website might expose information through relatively simple HTML pages, while another required navigating a multi-step search workflow before the relevant information could be retrieved.
A basic request such as:
$response = Http::get($url);
$html = $response->body();can work for simple sources.
But when building a production data collection system, I could not assume every source would respond consistently to the same approach.
The architecture needed to answer questions such as:
- What happens when a government website temporarily goes offline?
- What happens when an HTML structure changes?
- What happens when a request fails halfway through processing?
- How do I avoid duplicate records?
- How do I process multiple data sources without blocking the application?
- How do I manage source-specific sessions and request behavior?
- How do I prevent one failing source from stopping the entire pipeline?
- How do I detect that a source has changed before the data quality is affected?
That changed how I approached the entire system.
Working With NYC Government Data Sources#
One of the interesting challenges in Violerts was working with New York City public and government information systems.
These systems were not designed as a single unified API ecosystem.
Different portals had different:
- Navigation patterns
- Search parameters
- Record structures
- HTML layouts
- Session requirements
- Update schedules
For example, property or building-related information could require source-specific logic to locate, extract, and interpret the relevant records.
That meant I could not build a generic scraper and expect it to work everywhere.
Instead, each source needed to be treated as its own integration.
The system had to understand:
Source A
v
Search / Request Strategy A
v
Source-Specific Parser A
v
Normalized Record
Source B
v
Search / Request Strategy B
v
Source-Specific Parser B
v
Normalized RecordThis approach was much more maintainable than trying to create one large scraper containing logic for every website.
Why I Didn't Build One Giant Scraper#
One of the first architectural decisions was avoiding a single scraper responsible for every source.
That approach becomes difficult to maintain very quickly.
Instead, I treated each source as an independent integration.
Conceptually, the system looked like this:
Scheduler
v
Queue Jobs
v
Source-Specific Collectors
v
Request / Session Management
v
Validation Layer
v
Data Normalization
v
Deduplication
v
Database
v
ApplicationEach source-specific collector had one responsibility:
Retrieve and process data from one particular source and return it in a predictable format.
For example:
interface DataSource
{
public function collect(): Collection;
}Each source could then have its own implementation:
class GovernmentSourceCollector implements DataSource
{
public function collect(): Collection
{
// Source-specific collection logic
return collect();
}
}This separation made the system significantly easier to maintain.
If one source changed, I could fix that integration without affecting the others.
Managing Requests and Proxy Infrastructure#
Another important part of building a reliable collection system was request infrastructure.
When a system communicates with multiple external websites at scale, sending all traffic through a single network path can create reliability problems.
For appropriate and authorized use cases, I designed the collection workflow so that request routing could be managed independently from the scraper itself.
Conceptually:
Collector
v
Request Manager
v
Proxy / Network Routing
v
External SourceThis separation was important because the scraper should focus on:
- What data needs to be collected
- How the source should be parsed
- How the data should be normalized
While the request layer handles operational concerns such as:
- Connection reliability
- Request timeouts
- Network failures
- Controlled request distribution
- Session continuity where required
- Source-specific request configuration
The key architectural decision was making this configurable rather than embedding network logic directly into every scraper.
A simplified structure could look like:
class SourceRequestService
{
public function get(string $url)
{
return Http::timeout(30)
->retry(3, 5000)
->get($url);
}
}The actual production request strategy can then be configured independently based on the source and its supported access requirements.
This separation also made troubleshooting much easier.
When something failed, I could determine whether the problem was:
- The external source
- The network connection
- The request configuration
- The parser
- The data itself
When Simple HTTP Requests Were Not Enough#
Some sources were straightforward.
Others were built around modern browser behavior, JavaScript rendering, interactive sessions, or multi-step workflows.
In those situations, a simple HTTP request was not always enough to reproduce a legitimate user workflow.
The important lesson was:
Don't force every source through the same collection method.
The strategy needed to match the source.
That could mean:
- Official APIs where available
- Structured public datasets
- Standard HTTP requests for simple pages
- Browser automation for legitimate interactive workflows
- Source-specific session handling
- Alternative supported integration methods
The collection layer was designed so that the application did not care how a source was retrieved.
It only cared about receiving a predictable, normalized result.
Handling Anti-Automation and CAPTCHA Challenges#
Some of the external sources introduced additional challenges such as anti-automation systems, JavaScript checks, or CAPTCHA-gated workflows.
This meant that the collection architecture could not simply assume every request would receive the expected page.
A production pipeline needs to identify different response states:
Successful Response
v
Validate Data
Temporary Failure
v
Retry Strategy
Unexpected Response
v
Flag for Investigation
Access Restriction
v
Review Supported Access MethodThe important engineering challenge was not simply making a request.
It was understanding whether the request actually returned usable data.
For example, receiving an HTTP 200 response does not necessarily mean the collection succeeded.
The response might contain:
- An error page
- A challenge page
- An empty result
- An unexpected HTML structure
- A maintenance page
That is why response validation became part of the collection pipeline.
Normalizing Completely Different Sources#
One of the biggest problems was that different sources did not describe information in the same way.
One might return:
Name
Date
Location
Reference NumberAnother might return:
Organization
Created Date
Region
Case IDThe application should not need to understand every source.
So I created a normalization layer.
The responsibility of the collector was to retrieve source data.
The responsibility of the normalizer was to transform it into a consistent internal structure.
Conceptually:
$normalizedData = [
'title' => $sourceData['name'],
'reference' => $sourceData['reference_number'],
'published_at' => $sourceData['date'],
'location' => $sourceData['region'],
'source' => 'government_source',
];This meant the rest of Violerts could work with a predictable data model regardless of where the information originated.
That separation turned out to be one of the most important architectural decisions in the project.
Queues Were Essential#
Data collection should not block a user request.
Imagine one government source taking 30 seconds to respond while another source processes hundreds of records.
You do not want users waiting for that process.
Instead, collection tasks were moved into background jobs.
CollectSourceData::dispatch($source);Laravel queues allowed the system to:
- Process multiple sources independently
- Retry temporary failures
- Prevent long-running requests
- Schedule recurring collection
- Isolate source failures
- Scale workers independently
A failure in one source should not stop the entire pipeline.
That principle shaped the queue architecture.
BISWeb Source -> Success
Source B -> Retry
Source C -> Failed
Source D -> SuccessThe system could continue processing while individual problems were isolated and investigated.
Designing for Failure Instead of Assuming Success#
One of the biggest lessons from Violerts was that external systems fail.
Government and public websites can be:
- Slow
- Temporarily unavailable
- Under maintenance
- Inconsistent
- Changed without notice
So the pipeline needed to expect failure.
For temporary issues, retry strategies can help:
public $tries = 3;
public function backoff(): array
{
return [60, 300, 900];
}The goal was not to endlessly retry every failure.
The goal was to distinguish between:
- Temporary failures
- Network failures
- Invalid responses
- Permanent failures
- Source structure changes
- Unexpected access restrictions
Those categories require different responses.
A temporary network failure might justify a retry.
A changed HTML structure requires maintenance.
An unexpected response requires investigation.
That distinction makes a major difference when the system depends on external websites.
Deduplication Became a Data Problem#
Collecting data is only useful if the same record does not appear repeatedly.
Different collection runs could encounter:
- Previously processed records
- Slightly modified records
- Duplicate listings
- Updated source data
So the pipeline needed a way to determine whether something was genuinely new.
A simplified approach might use a unique source reference:
Record::updateOrCreate(
[
'source' => $source,
'reference' => $reference,
],
[
'title' => $title,
'published_at' => $publishedAt,
]
);However, real-world deduplication becomes more complicated when a source does not provide a stable identifier.
In those situations, the system may need to compare multiple attributes or generate a consistent internal fingerprint.
The important part was making deduplication part of the pipeline architecture—not something added later.
The Biggest Maintenance Problem: Government Websites Change#
The first version of a collector is rarely the final version.
External websites change:
- HTML structures
- CSS selectors
- Search flows
- URLs
- JavaScript behavior
- Data formats
A scraper can be working perfectly and still fail when the source changes.
That is why observability matters.
I wanted the system to make failures visible.
Useful signals include:
- Number of records collected
- Number of failures
- Processing duration
- Source response errors
- Unexpected response types
- Sudden drops in collected records
- Validation failures
For example, if a source normally returns hundreds of records and suddenly returns zero, that is not necessarily success.
Technically, the request may have completed successfully.
From a business perspective, something is probably wrong.
That difference between technical success and data success is critical in production data pipelines.
What I Learned Building the Violerts Data Pipeline#
The biggest lesson was that the difficult part of web scraping is rarely the first successful request.
The difficult part is everything that happens after that.
A production data collection system needs to survive:
- Source changes
- Temporary outages
- Slow responses
- Network problems
- Invalid data
- Duplicate records
- Background processing failures
- Changing external workflows
The architecture matters more than a clever scraper.
For Violerts, the solution was to build a system around separation of concerns:
Source Collection
v
Request Management
v
Validation
v
Normalization
v
Deduplication
v
Background Processing
v
MonitoringEach layer had a clear responsibility.
That made the system easier to debug, maintain, and extend as new government and public data sources were introduced.
Final Thoughts#
Building Violerts taught me that web scraping and external data collection are fundamentally reliability and architecture problems.
The data might be publicly available.
The first scraper might take an afternoon.
But building a system that can reliably work across complex government websites, changing source structures, network challenges, and multiple data formats requires a completely different level of engineering.
The most important things I focused on were:
- Treating every external source as an independent integration
- Building source-specific collectors
- Managing requests separately from parsing logic
- Using queues for long-running collection tasks
- Normalizing data early
- Preventing duplicate records
- Designing explicit failure and retry strategies
- Monitoring data quality, not just HTTP responses
- Building for changing external systems
The result was not just a collection script.
It became a more resilient data pipeline capable of processing fragmented external information and turning it into structured data that Violerts could actually use.
And that was the real engineering challenge: not scraping a website once, but building a system that can keep working when everything outside your application keeps changing.
Share this technical insight with your network
Share to LinkedIn or Facebook with key takeaways, featured media, and direct links.
Case Study: Violerts - Enterprise NYC PropTech Compliance & Violation Monitoring SaaS
Violerts is a PropTech SaaS platform that consolidates fragmented NYC municipal property data into a single compliance intelligence platform. I led the modernization of the React frontend and Laravel backend, building multi-agency data ingestion, GIS mapping, asynchronous scraping, real-time alerts, team collaboration, and Stripe-powered SaaS billing.
Related Technical Articles
View all articles →Beyond Documentation: What Production Engineering Teaches You About Software Architecture
Theoretical knowledge only takes you so far. Discover why real-world engineering begins in production navigating third-party latency, database scale, AI workflows, and architectural trade-offs.

Web Development Frameworks: A Practical Guide to Choosing the Right Framework in 2026
Choosing the right web development framework can shape your application's performance, scalability, maintainability, and long-term cost. This guide compares modern frameworks and the factors that actually matter when selecting one for a real-world project.
Have a complex technical project in mind?
Available for full-stack engineering, performance audits, cloud deployments, and high-concurrency systems architecture.

