Web automation looks cheap until the first full month of invoices lands. A scraper that ran fine on a laptop now needs proxy bandwidth, headless browser instances, a job queue, retry logic, and somewhere to put the data.

Costs rarely explode because the tools are expensive. They explode because the pipeline wastes what it buys: duplicate requests, discarded sessions, and premium IP addresses doing work a cheaper resource could handle just as well.

The fixes below are unglamorous. That’s usually where the savings hide.

Stop paying residential prices for datacenter work

Proxy spend is the single largest line item for most automation teams, and it’s the easiest one to get wrong. Residential and mobile IPs are billed per gigabyte, and that meter runs whether or not the target site ever cared about the IP.

Plenty of targets don’t care. Public product catalogs, documentation sites, government registries, and most open APIs will serve a datacenter IP without so much as a CAPTCHA. Routing that traffic through residential bandwidth can cost 20 to 40 times more for identical results.

A sensible default: send every new target through the cheapest tier first, then escalate only when blocks start appearing. Most teams skip that step and pay for it monthly.

Building the ladder that way (datacenter first, residential second, mobile only for genuine app traffic) ties proxy cost to difficulty rather than volume. Pools built from the cheapest datacenter proxy options can carry the easy majority of targets, which leaves the expensive IPs reserved for sites that actively fight back.

Cache like each request costs money

Cache like each request costs money

Because it does. Any request made twice is billed twice, and response caching with proper ETag and If-Modified-Since handling cuts repeat traffic hard on sites that update slowly. A lot of catalogs change once a day at most.

Scheduling matters as much as caching. Google’s crawl budget documentation makes the argument from the other side of the wire: hitting a server more often than its content changes wastes resources for everybody, including the crawler doing the hitting.

Measure how often each target genuinely changes over a week, then set polling intervals to match. A price tracker running hourly against a retailer that reprices twice a day is spending 12 times what the job requires.

Deduplication belongs upstream of storage, too. Hashing normalized responses before writing them stops the warehouse from filling with 400 identical copies of a product page that nobody edited.

Drop the headless browser habit

Headless browsers are the second-biggest drain, and they get used far more often than they’re needed. Each Chrome instance holds 200 to 500 MB of RAM and burns CPU rendering fonts, images, and analytics scripts that no parser will ever look at.

Check the network tab before reaching for Playwright or Puppeteer. A surprising share of so-called JavaScript-heavy sites load their real content from a JSON endpoint that a plain HTTP client can hit directly, at roughly 1% of the compute cost.

When rendering really is required, block images, fonts, media, and third-party trackers at the request level. Trimming page weight by 60% trims bandwidth billing and instance hours at the same time.

Right-size the compute underneath

Scraping workloads are interruptible by design, which makes them a natural fit for discounted spare capacity. Azure Spot Virtual Machines and equivalent offerings from other providers run at steep discounts on the condition that they can be evicted, and a queue with sane retry logic absorbs evictions without losing work.

Autoscale on queue depth rather than on a fixed schedule. Workers sitting idle overnight at full price is one of the most common and most avoidable costs in the whole stack.

Egress deserves a look as well. Pulling raw HTML into one cloud region and processing it in another quietly adds a per-gigabyte transfer charge to work that could have happened where the data already sat.

Where this is heading

Anti-bot vendors keep raising the price of brute force, so the teams that stay cheap will be the ones that stop treating every target identically. Tiered proxies, honest cache headers, and browsers reserved for pages that truly need them are becoming the line between a pipeline that pays for itself and one that quietly doesn’t.

A good place to start is measuring cost per successful record instead of cost per gigabyte. That single number tends to expose which targets deserve more investment and which ones should be cut loose.