Web scraping API vs Web Scraper Cloud: How to choose
September 28, 2026
data extraction, Web Scraper Cloud, scraping costs, data quality, Web scraping API
A scraping API, such as ScrapingBee's HTML API, fetches a page when your application sends a request. Web Scraper Cloud runs a tested sitemap that defines how to navigate a site and assemble records across pages. Both can use browser rendering and proxies. The practical difference is how much of the collection workflow your team builds and maintains.
For a few known URLs inside an existing application, an API can be simpler. For a recurring catalogue, directory or marketplace dataset, a reusable sitemap, hosted schedule and checks on the finished data can remove substantial work.
What does each product actually do?
Here, a scraping API means a third-party service that visits website pages on your behalf. It is different from a website owner's official data API, which exposes the owner's own data through supported endpoints.
A scraping API usually accepts a URL and options for rendering, proxies or extraction, then returns page content or selected fields. ScrapingBee is a useful example: its HTML API offers CSS or XPath extraction rules, AI extraction and JavaScript scenarios for clicks, waits and scrolling. It also has a separate CLI for crawling and batches. The exact features vary by vendor. Bright Data's Web Scraper API, for instance, offers discovery and validation, so it would be wrong to describe every scraping API as a raw HTML fetch.
With Web Scraper Cloud, you build and test a sitemap in the free browser extension first. The sitemap defines the links to follow, page interactions and fields to collect. Cloud then runs that compatible sitemap remotely. Its API launches an existing sitemap and lets an application retrieve job data; it does not invent an extraction workflow from any arbitrary URL. Both products can have APIs. The distinction is where the reusable, target-specific collection logic lives.
| Decision | Scraping API, using ScrapingBee as an example | Web Scraper Cloud |
|---|---|---|
| Starting point | Your application supplies URLs and request options. | You build and test a sitemap, then run it in Cloud. |
| Discovery and extraction | Your code or the vendor's extraction and crawling features identify pages and fields. | Selectors define navigation, pagination, interactions and record fields in one sitemap. |
| Repeat runs | Your application can schedule calls; ScrapingBee's CLI scheduling writes a cron entry on your own system. | Hosted schedules run the tested sitemap; the Cloud API can also launch jobs. |
| Checking and delivery | Your pipeline typically assembles and checks the dataset, although fuller API products may offer these functions. | Cloud provides job monitoring, collection-level quality checks and exports of completed data. |
| Natural fit | On-demand pages or an existing code-led data pipeline. | Repeated structured datasets from known website layouts. |
The table describes a division of work, not a universal feature limit. Compare the particular endpoint and companion tools you would buy, along with the work they leave to your team.
When a scraping API is the better fit
You already know the URLs. If your application has a current queue of product or property pages, it can request each one as it becomes due. Engineers can keep their existing deduplication, schema checks, storage and business rules. Building a separate sitemap for a small, changing list may add an unnecessary setup step.
You need a response inside an application flow. An enrichment service might receive one URL, fetch it and use the result immediately. A request-based API suits that interaction better than starting an asynchronous Cloud job and waiting for a finished dataset.
You want control at request level. ScrapingBee can return selected fields rather than only HTML, run page interactions and choose rendering or proxy configurations. Its Auto-Mode tries configurations until one succeeds and charges according to the tier used. Those options can spare a team from operating browsers and proxies while letting it retain control of its own crawler and pipeline.
An API is therefore a serious choice for a team that already has the surrounding system. The work to account for is how the team will find new URLs, join listing and detail data, schedule refreshes, detect incomplete runs and publish only accepted records. Some vendors offer additional products that cover parts of this work, so assess the complete product rather than the endpoint name.
Where Web Scraper Cloud does more of the dataset work
Imagine a retailer checking prices and availability across 20 public stores each morning. The required output is a dated table with store, product identifier, price, currency, availability and source URL. Some category pages already show every field. Others require opening product details. A page response alone does not say whether every product was discovered or whether a price belongs to the right item.
In Web Scraper, the team can build a sitemap for each different store layout. A sitemap can traverse listing pagination, produce multiple product records from one page and open detail pages where required. The team tests the selectors in the browser extension, then runs representative pages in Cloud using the appropriate driver and proxy. This is the sort of repeatable competitor price-monitoring workflow for which an inspectable sitemap is useful.
The Fast driver uses returned HTML and cannot run interaction-heavy sitemaps. Use FullJS when the necessary content or navigation requires browser execution. Neither driver guarantees access to a protected target. The configuration still needs repair if a website changes its layout or page state.
Once a sitemap works, Cloud can run it on a hosted schedule, inspect Failed and Empty pages, and apply a parser to format extracted values. Its quality rules can check minimum record counts, maximum Failed or Empty page percentages and required-field fulfilment. A quality failure can be reported even when the scraping job itself has finished. This helps expose a missing product category or a suddenly empty price column; it does not prove every value is correct. Sample the resulting rows against the source and apply your own acceptance rules before updating a production dataset.
Cloud can export the completed data to files or connected destinations, including Google Sheets and S3. An API-triggered job can fit into an engineering pipeline too. Its webhook signals the final job status and carries identifiers, not the scraped rows; the application retrieves those through the Cloud API. Your team still owns any downstream history, matching and final decision to publish.
A successful request can still produce the wrong data
Rendering, proxies and retries help with page access, but a technically successful response may be a consent screen, a different region's price, an empty listing or a partial product card. A target 404 can also be a meaningful deletion rather than merely a failure. This is why a 200 OK response is not proof of useful data.
Web Scraper Cloud retries Failed and Empty pages without consuming additional URL credits for those automatic retries. ScrapingBee offers its own retries and proxy options, including Auto-Mode. Its published billing rules also include some target error statuses, despite the short marketing description of paying for successful requests. A billed API response, a completed Cloud job and a correct dataset are three different outcomes.
Test the exact page types, region and access settings you plan to use. Check whether collection is appropriate under the target's terms, privacy obligations and other applicable rules; technical access alone does not settle that question.
How do the pricing models compare?
A Cloud URL credit represents a page loaded, whether it yields one record, many records or none. On Scale, Web Scraper charges for concurrent scraper capacity and has unlimited URL credits. ScrapingBee instead includes a monthly pool of API credits: a classic request without JavaScript uses one credit, while its default JavaScript-rendered request uses five. Premium proxy and AI options can raise that amount. One loaded page, one API call, one credit and one output record are different units.
Web Scraper's Scale calculator displays a $200 monthly total for two concurrent scrapers. Its calculator estimates 4.3 million Fast or 2.2 million FullJS URLs per month with those two scrapers. Those figures are estimated throughput, not URL allowances or guarantees. The table normalises each published plan at its own full estimated capacity or credit allowance, using US-dollar inputs checked on 29 September 2026. It assumes standard access, no AI extraction and no add-on proxy cost; annual billing is a separate option.
| Page processing | Web Scraper Scale, two scrapers | ScrapingBee |
|---|---|---|
| Without JavaScript | Fast: $200 monthly total; estimated 4.3m loaded pages. At that utilisation: $0.047 per 1,000 pages or $46.51 per 1m pages. | Startup: $99/month for 1m credits, enough for 1m classic non-JS calls. At full credit use: $0.099 per 1,000 billed calls or $99 per 1m calls. |
| With JavaScript | FullJS: $200 monthly total; estimated 2.2m loaded pages. At that utilisation: $0.091 per 1,000 pages or $90.91 per 1m pages. | Business+: $599/month for 8m credits, enough for 1.6m classic JS calls at five credits each. At full credit use: $0.374 per 1,000 billed calls or $374.38 per 1m calls. |
For example, the Fast figure is $200 ÷ 4.3 million × 1,000, rounded to three decimal places. The FullJS figure uses 2.2 million instead. These are effective costs at the calculator's estimated full utilisation, not per-page tariffs. If two Cloud scrapers process only one million pages in a month, the same $200 fixed capacity cost works out to $0.20 per 1,000 pages. Actual throughput depends on response times, delays, interactions, access failures and how continuously the jobs run.
The API figures divide plan spend by the number of calls its credit allowance supports in each configuration. A credit allowance is not a promise of correct records, just as a Cloud capacity estimate is not a guaranteed page count. The table compares different maximum volumes and operating models; it is not a like-for-like performance benchmark or a cost per million accepted records. An API's lower entry price can be attractive for small or irregular work, while a well-utilised capacity plan can have a lower effective page cost. Measure loaded listing and detail pages, valid records per page, proxy needs and engineering time on the actual target before choosing.
Run the same pilot before choosing
Test the hardest representative pages with both approaches, using the same fields, region and freshness target. Track four measures:
- Coverage: expected items and pages versus distinct records returned.
- Quality: required-field completion, duplicates, regional consistency and sampled values checked against the site.
- Consumption: loaded Cloud pages and proxy use versus API calls, configuration-weighted credits and billed statuses.
- Time and maintenance: elapsed run time, occupied capacity, retries and hands-on work when a layout changes.
An API plan's concurrent requests and Cloud's parallel tasks are different capacity measures. Each Cloud task runs a job and additional jobs queue; an API concurrency limit caps simultaneous calls from your client. Neither number alone predicts accepted records per hour. A pilot gives you the cost per accepted record and delivery time that a headline price cannot.
Which should you choose?
Choose a scraping API when your application already manages the URL list and dataset pipeline, you need request-level output, or the workload is small or irregular. ScrapingBee is a concrete example with extraction and browser options; other vendors may bundle more discovery and validation. Its lower entry price may suit a workload that would leave paid scraper capacity idle.
Choose Web Scraper Cloud when the recurring unit of work is a structured dataset across listings, pagination and detail pages. The sitemap makes navigation and fields inspectable before Cloud automation; hosted schedules, job-level quality checks and completed-data exports keep the repeated collection together. You still need to maintain the sitemap and validate actual values. Confirm compatibility first, especially for protected targets; Cloud is not intended for social platforms such as LinkedIn or large projects behind login.
An engineering team can still launch a tested Cloud sitemap through the API. If your goal is a recurring commercial dataset, build a small sitemap and run a representative Cloud job before committing to a schedule or price.