How to Fix 403 Client Error: Forbidden for URL in Python with One Controlled Benchmark

403 Client Error: Forbidden for URL means Python reached the server, received a 403 response, and raised an HTTPError when the status was checked. The error confirms that the request was refused, but it does not identify which part of the setup caused the refusal.
If the resource requires access your account does not have, correct the token, role, session, or permission first. If the page is public or opens in a normal browser, use NodeMaven’s open-source proxy-benchmark to test which part of the request setup changes the result.
The repository is part of NodeMaven’s GitHub, which provides open-source tools for proxy, browser, and connection testing. Our general guide to 403 Forbidden errors covers broader fixes for the status code; this guide focuses on the Python Requests exception and a reproducible benchmark workflow.
Why does Python Requests raise this 403 error?
Python Requests can connect to the server and receive a response successfully, then raise requests.exceptions.HTTPError because raise_for_status() found a 403 status.
In this example, requests.get() receives the response. The exception occurs on the following line, when Requests checks its status:
A 403 confirms that Python received an HTTP response, unlike a connection failure or timeout. The status still does not show why the server refused the request.
The Requests documentation explains this behavior: raise_for_status() raises HTTPError after an unsuccessful response. From there, proxy-benchmark can test whether the refusal follows a specific part of the setup. It also records transport or harness failures that produce no target verdict, while its gateway-health command checks whether the proxy gateway accepts the connection.
We used the same comparison method in our browser automation benchmark. With proxy-benchmark, the study ran 3,816 Amazon product-search attempts across eight automation setups, while keeping the target, host, gateway configuration, session pattern, and test harness fixed.
The testing scripts, run rows, and analysis are public, so developers can inspect the method or rerun the comparison with their own proxy. Raw rows are stored under data/runs/, generated tables appear in RESULTS.md, and NOTEBOOK.md documents the analysis.
How to run proxy-benchmark
Run the free direct baseline
Start with a plain HTTP request to DuckDuckGo without a proxy. This baseline requires no proxy account or browser download and does not consume proxy traffic.
Activate the virtual environment in Windows PowerShell:
On macOS or Linux:
Install the core requirements and run five direct attempts:
The command prints the test plan, one line for each attempt, and a final summary. It confirms that the harness runs correctly before you add a proxy or browser engine.
Check the available commands and preview a benchmark before sending traffic:
The dry run reports unavailable engines and their installation requirements. It also estimates traffic before the benchmark starts.
Add a proxy you already use
The repository includes a custom provider definition for a proxy with one host, port, login, and password. Add the following values to .env:
Check the gateway before sending traffic to a scraping target:
The health check opens ten CONNECT requests without contacting a scraping target. It can show whether the gateway is reachable and whether it accepts the supplied credentials.
Once the gateway responds, run the plain HTTP client through the proxy:
The basic custom definition does not contain country or sticky-session parameters. You can still compare clients, browser engines, targets, and direct versus proxied connections, but the benchmark cannot test rotation or country selection unless those settings are described in a provider definition.
If the credentials or username format are unclear, our proxy authentication guide covers common authentication patterns. The curl proxy guide provides standalone HTTP, HTTPS, and SOCKS examples.
Add a provider with country or session parameters
If the provider encodes countries, sessions, or other settings in the proxy username, copy the repository template:
Describe the provider’s parameter names and separators in that file. Then add its credentials to .env:
To compare it with NodeMaven, you must also add the NodeMaven credentials required by the included nodemaven provider definition. Then run:
This command also requires Patchright and its browser dependency to be installed. Run –dry-run first to confirm that the engine and both provider configurations are available.
The repository ships with provider definitions for NodeMaven, Decodo, Oxylabs, a Bright Data ISP zone, and a generic custom proxy. Each definition states whether its configuration was measured or transcribed from provider documentation. The dry run displays that provenance before the matrix starts.
Compare the same browser with and without a proxy
The :direct suffix runs an engine without the proxy inside the same benchmark matrix. This provides a direct control during the same test window as the proxied attempts.
To compare a custom proxy with a direct connection, run:
python scripts/benchmark.py –providers custom –engines chromium,chromium:direct,camoufox –targets amazon_search –queries 100 –batch 10
In this matrix:
- chromium runs through the custom proxy.
- chromium:direct runs the same engine without the proxy.
- camoufox adds a separate browser-engine comparison through the proxy.
If direct Chromium receives the requested page while proxied Chromium receives a block, the connection path is the variable that changed. The result does not automatically identify the provider or exit IP as the final cause, but it shows where to continue testing.
Compare HTTP clients and browser engines
The repository supports plain HTTP clients, unmodified browsers, patched browsers, CDP drivers, and browser automation frameworks.
| Engine | Role in the benchmark |
| http | Plain Requests client without a browser or JavaScript |
| curlcffi | Scriptless client using a Chrome-like TLS handshake |
| chromium | Unmodified Playwright Chromium control |
| patchright | Playwright-based Chromium with automation indicators patched |
| rebrowser | Playwright-based Chromium with the Runtime.enable leak patched |
| camoufox | Patched Firefox controlled through Playwright |
| cloak | Patched Chromium exposed through a Playwright browser |
| obscura | Rust browser controlled over CDP |
| seleniumbase | SeleniumBase UC mode using the installed Chrome browser |
| zendriver | Installed Chrome controlled through raw CDP |
| botasaurus | Installed Chrome controlled through a scraping framework and raw CDP |
Run only the engines available in your environment. The dry run names missing dependencies and provides the corresponding installation command.
The plain HTTP client provides a low-traffic baseline, while stock Chromium acts as the unmodified browser control. The remaining engines let you test whether a different client or browser implementation changes delivery.
How proxy-benchmark judges each response
proxy-benchmark evaluates registered targets using target-specific content markers instead of treating every HTTP 200 response as a success.
| Verdict | Meaning in the repository |
| ok | The target returned the requested page |
| captcha | A challenge appeared before the requested result |
| consent | A consent or cookie wall remained in front of the result |
| block | The target refused the address |
| throttle | Amazon returned its throttle page |
| empty | A body arrived without the requested result |
| error | The attempt did not complete, so no target verdict was produced |
This classification detects blocked pages that arrive with successful or inconsistent HTTP statuses. In the repository’s completed runs, the same Google reCAPTCHA page appeared with both 429 and 200 statuses. Amazon’s throttle page appeared with both 503 and 200 statuses.
A scraper that counts every 200 response as successful could therefore record a CAPTCHA, throttle page, or incomplete result as valid content.
Generate a report or inspect individual attempt rows after the benchmark:
Replace <stamp> with the timestamp from the relevant run filename. The report summarizes the comparison, peek.py displays individual rows, and calibrate.py measures transferred bytes for future traffic estimates.
Can proxy-benchmark test your own URL?
Yes, but you must register the target first. proxy-benchmark needs target-specific rules in nmbench/targets.py so it can distinguish the requested page from a CAPTCHA, consent wall, block, throttle page, or empty response. It cannot accurately classify an arbitrary URL from the URL alone.
Four targets carry the repository’s published benchmark results:
- google_serp
- amazon_search
- bing_serp
- ddg_serp
The registry also contains Walmart Search, Google Maps, Lazada, and Shopee, although these targets have fewer published runs. The ipinfo entry checks the connection path and exit address rather than page delivery.
To test another website, add it to nmbench/targets.py and define the content markers for each possible verdict. Build the rules from response bodies collected during real runs, since a successfully delivered page can still contain hidden challenge-related text.
What should you test after the benchmark identifies a variable?
If changing one variable changes the result, investigate that part of the setup next. For example, if the direct request works but the proxied request fails, check the proxy path. After making a change, rerun the same benchmark to confirm that the 403 no longer appears.
| What changed the outcome | What to test next |
| Direct versus proxied connection | Gateway health, exit type, session behavior, country, or provider |
| Plain HTTP versus browser | Client behavior, JavaScript execution, browser environment, or request flow |
| Browser engine | Run the leading and failing engines again while keeping the proxy, target, host, and time window fixed |
| Country | Repeat the same client and workload across the required locations |
| Host machine | Compare the runtime, browser build, network path, and machine-level environment |
| Provider | Repeat the interleaved provider matrix with the same target, engine, settings, and test window |
| Incomplete attempt with no target verdict | Inspect the saved failure reason, run gateway-health, and confirm that the selected engine can start |
A benchmark result identifies where the outcome changes, not every mechanism behind that change. Keep the other variables fixed while testing the affected part of the stack.
When the proxy changes the response
If the same client works directly but fails through the proxy, rerun the comparison with NodeMaven’s quality-filtered proxies before replacing the browser or rewriting the scraper.
Choose residential proxies for location targeting, rotation, and sticky sessions; mobile proxies for mobile-network IPs; or ISP proxies for static ISP-assigned addresses and longer sessions.
When connecting NodeMaven to your application, use the Python SDK or Rust SDK to build and validate the proxy configuration before sending a request. The SDKs catch unsupported or misspelled parameters locally, preventing country, session, or filtering settings from being silently ignored.
The benchmark also accepts proxies from other providers through its custom configuration. To compare provider-specific country or session settings, describe them in a TOML provider definition and check the planned matrix with –dry-run before sending traffic.
Run proxy-benchmark on your own stack
Clone nodemaven/proxy-benchmark and reproduce the failed request under controlled conditions. The repository includes the test harness, setup instructions, committed run data, and analysis commands needed to compare the results.
Start with the free direct baseline, configure the target and proxy path you want to test, and change one variable at a time. Each attempt is saved as a JSONL row so you can inspect which configuration produced the requested page, a block, a challenge, or an incomplete attempt.
For more developer tools, explore our selection of open-source GitHub repositories for web scraping.



