Начать пробный период
Назад

How to Fix 403 Client Error: Forbidden for URL in Python with One Controlled Benchmark

Обобщите эту статью с помощью предпочитаемого вами AI
Попробуйте наши премиум-прокси

Протестируйте наши премиум-прокси без ограничений по качеству.

  • Мобильные и резидентные прокси
  • Таргетинг на уровне ZIP
  • Статические и ротируемые IP
  • Встроенный фильтр качества
Попробовать сейчас

403 Client Error: Forbidden for URL means Python reached the server, received a 403 response, and raised an HTTPError when the status was checked. The error confirms that the request was refused, but it does not identify which part of the setup caused the refusal.

If the resource requires access your account does not have, correct the token, role, session, or permission first. If the page is public or opens in a normal browser, use NodeMaven’s open-source proxy-benchmark към test which part of the request setup changes the result.

The repository is part of NodeMaven’s GitHub, which provides open-source tools for proxy, browser, and connection testing. Our general guide to 403 Forbidden errors covers broader fixes for the status code; this guide focuses on the Python Requests exception and a reproducible benchmark workflow.

Why does Python Requests raise this 403 error?

Python Requests can connect to the server and receive a response successfully, then raise requests.exceptions.HTTPError because raise_for_status() found a 403 status.

In this example, requests.get() receives the response. The exception occurs on the following line, when Requests checks its status:

A 403 confirms that Python received an HTTP response, unlike a connection failure or timeout. The status still does not show why the server refused the request.

Зона Requests documentation explains this behavior: raise_for_status() raises HTTPError after an unsuccessful response. From there, proxy-benchmark can test whether the refusal follows a specific part of the setup. It also records transport or harness failures that produce no target verdict, while its gateway-health command checks whether the proxy gateway accepts the connection.

We used the same comparison method in our browser automation benchmark. With proxy-benchmark, the study ran 3,816 Amazon product-search attempts across eight automation setups, while keeping the target, host, gateway configuration, session pattern, and test harness fixed.

Зона testing scripts, run rows, and analysis are public, so developers can inspect the method or rerun the comparison with their own proxy. Raw rows are stored under data/runs/, generated tables appear in RESULTS.md, and NOTEBOOK.md documents the analysis.

How to run proxy-benchmark

Run the free direct baseline

Начните с plain HTTP request to DuckDuckGo without a proxy. This baseline requires no proxy account or browser download and does not consume proxy traffic.

Activate the virtual environment in Windows PowerShell:

На macOS или Linux:

Install the core requirements and run five direct attempts:

The command prints the test plan, one line for each attempt, and a final summary. It confirms that the harness runs correctly before you add a proxy or browser engine.

Check the available commands and preview a benchmark before sending traffic:

The dry run reports unavailable engines and their installation requirements. It also estimates traffic before the benchmark starts.

Add a proxy you already use

The repository includes a custom provider definition for a proxy with one host, port, login, and password. Add the following values to .env:

Check the gateway before sending traffic to a scraping target:

The health check opens ten CONNECT requests without contacting a scraping target. It can show whether the gateway is reachable and whether it accepts the supplied credentials.

Once the gateway responds, run the plain HTTP client through the proxy:

The basic custom definition does not contain country or sticky-session parameters. You can still compare clients, browser engines, targets, and direct versus proxied connections, but the benchmark cannot test rotation or country selection unless those settings are described in a provider definition.

If the credentials or username format are unclear, our руководство по аутентификации прокси covers common authentication patterns. The curl proxy guide provides standalone HTTP, HTTPS, and SOCKS examples.

Add a provider with country or session parameters

If the provider encodes countries, sessions, or other settings in the proxy username, copy the repository template:

Describe the provider’s parameter names and separators in that file. Then add its credentials to .env:

To compare it with NodeMaven, you must also add the NodeMaven credentials required by the included nodemaven provider definition. Then run:

This command also requires Patchright and its browser dependency to be installed. Run –dry-run first to confirm that the engine and both provider configurations are available.

The repository ships with provider definitions for NodeMaven, Decodo, Oxylabs, a Bright Data ISP zone, and a generic custom proxy. Each definition states whether its configuration was measured or transcribed from provider documentation. The dry run displays that provenance before the matrix starts.

Compare the same browser with and without a proxy

The :direct suffix runs an engine without the proxy inside the same benchmark matrix. This provides a direct control during the same test window as the proxied attempts.

To compare a custom proxy with a direct connection, run:

python scripts/benchmark.py –providers custom –engines chromium,chromium:direct,camoufox –targets amazon_search –queries 100 –batch 10

In this matrix:

  • chromium runs through the custom proxy.
  • chromium:direct runs the same engine without the proxy.
  • camoufox adds a separate browser-engine comparison through the proxy.

If direct Chromium receives the requested page while proxied Chromium receives a block, the connection path is the variable that changed. The result does not automatically identify the provider or exit IP as the final cause, but it shows where to continue testing.

Compare HTTP clients and browser engines

The repository supports plain HTTP clients, unmodified browsers, patched browsers, CDP drivers, and browser automation frameworks.

EngineRole in the benchmark
httpPlain Requests client without a browser or JavaScript
curlcffiScriptless client using a Chrome-like TLS handshake
chromiumUnmodified Playwright Chromium control
patchrightPlaywright-based Chromium with automation indicators patched
rebrowserPlaywright-based Chromium with the Runtime.enable leak patched
camoufoxPatched Firefox controlled through Playwright
cloakPatched Chromium exposed through a Playwright browser
obscuraRust browser controlled over CDP
seleniumbaseSeleniumBase UC mode using the installed Chrome browser
zendriverInstalled Chrome controlled through raw CDP
botasaurusInstalled Chrome controlled through a scraping framework and raw CDP

Run only the engines available in your environment. The dry run names missing dependencies and provides the corresponding installation command.

The plain HTTP client provides a low-traffic baseline, while stock Chromium acts as the unmodified browser control. The remaining engines let you test whether a different client or browser implementation changes delivery.

Как proxy-benchmark judges each response

proxy-benchmark evaluates registered targets using target-specific content markers instead of treating every HTTP 200 response as a success.

ВердиктMeaning in the repository
okThe target returned the requested page
КАПЧАA challenge appeared before the requested result
consentA consent or cookie wall remained in front of the result
blockThe target refused the address
throttleAmazon returned its throttle page
emptyA body arrived without the requested result
ошибкаThe attempt did not complete, so no target verdict was produced

This classification detects blocked pages that arrive with successful or inconsistent HTTP statuses. In the repository’s completed runs, the same Google reCAPTCHA page appeared with both 429 and 200 statuses. Amazon’s throttle page appeared with both 503 and 200 statuses.

A scraper that counts every 200 response as successful could therefore record a CAPTCHA, throttle page, or incomplete result as valid content.

Generate a report or inspect individual attempt rows after the benchmark:

Replace <stamp> with the timestamp from the relevant run filename. The report summarizes the comparison, peek.py displays individual rows, and calibrate.py measures transferred bytes for future traffic estimates.

Мочь proxy-benchmark test your own URL?

Yes, but you must register the target first. proxy-benchmark needs target-specific rules in nmbench/targets.py so it can distinguish the requested page from a CAPTCHA, consent wall, block, throttle page, or empty response. It cannot accurately classify an arbitrary URL from the URL alone.

Four targets carry the repository’s published benchmark results:

  • google_serp
  • amazon_search
  • bing_serp
  • ddg_serp

The registry also contains Walmart Search, Google Maps, Lazada, and Shopee, although these targets have fewer published runs. The ipinfo entry checks the connection path and exit address rather than page delivery.

To test another website, add it to nmbench/targets.py and define the content markers for each possible verdict. Build the rules from response bodies collected during real runs, since a successfully delivered page can still contain hidden challenge-related text.

What should you test after the benchmark identifies a variable?

If changing one variable changes the result, investigate that part of the setup next. For example, if the direct request works but the proxied request fails, check the proxy path. After making a change, rerun the same benchmark to confirm that the 403 no longer appears.

What changed the outcomeWhat to test next
Direct versus proxied connectionGateway health, exit type, session behavior, country, or provider
Plain HTTP versus browserClient behavior, JavaScript execution, browser environment, or request flow
Browser engineRun the leading and failing engines again while keeping the proxy, target, host, and time window fixed
СтранаRepeat the same client and workload across the required locations
Host machineCompare the runtime, browser build, network path, and machine-level environment
ПоставщикRepeat the interleaved provider matrix with the same target, engine, settings, and test window
Incomplete attempt with no target verdictInspect the saved failure reason, run gateway-health, and confirm that the selected engine can start

A benchmark result identifies where the outcome changes, not every mechanism behind that change. Keep the other variables fixed while testing the affected part of the stack.

When the proxy changes the response

If the same client works directly but fails through the proxy, rerun the comparison with NodeMaven’s quality-filtered proxies before replacing the browser or rewriting the scraper.

Выберите резидентские прокси for location targeting, rotation, and sticky sessions; мобильные прокси. for mobile-network IPs; or ISP прокси for static ISP-assigned addresses and longer sessions.

When connecting NodeMaven to your application, use the Python SDK или Rust SDK към build and validate the proxy configuration before sending a request. The SDKs catch unsupported or misspelled parameters locally, preventing country, session, or filtering settings from being silently ignored.

The benchmark also accepts proxies from other providers through its custom configuration. To compare provider-specific country or session settings, describe them in a TOML provider definition and check the planned matrix with –dry-run before sending traffic.

Беги proxy-benchmark on your own stack

Clone nodemaven/proxy-benchmark и reproduce the failed request under controlled conditions. The repository includes the test harness, setup instructions, committed run data, and analysis commands needed to compare the results.

Start with the free direct baseline, configure the target and proxy path you want to test, and change one variable at a time. Each attempt is saved as a JSONL row so you can inspect which configuration produced the requested page, a block, a challenge, or an incomplete attempt.

For more developer tools, explore our selection of open-source GitHub repositories for web scraping.

Часто задаваемые вопросы

raise_for_status() checks the received HTTP status and raises HTTPError when the response is 403. The exception shows the status and URL, but it does not explain why the server refused the request. Log the response body, headers, and final URL before calling raise_for_status().

Нет. A 403 does not identify the proxy as the cause. Compare the same client directly and through the proxy under the same conditions. If only the proxied request fails, investigate the gateway, exit IP, country, and session configuration. If both fail, check permissions, the client, the host machine, and the target itself.

Chrome and Requests use different browser capabilities, connection behavior, and TLS handshakes. Run the plain HTTP client and stock Chromium through the same proxy in proxy-benchmark. If Chromium receives the requested page while the HTTP client receives a block, continue testing the client path rather than replacing the proxy first.

Yes, after the target has been registered. Add the site to nmbench/targets.py and define content markers from real response bodies so the benchmark can distinguish the requested page from challenges, blocks, consent walls, and empty responses. You cannot paste an arbitrary URL into the tool and receive an accurate verdict without these rules.

Нет. The direct baseline requires no proxy account, and the generic custom provider accepts the host, port, login, and password of a proxy you already use. Providers that encode country or session settings in the username can be added through a TOML definition.

Нет. A User-Agent is only one signal in the request. Changing it may affect one target without resolving refusals caused by permissions, the connection path, browser behavior, the host machine, or another target-side decision. Test one change at a time and compare the returned content, not only the status code.

Вам также могут понравиться эти статьи

Этот сайт использует Файлы cookie чтобы улучшить ваш опыт. Продолжая, вы соглашаетесь на использование файлов cookie.