Skip to main content

The edge

The proxy in front of every environment — what it refuses on its own, what you can switch on, and taking a site offline.

Every request to your site passes through a proxy on the machine before it reaches your application. It terminates HTTPS, routes the hostname to the right environment, and applies the rules below. Some of them are the platform's and apply to everyone. The rest are yours: most are set once for the whole project on its Edge tab, and who may reach each environment is set on that environment's Routing tab.

In order

A request meets the checks in this order, and the first one that refuses it answers:

  1. HTTPS. Plain HTTP is redirected to HTTPS — see Domains.
  2. Probe guard — the platform's, what is refused by default for your kind of project, and your own probe rules, below.
  3. Blocked addresses, allowed addresses and blocked user agents.
  4. Password protection.
  5. Rate limit.
  6. Compression and HSTS, on the way back out.
  7. The canonical redirect, if the hostname is not the canonical one.

So a scanner is turned away before it is ever asked for a password, and an address you have blocked never gets as far as counting against a rate limit.

What is refused for everyone

The probe guard turns away requests for files no site should serve, before they reach your application: the ones an automated scanner asks every server on the internet for, hoping one of them answers. Hidden files, paths that try to climb out of the web root, and the well-known names of configuration and credential files — along with backup copies of those.

/.well-known/ is the exception, because certificates and app links live there. Matching ignores case.

A refused request gets a 404, not a 403. A 403 tells a scanner the file exists and is worth coming back for; a 404 says there is nothing here, which is also true of what it can reach.

It is deliberately narrow, and it is not a firewall that understands your application: it will not stop an attack on a URL your application really serves. What it does is take the noise away, so your logs and your PHP workers are spent on visitors.

You cannot turn it off or change what it covers. If a path your application genuinely needs is refused, tell us. You can add to it, though — below.

Refused by default

Three more, decided by what your project is, each on unless you turn it off on the project's Edge tab. They are refused the same way — a 404 and your probe page — and they apply on shared machines too, where your other proxy options do not.

Refused Where What it covers
Requests with no user agent Every project A request that names no user agent at all. Every browser, crawler, HTTP library — curl, Python's requests, Go, Node's fetch — and every payment or Git webhook sends one; a request without is a script written to scan
WordPress paths Projects that are not WordPress /xmlrpc.php; anything under /wp-admin, /wp-includes, /wp-json, /wp-content/plugins, /wp-content/themes and /wp-content/mu-plugins; wp-*.php anywhere (wp-login.php, wp-cron.php); /wp-admin/ and /wp-includes/ under a subdirectory; code files (.php, .phtml, .phar) under /wp-content/; and PHPUnit's vendor/phpunit/ and eval-stdin.php
.php Projects that run Node.js or Go A .php, .phtml or .phar file anywhere in the path, /index.php/admin included

Not refused, because real sites serve them: /wp-content/uploads/ (a site moved off WordPress keeps those links), /vendor/ as a whole (Laravel publishes its packages' files there), /config.json (single-page applications load one) and /robots.txt.

SQL injection

The probe guard also knows the shapes of SQL injection a scanner tries — UNION SELECT, ' OR 1=1, SLEEP(5), ; DROP TABLE — in the address: the path and the query string, never what a form posts. On the Edge tab, under Refused by default, choose what it does with them:

Setting What happens
Count only The default. The request reaches your site, and the Metrics tab says how many carried it
Refuse Answered as a probe, a 404 with your probe page, and counted as refused
Off Neither

Count first, and look at what was counted: if it is all scanners, refuse. It stops a scanner working through its list, not somebody who has studied your site — queries with bound parameters are what keep a database safe.

Test a path on the Edge tab says when one of these is what refuses it, and which switch to turn off. Turn one off if a client of yours really does call your site without a user agent, or if your site serves a path from the list.

Your own probe rules

On the project's Edge tab, under Probe rules: paths your site never serves and scanners ask it for anyway — /xmlrpc.php on a WordPress that does not use it, an admin path you moved, backups left lying around. They are refused the same way as the platform's — a 404 and your probe page — on every environment, after the platform's own rules.

Kind Example Matches
Exact paths /xmlrpc.php That path, ignoring case and a trailing slash
Path prefixes /old-admin/ Everything beneath it, ignoring case
Globs *.bak, /wp-json/*/users * within one segment, ** across segments, ? one character, ignoring case. Without a slash it matches the last segment at any depth; with one, the whole path from the start
Patterns ^/api/v[0-9]+/debug A regular expression in RE2 syntax, anywhere in the path unless anchored, case-sensitive unless you add (?i)

One rule per line, at most 50 in all, each up to 200 characters. A rule that would refuse your home page — and with it the site — is refused, and so are the pattern features RE2 lacks: look-ahead, look-behind, back-references.

Test a path under the rules says which rule would refuse it, yours or the platform's, against the rules as saved.

Some rules worth having, by what they are for:

Kind Rule Refuses
Exact path /xmlrpc.php WordPress's XML-RPC endpoint, on a site that does not publish from an app
Exact path /user/register Drupal's sign-up form, on a site where nobody signs up
Path prefix /old-admin/ An admin area you moved, which bots still try
Path prefix /wp-json/wp/v2/users WordPress's user list, which hands out login names
Glob *.bak, *.old, *.orig Backup copies of any file, at any depth
Glob *.sql, *.sql.gz, *.zip Database dumps and archives left in the web root
Glob /wp-content/uploads/**/*.php A PHP file in uploads, which is never one you put there
Pattern ^/api/v[0-9]+/debug A debugging route in every API version
Pattern `(?i)/(phpmyadmin pma

Test a few real paths of your own site afterwards — a rule that is broader than meant refuses pages your visitors want, with a 404 that looks like a missing page.

A machine needs agent 0.5.136 or later to apply them. Until it has upgraded, which it does on its own, the tab names it and it keeps the platform's rules alone.

Your rules

Two places, because they answer two different questions:

Set on What Applies to
The project's Edge tab Blocked addresses, blocked user agents, rate limit, compression, HSTS, your own probe rules, and the pages the edge serves in place of your site Every environment of the project
An environment's Routing tab Allowed addresses, password protection, keeping a visitor on one machine That environment

Who is turned away and how fast a visitor may ask are decisions about the site, so they are made once. Who may reach a copy of it differs between copies: staging can be locked to the office while production is open.

The Edge tab is one form with one Save: the options and the pages are checked together, and if anything is refused nothing is kept.

Not on shared infrastructure. On a machine the platform shares between customers the proxy is not yours alone, so none of these apply there — neither the project's options nor an environment's own. The Edge tab names any environment they skip, and an environment's Routing tab says so. Your own machines, dedicated or shared between your own environments, have them all. The pages the edge serves apply everywhere.

Allowed and blocked addresses

One address or range per line, IPv4 or IPv6 — 203.0.113.7, 203.0.113.0/24, 2001:db8::/32.

  • Allowed addresses, per environment: when there are any, only those can reach the environment. Empty means everyone. This is how a staging site is kept to the office.
  • Blocked addresses, for the whole project: refused on every environment even if its allow list permits them. That is what you want when the address is a scraper rather than a stranger.

A refused visitor gets a 403.

# Allowed addresses on staging: the office, and the VPN everybody else uses
203.0.113.0/24
198.51.100.14

# Blocked addresses on the project: a scraper's network, IPv4 and IPv6
192.0.2.0/24
2001:db8:42::/48

Behind a CDN, the address is the one the CDN reports. Every request arrives from the CDN, so the rules are matched against the visitor address it passes on — the platform CDN's, or the header of the network you named on your domain — and only on requests that came from that CDN's published edge addresses. Anything connecting directly is judged by its own address, so the header cannot be forged past the rules. For a CDN that publishes no edge addresses — Akamai, or one the platform has no entry for — the header is never believed: a block list stops nobody and an allow list stops everybody. Domains has the details.

Blocked user agents

One per line, matched anywhere in the user agent and ignoring case — GPTBot is enough. A match gets the same 403.

This is for crawlers that name themselves honestly. Scanners send an ordinary browser's user agent, so blocking by name will not stop them — that is what the probe guard and the address lists are for.

# Crawlers gathering text to train language models
GPTBot
ClaudeBot
CCBot
Bytespider
meta-externalagent

# SEO tools crawling for somebody else's report
AhrefsBot
SemrushBot
MJ12bot

Leave Googlebot and bingbot alone unless you mean to leave the search results: blocking them is what that does.

Password protection

One user:hash per line, the format htpasswd writes:

htpasswd -nB reviewer

A plain password is refused. It would be readable by everyone who can open the Routing tab.

The line htpasswd -nB reviewer prints, once it has asked for the password, is what goes in the box:

reviewer:$2y$05$4kcvXm6qX1xJt8Lr2dYQ3e7oQ0Sx9uM3cB1tFhZ5u3pQe6rA2wC1y

One per person who reviews the site, so one can be taken away without telling everybody a new password.

It covers every hostname of the environment, the platform hostname included — a password that only guarded your own domain would leave the site open at its vallic.cloud address. The two things that pass without it are the challenge a certificate authority uses to issue your certificate, and the platform's own health check.

Use -B (bcrypt). The form accepts some older hash formats as well, but bcrypt is the one to rely on.

Rate limit

Requests per second a single visitor can average, and a Burst it may go over that by for a moment. Zero means no limit; a burst left at zero is the same as the rate. A client over the limit is answered 429 Too Many Requests until it slows down.

The address counted is your visitor's. With nothing in front, that is the address that connects. Behind our CDN, Cloudflare or Fastly it is the visitor address the CDN reports — believed only from that CDN's own edge servers, so somebody going round the CDN cannot pick an address to be counted as.

Reasonable starting points, to watch and adjust:

Site Requests per second Burst
A brochure site or a blog 10 50 — one page and its images arrive together
A shop, behind a CDN that serves the images 15 40
An API that apps call 5 10

Set it with the site's own traffic in mind: a visitor loading a page that pulls thirty images and scripts from your origin needs a burst of thirty to see it whole.

Behind any other CDN the platform has no way to tell the visitor's address from a forged one, so the limit counts per CDN edge server: everyone that server carries shares the budget. The page says so when it applies; set the limit with that in mind, or leave rate limiting to the CDN.

The limit is set once for the project and counted on each environment separately: a visitor's requests to staging do not use up their allowance on production.

Compression

Compress responses has the proxy compress what your application sends — gzip, Brotli or zstd, whichever the browser asks for. On unless you turn it off. A response your application or a CDN already compressed says so, and is passed on as it is: nothing is compressed twice.

HSTS

HSTS max-age, in seconds tells browsers to refuse plain HTTP for your site for that long. Zero, the default, does not send the header.

Browsers remember it, so a value set by mistake outlives the mistake. 604800 (a week) is a sensible first step; raise it once you are sure.

Value For
300 Five minutes: trying it out, on a site you can afford to be wrong about
604800 A week: the first real step
31536000 A year: once every hostname of every environment has served HTTPS for a while
sent without includeSubDomains or preload, and there is no setting for
either.

Keeping a visitor on one machine

Only offered where an environment answers from more than one web machine. See the note on the form: it is for an application that keeps sessions on local disk and cannot move them, and it costs you some of what the second machine was for.

Taking a site offline

Show the offline page instead of the site, on the Routing tab. Visitors get a 503 with a Retry-After, which tells search engines the outage is temporary rather than the page gone. You can add a one-line note for visitors.

The site keeps running behind it. Deploys, the shell and restores all still work, so this is the switch for work you would rather nobody watched — a large migration, a restore, a content freeze.

There is no bypass: no address, cookie or hostname that sees the site while it is offline. Check your work on staging, or turn the page off again to look.

It needs the Admin role. Deploys do not use it: a deploy never takes your site offline on its own.

Your own pages

On the project's Edge tab, under Pages the edge serves: what a visitor sees when the edge answers instead of your site — while it is offline or being deployed (503), when a request is blocked (403) or refused as a probe (404), and when your application does not answer (502) or answers too slowly (504). Leave one empty and visitors get the platform's page, which is branded Vallic.

Each is complete HTML, served as written. It must not load anything from the site it stands in for — that is what is unreachable — so inline the styles and embed any image. A page is sent to every machine the project runs on, so each has a size limit, and the form says which is too large. The form lists the placeholders each page can use, such as the time and the region, and links to the platform's own page for comparison.

When your application is not answering

If your application returns a 502, 503 or 504, or does not answer at all, visitors get the platform's page for it instead of your application's response — or your own page for it. That includes a 503 your application sends deliberately: its own maintenance page is replaced.

When a change takes effect

Saving does not restart anything. The proxy's configuration is rebuilt from what is saved and sent to the machine, and the proxy reloads it without dropping connections.

Saving sends it: every machine the environment runs on is asked to rebuild its configuration, and picks that up within about a minute. That applies to everything on this page: the project's Edge tab, an environment's Routing tab, and taking a site offline.

Who can do what

Role needed
The project's Edge tab — blocks, rate limit, compression, HSTS, probe rules, pages Admin
An environment's own — allowed addresses, password, one machine per visitor Developer — Owner on a protected environment
Taking the site offline Admin

See Teams.

Next

  • Domains — the hostnames these rules apply to
  • CDN — caching and blocking by country, further out