Build docker image and push to registry.bitdeals.org / main-build-job (push) Successful in 37s
An empty key switches the pseudonym off: no X-Client-Id is sent, and a rate limit downstream falls back to one bucket shared by every visitor. That is the one state nobody chooses on purpose and the easiest to reach by forgetting a line in a .env — so the entrypoint now fills the key in with `openssl rand -base64 32` when nothing else did. Nobody picks this value, nothing outside the container needs to know it, and no two deployments need the same one, which is what makes generating it the right default rather than a convenience. A fresh key per container start costs a reset of the downstream rate-limit buckets — invisible against a one-minute window — and makes pseudonyms from before and after unlinkable, which is the property the key exists for rather than a loss. Passing one explicitly still wins, for whoever wants pseudonyms stable across restarts or identical on two proxies. base64 and not hex, deliberately: HAProxy's hmac() decodes the key as base64, and hex would be accepted and silently decoded into something else — a usable key, but it would quietly cost the guarantee that a malformed one stops the container at configuration parsing. The base image's entrypoint is the haproxy binary with no shell in between, so the wrapper is the whole chain and execs the same binary with the same arguments. CMD is restated rather than inherited. Verified by building the image and checking the config in all three states: unset (wrapper reports it generated one, config parses), set and valid (wrapper silent, config parses), set and not base64 (`[ALERT] invalid args in converter 'hmac' : failed to parse key`, container refuses to start). READMEs updated in both languages, and "address" is spelled "IP address" throughout — it was never anything else.
238 lines
13 KiB
Markdown
238 lines
13 KiB
Markdown
# Intro
|
|
|
|
> Русская версия: [README.ru-RU.md](README.ru-RU.md)
|
|
|
|
[HAProxy](https://www.haproxy.org/) is a TCP/HTTP load balancer and reverse proxy. Here it is the public edge of a BitDeals site: it terminates TLS on 443, forwards everything to the web container, and routes ACME challenges to certbot.
|
|
|
|
HAProxy running in a docker container with a baked-in configuration.
|
|
|
|
This repository covers the docker deployment only. The image is `bitnami/haproxy` with one file copied into it.
|
|
|
|
# Usage
|
|
|
|
The container has two ports, **80** and **443**, and both are the public site.
|
|
|
|
The third channel is not a port: the HAProxy runtime API listens on a unix
|
|
socket at `/var/lib/haproxy/admin.sock`, on a volume shared with the
|
|
[certbot](https://git.bitdeals.org/private/certbot) container, which uses it to
|
|
install a renewed certificate into the running process without a restart. The
|
|
API is `level admin` and has no authentication, so who can open it is decided by
|
|
file permissions — see Notes.
|
|
|
|
There is one environment variable, `XFF_HMAC_KEY`, and it needs no setting: the
|
|
entrypoint generates one when it is absent. Everything else is in
|
|
`docker/haproxy.cfg`, which is copied into the image at build time, so changing
|
|
the routing means rebuilding and redeploying.
|
|
|
|
The certificate is read from `/usr/local/etc/haproxy/certificates/site.pem`,
|
|
mounted **read-only** from a volume shared with certbot. It must exist before
|
|
the container starts — see Notes.
|
|
|
|
## docker-compose
|
|
|
|
```yaml
|
|
services:
|
|
haproxy:
|
|
build:
|
|
context: https://git.bitdeals.org/private/haproxy.git
|
|
dockerfile: ./docker/Dockerfile
|
|
image: registry.bitdeals.org/haproxy
|
|
restart: unless-stopped
|
|
depends_on:
|
|
- nginx
|
|
- certbot
|
|
volumes:
|
|
- certificates:/usr/local/etc/haproxy/certificates:ro
|
|
- haproxy_admin:/var/lib/haproxy # runtime API socket — certbot only
|
|
ports:
|
|
- "80:80"
|
|
- "443:443"
|
|
|
|
volumes:
|
|
certificates:
|
|
haproxy_admin:
|
|
```
|
|
|
|
The two backends are named after the services they reach: `nginx:80` for the
|
|
site and `certbot:380` for ACME challenges. Both names have to resolve inside
|
|
the compose project, so those services must share a network with this one.
|
|
|
|
## docker cli
|
|
|
|
```sh
|
|
docker run -d \
|
|
-p 80:80 \
|
|
-p 443:443 \
|
|
-v certificates:/usr/local/etc/haproxy/certificates:ro \
|
|
registry.bitdeals.org/haproxy
|
|
```
|
|
|
|
Anything after the image name replaces the daemon's own arguments, so a config
|
|
check against a mounted file needs no new image:
|
|
|
|
```sh
|
|
docker run --rm -v "$PWD/docker/haproxy.cfg:/tmp/haproxy.cfg:ro" \
|
|
registry.bitdeals.org/haproxy -c -f /tmp/haproxy.cfg
|
|
```
|
|
|
|
## build and publish
|
|
|
|
A push to `main` builds and publishes the image
|
|
(`.gitea/workflows/build.yaml`), tagging it three ways: `<version>.<sha7>` to
|
|
deploy by, `<version>` to read, and `latest` for compose and Watchtower. A
|
|
nightly cron rebuilds from the same sources. By hand, when the registry
|
|
credentials are at hand:
|
|
|
|
```sh
|
|
docker build . --file docker/Dockerfile --tag registry.bitdeals.org/haproxy
|
|
docker push registry.bitdeals.org/haproxy
|
|
```
|
|
|
|
**The build context is the repository root**, not `docker/`: the Dockerfile
|
|
copies `./docker/haproxy.cfg`, so a context of `./docker` cannot see it and the
|
|
build fails on the `COPY`.
|
|
|
|
# Parameters
|
|
|
|
Container images are configured using parameters passed at runtime.
|
|
|
|
|Parameter|Function|
|
|
|:--------|:-------|
|
|
|-p 80|Plain HTTP. Redirects to HTTPS with a 301, except the ACME challenge path, which must stay reachable here for renewals to work|
|
|
|-p 443|HTTPS. Needs `site.pem` in the certificates volume before the container starts|
|
|
|-v /usr/local/etc/haproxy/certificates|Certificate directory, read-only. Only `site.pem` is read, at bind time. certbot writes it through the same volume mounted read-write at `/etc/certificates`|
|
|
|-v /var/lib/haproxy|Runtime API socket (`admin.sock`, `level admin`, **no authentication**). Mount it into certbot and nothing else — see Notes|
|
|
|-e XFF_HMAC_KEY|Optional, base64. The visitor's IP is replaced by an HMAC of it in `X-Client-Id` and never passed on. Leave it unset and the entrypoint generates one per container start; pass one only to keep pseudonyms stable across restarts or identical on two proxies — see Notes|
|
|
|
|
Routing, timeouts and TLS settings are not parameters: they live in
|
|
`docker/haproxy.cfg` and ship inside the image.
|
|
|
|
# Notes
|
|
|
|
- **`site.pem` must exist before the container starts.** `bind ... ssl crt` is
|
|
resolved while the configuration is parsed, so an empty volume is a fatal
|
|
start-up error, not a warning — HAProxy exits, and without a restart policy it
|
|
stays down. certbot writes a self-signed placeholder on its own first start
|
|
precisely to break this circle, which is why the service ships with
|
|
`restart: unless-stopped`; order it after certbot with `depends_on` in a
|
|
project that defines one.
|
|
- **A certificate installed over the runtime API lives in memory only.** That is
|
|
why the volume is mounted read-only here: `set ssl cert` + `commit ssl cert`
|
|
never write to disk. The file on the volume is certbot's copy, and it is what
|
|
HAProxy re-reads after a restart — so the two paths agree without HAProxy
|
|
needing write access.
|
|
- **The runtime API is a full administrative channel with no password.** Anyone
|
|
who can open it can install a different certificate and private key, redirect
|
|
a backend to another address, or take servers out of rotation — that is,
|
|
silently man-in-the-middle the site. Treat access to `admin.sock` as
|
|
equivalent to holding the TLS private key, and mount that volume into certbot
|
|
and nothing else. `expose-fd listeners`, which would additionally hand a
|
|
client of the socket the listening sockets themselves, is deliberately **not**
|
|
set: it exists for seamless reloads, which this image never performs.
|
|
- **A unix socket, because a port cannot be restricted.** `expose:` publishes
|
|
nothing to the host but restricts nothing either, and docker networks have no
|
|
per-port rules — so a TCP runtime API is open to every container sharing a
|
|
network, which here includes nginx, since HAProxy must be able to call *it*.
|
|
A socket on a volume is reachable only by containers that mount the volume,
|
|
and that is the whole access-control story. It also keeps the private key,
|
|
which crosses this channel on every renewal, off the network.
|
|
- **HAProxy needs write access to the socket's directory, not just the file.**
|
|
It binds by creating `<path>.<pid>.tmp` and renaming it over the target — so
|
|
the image creates `/var/lib/haproxy` owned by uid 1001, and docker carries
|
|
that ownership onto an empty named volume mounted there. The rename is also
|
|
why a stale socket left by a previous run is harmless. certbot connects as
|
|
root and is unaffected by the `mode 660`.
|
|
- **The redirect to HTTPS carries one exception, and it is load-bearing.**
|
|
Port 80 answers 301 for everything except `/.well-known/acme-challenge/`,
|
|
which Let's Encrypt validates over plain HTTP — redirect that and every
|
|
renewal stops. The rule is written above `use_backend` because that is the
|
|
order it runs in: `http-request` rules are evaluated before backend selection
|
|
whatever the file says, and HAProxy warns when the two disagree.
|
|
- **HSTS is one day, not the customary year.** It is a one-way door: a browser
|
|
that has seen the header refuses plain HTTP to this host until it expires, and
|
|
nothing server-side can call that back. A day keeps a lapsed certificate
|
|
recoverable. Raise it in steps — 86400, 2592000, 31536000 — once renewals have
|
|
been seen to work. `includeSubDomains` and `preload` are deliberately absent:
|
|
the first binds names this proxy does not serve, the second is effectively
|
|
permanent.
|
|
- **Backend addresses are re-resolved, and that is not the default.** Both
|
|
`server` lines carry `resolvers docker`, so the `nginx` and `certbot` names
|
|
are looked up again while HAProxy runs. Without it a name is resolved once at
|
|
boot and kept for the life of the process, and a container recreated on a new
|
|
IP — which is what Watchtower does on every deploy — is never noticed.
|
|
`init-addr libc,none` is the other half: it lets HAProxy start when a backend
|
|
is not up yet, instead of refusing to parse a name it cannot resolve.
|
|
- **Logging goes to stdout, and `option dontlog-normal` makes it errors-only.**
|
|
`log stdout format raw local0` needs no syslog daemon — `docker logs` collects
|
|
it. A successful request writes nothing; a 503, a backend with no server, a
|
|
refused handshake do. Drop `dontlog-normal` deliberately if a full access log
|
|
is wanted, and understand that it is also what keeps the volume down.
|
|
- **There are two loggers, and forgetting the second one leaks IP addresses.**
|
|
`option httplog` is never used: its default format opens with `%ci:%cp`, which
|
|
would put every visitor's IP address into `docker logs` and undo the pseudonym
|
|
the frontends mint. A hand-written `log-format` puts the pseudonym in that first
|
|
field instead. The trap is `error-log-format`, which covers what happens
|
|
*before* a transaction exists — a refused TLS handshake, and TLS 1.2 is now
|
|
the floor — and whose default opens the same way. Both are set here. The
|
|
pseudonym is therefore computed by a `tcp-request connection` rule on accept,
|
|
in `sess` scope, because an http-phase rule would not have run yet when a
|
|
handshake fails.
|
|
- **Only the method and path are logged, never the query string.** `%{+Q}r`
|
|
would carry it, and a token that ever appeared in a URL would be written down
|
|
for as long as the log is kept.
|
|
- **The visitor's IP address stops here.** There is no `option forwardfor`:
|
|
`X-Forwarded-For` is deleted in both frontends and never filled in, so nothing
|
|
behind this proxy can log an IP it was never given. `X-Client-Id` carries a
|
|
pseudonym instead — HMAC-SHA256 of the IP address under `XFF_HMAC_KEY`. Being
|
|
one-to-one with the IP it is exactly as good a rate-limiting key, and without
|
|
the key it is not reversible. HMAC rather than a bare digest because IPv4 is
|
|
2^32 values and an unkeyed hash of an IP address is brute-forced in seconds.
|
|
- **`XFF_HMAC_KEY` is generated when absent, not left empty.** An empty key
|
|
disables the feature — no `X-Client-Id` at all, and a rate limit downstream
|
|
falls back to one bucket shared by every visitor, which is the one state
|
|
nobody chooses on purpose and the easiest to reach by forgetting a line in a
|
|
`.env`. So the entrypoint fills it in with `openssl rand -base64 32` when
|
|
nothing else did. Nobody picks this value, nothing outside the container needs
|
|
to know it, and no two deployments need the same one.
|
|
A fresh key per container start costs a reset of the downstream rate-limit
|
|
buckets — invisible against a one-minute window — and makes pseudonyms from
|
|
before and after unlinkable, which is the property the key exists for rather
|
|
than a loss. Pass a value explicitly only to keep pseudonyms stable across
|
|
restarts, or identical on two proxies.
|
|
The config still handles an empty key, because `haproxy.cfg` can be run
|
|
outside this image. And a value that is not valid base64 still stops the
|
|
container at configuration parsing — it cannot degrade quietly, which is also
|
|
why the generated one is base64 and not hex: hex would be accepted here and
|
|
silently decoded as base64 into something else.
|
|
- **Both deletes are unconditional.** `X-Forwarded-For` and `X-Client-Id` are
|
|
dropped whether or not a key is configured, so a header a client sent can
|
|
never be mistaken downstream for one this proxy minted. Same for
|
|
`X-Forwarded-Proto`, which each frontend sets to its own scheme rather than
|
|
passing on the client's claim.
|
|
- **The consumer must still be told to use it.** A downstream rate limit keyed
|
|
on the socket IP address — nginx's `$binary_remote_addr`, ДС's
|
|
`request.client.host` — sees this proxy's IP for every request and degenerates
|
|
to one shared bucket. It has to key on `X-Client-Id`, and trust that header
|
|
only from this proxy's IP; `frontend/docker/rate-limit.conf` in the bitdeals-ng
|
|
repository is the worked example.
|
|
- **TLS is pinned in `global`, not left to OpenSSL.** TLS 1.2 is the floor,
|
|
the cipher list is ECDHE-only in both ECDSA and RSA variants — certbot issues
|
|
ECDSA, the self-signed placeholder is RSA — and session tickets are off so
|
|
forward secrecy is not undone by a long-lived ticket key. `alpn h2,http/1.1`
|
|
on the bind offers HTTP/2 to browsers; the backend stays HTTP/1.1 and HAProxy
|
|
translates. **No HSTS header is sent**, deliberately: it would be premature
|
|
while port 80 still serves the site rather than redirecting, and it is hard to
|
|
take back once browsers have cached the policy.
|
|
- **`timeout http-request 10s` is what bounds the header phase**, and
|
|
`timeout client` cannot stand in for it: that one is an *inactivity* timeout
|
|
and resets on every byte received, so a client dripping a byte at a time holds
|
|
a connection open indefinitely. This one is absolute.
|
|
- **The process runs as uid 1001 and still binds 80 and 443.** That works
|
|
because Docker sets `net.ipv4.ip_unprivileged_port_start=0` in containers by
|
|
default; a host or runtime that restores the traditional value will make the
|
|
container fail to bind.
|
|
- **The base image is unpinned.** `FROM bitnami/haproxy` means `:latest`, and
|
|
the nightly rebuild cron picks up whatever that tag points at — a HAProxy
|
|
minor version can change under a build nobody triggered, and Watchtower then
|
|
rolls it out. Pin `FROM bitnami/haproxy:<version>` for reproducible builds.
|