Illia Nikitchenko

Home / Cases

Caddy in Docker, and the www certificate that wouldn't issue

I moved this site's web server off the host and into a container, so I can experiment on the VPS without risking the site. Then I put it on a domain with HTTPS, and the apex got a certificate in seconds while www kept failing.

Why move it at all

The VPS has two jobs: it serves this site, and it's a machine I want to try things on. With Caddy installed as a host package, those jobs overlap. The config lives in /etc/caddy, the content in /var/www/html, and the service is a systemd unit, so package upgrades, a second web server, firewall changes or Docker networking tests all touch the same system the site depends on.

The goal was for the site to depend only on one container and a few folders, so that anything else on the host can break without taking it down.

Starting point

PartBefore
HostDebian 13 VPS, SSH keys only, ufw allowing 22, 80, 443
Web serverCaddy from the official apt repository, run by systemd
Content/var/www/html
ExposurePlain HTTP on port 80, reached by IP. No domain yet.

Target layout

Browser nikitchenko.com VPS · Debian 13 ufw: 22, 80, 443 caddy container image caddy:latest restart: unless-stopped 80 / 443 ~/caddy/Caddyfile → /etc/caddy/Caddyfile ~/caddy/site → /srv (the site) caddy_data volume → /data (certificates) The container can be deleted at any time. Everything it needs lives outside it.
New content gets onto the server with scp into ~/caddy/site. The container picks it up without a restart.

Steps

1. Install Docker

curl -fsSL https://get.docker.com -o get-docker.sh
sudo sh get-docker.sh
sudo usermod -aG docker $USER   # then log out and back in
docker run hello-world

Adding my user to the docker group means I don't need sudo for every command, but it's effectively root access to the host, since anyone in that group can start a container that mounts /. That's acceptable on a single-admin VPS. I'd think twice about it on a shared one.

2. Free ports 80 and 443

sudo systemctl stop caddy
sudo systemctl disable caddy
sudo ss -tlnp | grep -E ':80|:443'   # no output = ports are free

Only one process can listen on a port, so the host's Caddy had to stop before the container could start. That leaves a gap of a few seconds with nothing serving the site. Here it didn't matter because nothing was live yet. On a real service I'd start the container on a spare port first, test it, and only then switch.

3. Keep config and content on the host

mkdir -p ~/caddy/site

First version of ~/caddy/Caddyfile, still serving HTTP by IP:

:80 {
    root * /srv
    file_server
}

/srv is a path inside the container. The next step maps the host folder onto it.

4. Run the container

docker run -d \
  --name caddy \
  --restart unless-stopped \
  -p 80:80 \
  -p 443:443 \
  -v ~/caddy/Caddyfile:/etc/caddy/Caddyfile \
  -v ~/caddy/site:/srv \
  -v caddy_data:/data \
  caddy:latest
FlagWhy
--restart unless-stoppedComes back after a crash or a reboot. Stays down only if I stop it myself.
-p 80:80 -p 443:443Host port → container port. 443 was added in step 5, when HTTPS came in.
-v ~/caddy/Caddyfile:…Config is a file on the host that I edit with nano. It isn't baked into the container.
-v ~/caddy/site:/srvContent. Deploying means copying files here, with no restart and no rebuild.
-v caddy_data:/dataNamed volume for certificates and the ACME account. Without it, every docker rm throws them away, and re-requesting them repeatedly runs into Let's Encrypt's rate limits.

Deploying the site itself, from the laptop:

# run on the laptop, where the files are
scp -r site/* user@my-vps:~/caddy/site/

5. Domain and HTTPS

At the registrar, I replaced the default ALIAS record on the apex (it pointed at the registrar's parking page) with an A record pointing at the VPS. I also checked it with dig +short nikitchenko.com instead of trusting the browser, which kept showing the registrar's proxy error for a few minutes from cache.

Then I swapped the Caddyfile's :80 for real names, which is all Caddy needs to request and renew certificates on its own:

nikitchenko.com, www.nikitchenko.com {
    root * /srv
    file_server
}

I recreated the container with port 443 and the caddy_data volume (the command in step 4).

Symptom

nikitchenko.com got a certificate within seconds. www.nikitchenko.com failed both challenge types Caddy tried (log trimmed):

challenge failed  identifier=www.nikitchenko.com  challenge_type=tls-alpn-01
  Cannot negotiate ALPN protocol "acme-tls/1" for tls-alpn-01 challenge

challenge failed  identifier=www.nikitchenko.com  challenge_type=http-01
  207.207.210.126: Fetching https://nikitchenko-com.l.ink/.well-known/acme-challenge/…:
  Error getting validation data

Hypotheses I checked

HypothesisHow I checkedResult
Port 443 isn't reachableThe apex validated over tls-alpn-01 on the same port and the same containerRuled out
www is missing from the CaddyfileStartup log lists both names under "automatic TLS certificate management"Ruled out
www resolves to some other serverThe http-01 error names the host Let's Encrypt actually reached: the registrar's l.ink service at 207.207.210.x. That isn't my server.Confirmed

Root cause

A new domain at my registrar comes with default records: an ALIAS on the apex and a wildcard CNAME *.nikitchenko.com → uixie.porkbun.com, both pointing at the registrar's parking/link-in-bio service. I replaced the apex record and left the wildcard alone.

A wildcard matches any name that has no record of its own, and that includes www. So Let's Encrypt, which validates by connecting to wherever the name resolves, landed on the registrar's servers. Those servers don't speak acme-tls/1 and have never heard of Caddy's challenge token, so both challenges failed. Caddy and Docker were working correctly. The DNS record for that one name was wrong.

Fix

One record at the registrar:

A    www.nikitchenko.com    →  (VPS address)

A record for a specific name takes priority over a wildcard, so www now resolves to the VPS. I didn't have to do anything on the server, because Caddy retries failed issuance with backoff (60 s, then 120 s). On the retry it first ran the challenge against Let's Encrypt's staging environment, and only after that succeeded did it request the real certificate:

obtaining certificate     identifier=www.nikitchenko.com  ca=acme-staging-v02…
authorization finalized   authz_status=valid              ca=acme-staging-v02…
authorization finalized   authz_status=valid              ca=acme-v02…
certificate obtained successfully  identifier=www.nikitchenko.com

That staging step keeps a broken setup from burning through production rate limits. I also ran docker restart caddy "to force a retry", but the logs showed the certificate had already been issued about two minutes before I did.

Result

PartBeforeAfter
Web serverHost package + systemd unitOne container, restarts itself
Config/etc/caddy/Caddyfile~/caddy/Caddyfile, mounted in
Content/var/www/html~/caddy/site. Deploy = scp, no restart
TLSNone, HTTP by IPLet's Encrypt for apex and www, renewed automatically, stored in a volume

Smaller things that went wrong

What I learned

Loose ends

← Back to all cases